powering productive workplaces
Front page
InterviewService & Connectivity1h · 15:01 BST · 6 min read

Why Network Redundancy Does Not Always Prevent Outages

Greg LaBrie of WEI explains why secondary connections may not activate when a primary route degrades, leaving users offline despite built-in redundancy. He outlines the monitoring, documentation and failover practices needed to restore services faster

A backup connection only protects the business if the network can recognise the primary route has failed and switch traffic over fast enough.

That is not always straightforward. A circuit can be degraded enough to disrupt users without appearing completely unavailable to the network, preventing an automatic failover and leaving IT teams to intervene manually.

Greg LaBrie, VP and General Manager at WEI, describes this as a connection being “sick but not dead”.

“An outage of a primary connection oftentimes did not seamlessly transition over to the backup,” LaBrie told UC Today.

The result is that organisations can have redundancy in place but still face downtime while teams validate the fault, force traffic onto a secondary path and work to restore users.

Restore Users Before Fixing the Primary Circuit

LaBrie’s experience of managing network environments before joining WEI highlighted how easily a backup plan can become a manual recovery process.

In one role, primary and backup connectivity were in place alongside Cisco networking equipment and monitoring tools. But a primary connection did not always fail in a way that automatically triggered the secondary route.

A circuit could be impaired by routing issues or other problems without the physical connection going down. In those circumstances, the router might continue to treat the path as viable even while users were unable to access the services they needed.

The immediate priority was therefore not to diagnose and repair the primary connection. It was to establish which circuits were affected and restore users through the backup path.

“The first thing that we would do would be to validate which specific circuits were impacted and then immediately try to get the backup working,” LaBrie said.

That approach puts business continuity ahead of root-cause analysis. Once the backup connection is active, IT teams can investigate the primary route, work with the relevant provider and arrange a controlled switch back when it is safe to do so.

During an outage, teams could receive alerts at the same time as employees began reporting that the network was not working. Engineers might need to access a router, check the circuit state and force the primary interface down before the backup connection would take over.

That is a labour-intensive process at precisely the moment an organisation needs speed and clarity.

The Commercial Reality of Failover

Backup connectivity can also introduce a difficult commercial decision.

LaBrie said organisations may choose lower-cost backup options that incur charges when used. In an incident, IT teams can find themselves weighing the cost of activating a backup service against the impact of leaving users without connectivity.

“I know I can do backup, but does the company want backup now at $5 a minute or $10 a gigabyte or whatever it is that you were having to pay for that usage?” he said.

That decision needs to be made before an incident, not during one.

Organisations should understand how their backup connectivity is priced, who has authority to activate it and which services or groups of users should take priority. If a backup link is expensive or capacity-constrained, the business must decide in advance whether it is intended to support every workload or only critical operations.

The cost of backup capacity should also be weighed against the impact of downtime. A connectivity outage can affect employees, cloud applications, customer-facing services and operations across branch locations. The price of using an alternative route may be easier to justify when the operational consequences are clear.

Runbooks should set out escalation paths, decision-makers and technical steps, so that the response is not held up by uncertainty over spending or responsibility.

Monitoring, Documentation and Visibility

LaBrie said no single gap is usually responsible for a slow or ineffective response. Monitoring, failover mechanisms, documentation, provider accountability and network visibility all matter.

“Timely notifications, the ability of hardware to sense the fact that there’s an outage, to automatically go to a backup, to have those tools and to kind of understand your network topology in a way down to the details” are all important, he said.

Documentation becomes especially valuable when an outage is underway. Teams may need immediate access to circuit information, device addresses, provider contacts and network topology. If that information is outdated or difficult to find, restoration takes longer.

“A lot of times people spend time flipping through network diagrams, asking, ‘What was the address of that? I need it right now,’ but you don’t have it.”

Accurate network diagrams and operational records should therefore be treated as ongoing requirements rather than project documentation that is only updated during a deployment.

Communications processes are equally important. LaBrie said previous teams relied mainly on email-based notifications, with priority levels to signal issues that required an immediate response. The platform may differ today, but the principle remains relevant: stakeholders need consistent, timely updates without IT teams overpromising on restoration times.

The right people also need a shared view of the issue. That can include internal IT staff, leadership, managed service providers and connectivity suppliers. Without clear ownership and a common understanding of the current network state, incident response becomes fragmented.

SD-WAN Brings a More Proactive Approach

LaBrie believes SD-WAN and software orchestration have improved the way enterprises manage these risks over the past five years.

Rather than relying on separate tools and manual interventions, organisations can use a centralised platform to monitor network health, generate alerts and automate responses. It gives network teams and providers a common operational view.

“We’re all looking at the same tool. We’re all looking at the same status,” LaBrie said.

A correctly deployed orchestration platform can help reduce the likelihood of a “sick but not dead” condition preventing automatic failover. It cannot eliminate every outage, but it can identify degraded connections more effectively and automate tasks that previously required engineers to react manually.

That allows network teams to move beyond continuously validating connections and handling routine failures. Instead, they can spend more time improving network performance, planning capacity and supporting changing business requirements.

AI and New Connectivity Options

Looking ahead, LaBrie expects AI to contribute to stronger network operations by helping systems detect more issues and support faster decisions.

He warned that vendors can overstate their AI capabilities, but said the technology will help enterprises sense problems more effectively. For network teams, that could mean earlier warning of degradation and more intelligent responses before a service interruption becomes widespread.

The essential lesson is that resilience is not achieved by simply buying a secondary circuit.

Organisations need to test how their network responds when a primary route is degraded, confirm that failover will happen when required and ensure teams know how to act if it does not.

Network redundancy is only valuable when it restores users at the moment they need it most.

rate this story
helps rank stories across uc today
The discussion0 takes · attributed & checked

Does this reflect your experience?

opening the room…
Read nextordered by techtelligence · every pick explained
same beat · Service & Connectivity

Can Better Integration Save Customer Trust? Xurrent Is Betting on It

11 Aug 2026
same beat · Service & ConnectivityThe Future of UC Onboarding Is Self-Service10 Aug 2026same beat · Service & ConnectivityTata Communications Global Network: Is Network Fabric Ready for AI Service Management?7 Aug 2026