In today’s digital economy, the data center is the beating heart of business operations. From e-commerce and finance through to heavy duty AI workloads, almost every organization now relies on uninterrupted network delivery to maintain customer trust, operational activities and profitability. Downtime, even when measured in seconds, can result in lost revenue, SLA violations and reputational damage.
Indeed, in a recent study by independent technology research company Futurum, ‘The data center networking imperative: Key trends driving the next era of data centers’, a staggering 86 percent of IT infrastructure leaders selected reliability as the single most important factor when building their next data center network, outranking ease of integration (82 percent), ease of operations (74 percent) and automated operations (70 percent).
To understand how enterprises can improve their network reliability, DCD spoke to Anthony Peres, director, product marketing, data center networks at Nokia, who has spent decades studying the evolution of mission-critical IP networking.
Peres confirms that while predictable performance has always been a priority, the spotlight is now on next-generation applications and use cases. “While reliability applies to traditional workloads too, with AI the need for reliability and highly available infrastructures – and by consequence highly available networks – is now paramount.”
Legacy architectures
For operators seeking guaranteed availability, the first step is recognizing the systemic vulnerabilities embedded in their current – or legacy – architectures, referred to in reliability modeling as the Present Mode of Operation (PMO). Typically, this comprises outdated hardware and software, manual processes and limited automation.
According to Nokia’s Peres, merely running the network on legacy software compromises resilience from the outset. “Many of these deployments today actually have operating system software that was built decades ago,” he says, adding that “the legacy PMO is characterized by a set of capabilities and attributes that are lagging.”
These older operating systems lack the architecture and software design quality that’s essential in modern data center buildouts. The lack of flexibility and reliance on cumbersome Command Line Interfaces (CLIs) for management mean that any change or upgrade is inherently high-risk, leading to manual processes that fail to scale and are prone to errors.
Conversely, a modern data center fabric comprises a series of leaf and spine switches for connecting traditional and AI-based applications installed on servers. The switch hardware provides the physical networking interfaces, switching matrix, control complex, fans and other physical components while network operating system (NOS) software provides the intelligence, routing capabilities, telemetry and other programmable capabilities required to operate the switch.
“All of these need to be managed in an easy and succinct way which is where the third element, the network management platform – sometimes referred to as the operations or automation platform – comes in,” says Peres. “A full data center fabric comprises all three – hardware, software and a management platform.”
For example, the Nokia Data Center Fabric solution comprises high-performance hardware platforms (such as the Nokia 7220 IXR and 7250 IXR), which run Nokia SR Linux NOS and are managed by Nokia Event-Driven Automation (EDA) – an automation platform that combines speed with reliability, while providing guardrails for detecting errors caused by automation.
Measuring reliability
Inevitably, for network operators, one of the main challenges is how to measure the potential benefits of moving from a legacy PMO system to a Future Mode of Operation (FMO) with the latest data center fabric. In other words, how do they know they will see a return on their investment?
While most now acknowledge the importance of reliability ”…there hasn’t been a way to quantify what’s good versus what’s better,” says Peres. “We’ve had a lot of press coverage recently about outages, but how do you quantify this? What’s the impact of that downtime?” he adds.
Enter Nokia’s Data Center Fabric Reliability Study. Put together by Bell Labs Consulting in collaboration with Futurum, it’s been designed to move beyond anecdotal evidence to provide clear, quantifiable metrics that operators can understand.
Using a rigorous reliability model developed by Bell Labs Consulting, it compares the PMO with a FMO based on the Nokia Data Center Fabric solution. “Bell Labs Consulting does plenty of these studies and they come with a lot of credibility because of the way the models are created, the parameters that are used and the work done with actual customers,” explains Peres.
Complementing Bell Labs technical abilities, Nokia turned to Futurum to use its industry experience to “…help quantify what a modern data center should look like,” explains Peres. “We wanted to make sure that it wasn’t just us saying it. We wanted an external party to come in, challenge us and say, ‘Okay, you're not thinking of this, or could you add this into the model?’”
Human error zero
Inevitably the greatest cause of network downtime is operational error stemming from human intervention. Manually managing thousands of network elements is fraught with risk, where simple configuration inconsistencies can cascade into devastating outages. According to The Data Center Networking Imperative report, more than 80 percent of organizations admit that human error at least occasionally impacts service continuity with 17.5 percent claiming it is a ‘frequent top cause for disruption.’
With its ‘Human Error Zero’ mission, Nokia directly addresses this systemic weakness. This philosophy seeks to eliminate errors originating from both vendor software shortcomings and, crucially, user mistakes. “Human Error Zero isn’t just a rallying cry or a banner, it’s what we are about,” says Peres. “Vendor errors and operations errors – that’s what human error zero is trying to reduce,” he adds. “We’re not there yet but we’re getting close.”
Critically, the Bell Labs methodology recognized that reliability is overwhelmingly an operational challenge, not purely a hardware limitation. Reliability gains can be achieved through enhanced software and automation, which complement the resilient hardware platforms.
However, while automation is essential for scaling the data center, many operators still hesitate. “AI and cloud providers automate more, but for others it’s more like 25 to 30 percent of tasks which are automated,” says Peres. The core barrier is that if the underlying system is fragile, automation merely accelerates the process of breaking the network, resulting in amplified, widespread failures.
“While speed is an important element, you can’t have it without reliability. If I'm performing an operations task, I need to know it’s going to work as intended and not break something,” explains Peres.
The leap to 5.1 nines
Importantly, to make the findings directly relevant to network teams, the Data Center Fabric Reliability Study modeled specific operations use cases across the network lifecycle, including the design, deployment and operations phases. “A foundational model took each element and split it across hardware, software, configuration, before overlaying operations use cases that relate to the different phases,” explains Peres.
What the results of the study found was that the PMO baseline – reflective of a traditional, manually operated network – achieved an availability of only 99.981736 percent (3.7 nines), resulting in a cumulative annual downtime of approximately 96.1 minutes. However, the FMO solution leveraging SR Linux and EDA achieved 99.999235 percent (5.1 nines) availability, reducing total annual downtime to just 4.0 minutes.
In other words, the Nokia Data Center Fabric solution delivers 23.9 times less downtime, equating to a 96 percent reduction in annual downtime compared to the legacy solution.
Furthermore, the FMO delivers its greatest gains in the operational phases which are the most prone to human error. For example, downtime is reduced by up to 95 percent for configuration and provisioning tasks and up to 99 percent for operations and monitoring tasks as well as scheduled maintenance tasks. According to Peres, these gains are enabled by a holistic architecture that addresses failure at three levels: the NOS, the automation platform, and the operational process itself.
He emphasizes that combining the components to provide a complete ‘holistic architecture’ is key. “You can just take SR Linux by itself, but if you can combine it with EDA that’s where you will see the most benefit,” he says. “It’s not just about good hardware and software, but the automation toolkit.”
Modular approach
Unlike legacy monolithic NOS architectures, SR Linux was built from the ground up on an unmodified Linux kernel, giving it intrinsic stability and resilience. Whereas in a monolithic NOS, failure in one process can bring down the entire switch, SR Linux mitigates this by isolating every component – from routing protocols to service features – as an independent module.
Peres explains: “Modularity is good because it can reduce the impact of failure. If you have one monolithic element and something happens, everything goes down. But if you have modularity and a particular module or element goes down, everything else is still running.”
Complementing SR Linux NOS in the Nokia Data Center Fabric solution is an intelligence layer provided by EDA. A cloud-native management solution built on Kubernetes, EDA is, as Peres puts it, ‘the embodiment of Human Error Zero,’ constantly monitoring the network’s actual condition against its intended state.
Undoubtedly, the most transformative feature within EDA is the digital twin. Addressing the core operational barrier – the fear of making a change – the digital twin is a virtualized, containerized replica of the production network. This virtual environment allows network operators to perform critical risk mitigation before any configuration touches live hardware by validating new configurations, OS upgrades and policy changes within the replica environment.
As Peres explains, a digital twin eliminates the trial-and-error approach that’s no longer commercially viable: “It allows you to simulate or validate changes or configurations in a virtual environment before you apply it to the real-world network. If I make a change in a live environment, I don’t know what’s going to happen, but with the digital twin I can make changes and validate them without worrying.”
Reducing costs, increasing morale
Importantly, improved data center reliability translates into major cost reductions. According to Bell Labs, Nokia Data Center Fabric delivers savings across three critical categories: reducing exposure to SLA penalties and operational costs by up to 60 percent; reducing potential revenue loss during downtime and churn by up to 53 percent; and minimizing reputational damage and brand loss by up to 44 percent.
Beyond the balance sheet, the FMO also addresses the severe problem of operational fatigue and team morale. Whereas in legacy environments, engineers are constantly trapped in a reactive posture, suffering from ‘alert fatigue’. By improving network reliability, the FMO helps teams achieve their Mean Time to Innocence (MTTI) goals – in other words, the time it takes for network teams to prove that an outage was not caused by the network.
Peres directly links reliability to team efficiency and well-being. “This really gets to the crux of what we can do to make life better and easier for that network engineer,” he states. “The end goal is to empower operations teams and allow them to spend more time doing the things they want to be doing outside of work.”
Time to switch?
For network operators, the question is whether they can afford to remain trapped in the operational fragility of the PMO, especially as AI demands absolute continuity. The quantifiable proof provided by the Bell Labs study (the shift from 3.7 nines to 5.1 nines) is the guarantee needed to sustain costly AI training jobs and sensitive financial operations.
Furthermore, the FMO delivers a 96 percent reduction in downtime, enabling operators to finally embrace the concept of ‘Human Error Zero’ through architectural resilience and the predictive power of the digital twin.
Finally, the real measure of this enhanced reliability is the reduction in the daily operational burden. Speaking about a customer who deployed Nokia Data Center Fabric alongside a competitor’s solution, Peres highlighted the stark operational difference: “As one customer shared, ‘I’ve only raised one ticket with you in the same time as I’ve raised 75 with another equipment vendor.’”
For a deeper dive into how next-generation data center fabrics can dramatically boost reliability and slash downtime, explore Data Center Fabric Reliability Study by Bell Labs and Futurum. Read the executive summary here.


Comments