The UK Must Move Away from Single Points of Failure in its Digital Infrastructure
Mon 24 Aug 2026 | Mike Hellers

Despite the Internet being designed as a distributed network, many organisations have gradually concentrated their operations around a very small number of networks, cloud platforms and data centres. The result of this is an uncomfortable contradiction. While the services themselves may appear distributed, the underlying infrastructure supporting them still depends upon a single critical component.
When a configuration-related failure propagated across Cloudflare’s network and disrupted access to services globally in November 2025, the takeaway was not that disruption can be eliminated entirely, rather the necessity of ensuring that one failure cannot cascade across an entire organisation or across every service on which millions of people depend.
Why Resilient Often Isn’t
Establishing a clear distinction both in principle and practice between cybersecurity and resilience is crucial, as organisations often use these terms interchangeably. But they are not the same thing.
Threat detection tools, firewalls, and incident response policies are essential cybersecurity tools for protecting systems from malicious activity. However, they are incapable of maintaining connectivity if the underlying route carrying an organisation’s traffic becomes unavailable.
Resilience, on the other hand, is not only about preventing an incident from occurring at all but is about keeping these essential services functioning when an incident inevitably occurs. Organisations may have detailed continuity plans, multiple cloud applications, and extensive security controls, yet all of these are redundant if the entire ecosystem is reliant upon one network fabric to exchange its Internet traffic. If that fabric fails, the organisation will lose access to its customers, suppliers, and cloud services all simultaneously, regardless of how sophisticated the other protections are.
It’s important to understand that redundancy can also be misleading. Two separate connections may appear independent while passing through the same carrier, router, fibre duct, data centre, and software environment, and in this case the organisation has duplicated its two connections without removing the underlying point of failure. This means two routes sharing the same dependency are effectively one route in disguise.
It’s crucial, therefore, for organisations to map the full path their traffic navigates and identify shared physical, technical, and operational dependencies, not simply count the number of services that they have purchased to mitigate risks.
The Critical National Infrastructure Risk
Failure to do so can result in consequences that extend far beyond an individual’s company website becoming temporarily unavailable or offline. Healthcare providers, government services, banks, payment systems, streaming platforms, and communication services all crucially depend on interconnected layers of digital infrastructure.
While these sectors may not all use precisely the same systems, they nonetheless rely on overlapping networks, data centres, cloud providers, and technology suppliers. A failure within one heavily relied upon provider will catastrophically ripple through multiple organisations and sectors.
For members of the public, this manifests in an inability to access medical services, make payments, contact organisations or obtain important information, and what begins as a technical fault quickly spirals into a socio-economic spiral.
The UK’s Cyber Security and Resilience (Network and Information Systems) Bill reflects growing recognition of this systemic risk. This legislation would bring data centres into scope as essential services and extend regulation to medium and large managed service providers and strengthen incident reporting duties.
Regulation needs to provide a stronger baseline, but compliance alone does not make infrastructure resilient, and organisations must be in the position to demonstrate how services will remain available even when a critical component fails.
What Genuine Resilience Looks Like
For organisations exchanging significant volumes of Internet traffic, the starting point needs to be a genuinely independent alternative route. Dual fabric infrastructures allow organisations to connect to two separate network fabrics so that traffic can move to the second if the primary fabric becomes compromised or unavailable.
Independence is the crucial element. Fabrics should not rely on the same hardware, software, or network operating environment, as a single technical fault can affect both simultaneously. The London Internet exchange operates two independent peering LANs LON1 and LON2, using diverse infrastructure, which allows connected networks to exchange traffic across both rather than placing their entire reliance on one fabric.
The secondary route must have sufficient capacity to handle traffic during an incident. A backup that becomes overwhelmed as soon as it’s needed provides little to no meaningful protection, and these failover arrangements should be configured in advance and tested very regularly. Organisations should not discover during an outage that traffic cannot reroute automatically or that configurations have drifted, or that the internal teams do not know how to activate recovery procedures.
Controlled failover exercises, continuous route monitoring, and clear recovery objectives turn redundancy from a procurement claim into an operational capability. Whilst this requires significant investment, this cost must be compared with the financial, operational, and reputational consequences of prolonged downtime, not necessarily with the cost of operating a single connection under ideal conditions.
Foot the Bill
The Cyber Security and Resilience (Network and Information Systems) Bill should prompt every organisation to ask one simple question: what happens if our primary route disappears tomorrow?
If the answer to this question is that services grind to a halt, customers lose access, and recovery depends entirely on hurried manual intervention, the organisation is far from resilient.
While the UK cannot prevent every outage, it can design infrastructure that contains failures, rather than amplifying them, and contains procedures and infrastructure to ensure that we’re best prepared to navigate the challenges when they inevitably arise. Removing avoidable single points of failure must become a basic expectation of responsible digital infrastructure, not just an optional upgrade, and we will never achieve resilience until this is the case.

