While AWS and Microsoft Azure dominate the enterprise cloud market, both have faced reliability incidents in recent years. Gartner Vice President Analyst Lydia Leong argues that these events are not evidence that the cloud is unreliable — only that it is not infallible.
“These events highlight an important truth,” Leong wrote on Gartner’s site. “Cloud disruptions happen, but they are not evidence that the cloud is inherently unreliable.”
She cautioned against knee-jerk reactions like “repatriation,” moving workloads back on-premises, or “geopatriation,” shifting to smaller regional or sovereign clouds. “These moves often introduce new risks and may even slow down recovery when things go wrong,” Leong warned.
She also dispelled the notion that adopting a multicloud strategy automatically guarantees uptime. “Pursuing multicloud resilience can cost more than it saves, introducing technical complexity without truly eliminating systemic risk,” she said.
Leong’s central message is one of perspective: “Cloud outages make headlines because they affect so many people at once, but context matters. Every major provider — from Microsoft to Google — has experienced similar events. The real differentiator is how well your organization plans for and recovers from inevitable disruption.”
You Can’t Engineer Away All Risk
That inevitability is a recurring theme among industry experts. Shawn Michels, vice president of product management at Akamai Technologies, said the digital world’s increasing complexity ensures that failures will continue, regardless of scale or sophistication.
“From cloud platform outages to undersea cable cuts, even the most advanced systems can experience failures,” Michels told TechNewsWorld. “A lot of organizations still assume that because something runs in the cloud, it’s automatically resilient — but that’s not the case. Even the biggest clouds don’t have perfect uptime.”
According to Michels, resilience comes not from eliminating failure, but from reacting to it quickly. “What separates the best from the rest is how well a system reacts to small failures to prevent a larger outage,” he said. “You can’t stop every component from breaking, but you can design systems to recover so quickly that customers barely notice.”
He added that true resilience is as much about people and process as it is about technology. “It’s how teams prepare for failure, respond under stress, and learn from every incident,” he said.
Hyperscale Doesn’t Mean Equal Reliability
Not all cloud giants handle reliability the same way, said Rich Mogull, chief analyst at the Cloud Security Alliance. While AWS, Microsoft Azure, and Google Cloud all operate at hyperscale, their infrastructure and regional architectures vary significantly.
“Enterprises tend to gloss over these differences,” Mogull explained. “For example, AWS rarely has cross-region failures, and when they do, they tend to be limited. You can largely plan around this potential. Azure, by comparison, is more likely to experience global failures due to how their infrastructure is designed.”
These structural differences can have major implications for customers running mission-critical applications. A global authentication or DNS issue, for example, can cascade across regions, taking down services even if local compute nodes remain operational.
No Cloud Is Immune to Downtime
The assumption that redundancy equals invulnerability is one of the most persistent misconceptions about cloud reliability, said Ensar Seker, CISO of SOCRadar. “Redundancy mitigates risk, but it doesn’t eliminate it,” he told TechNewsWorld.
“Even hyperscalers like AWS or Azure operate in a complex web of dependencies across regions, zones, and third-party services. An issue in one layer — like identity federation, DNS propagation, or load balancer routing — can still ripple out and break critical functionality, even if the core compute nodes are up.”
Seker pointed to the 2023 AWS outage that disrupted banking portals, hospital systems, and consumer platforms alike. “That incident didn’t happen because AWS lacked redundancy,” he said. “It happened because many enterprises hadn’t built their apps to withstand regional or service-specific degradation.”
His takeaway is blunt: “Cloud outages are inevitable, not hypothetical. The question isn’t if, but how often — and how prepared your organization is.”
Complex Systems Breed Complex Failures
John Strand of Strand Consulting, a telecom-focused research firm in Denmark, echoed that sentiment. “The day that there are clouds with 100% uptime is the day when all problems in this world are eliminated,” he said.
As cloud providers race to expand capacity, complexity itself becomes a source of fragility. “Everyone — and especially hyperscalers — is building tons of new data centers across the world,” Strand observed. “The size and complexity of these centers is exploding, and when that happens, the risk of something going wrong increases. Some problems will be solved over time, but new ones will always arise.”
That complexity isn’t limited to physical infrastructure. Cloud-native architectures often involve intricate combinations of microservices, APIs, and third-party integrations. A disruption in one layer can quickly cascade across the entire stack — a phenomenon that traditional risk assessments often fail to capture.
The Shared Responsibility Reality
Some experts argue that enterprises don’t necessarily overestimate reliability, they just misunderstand who is responsible for it.
“The cloud isn’t a silver bullet,” said Sergiy Balynsky, vice president of engineering at Spin.AI, a cybersecurity company specializing in SaaS protection. “It’s a shared responsibility model.”
Cloud providers offer the building blocks for resilience — regional redundancy, failover mechanisms, and replication — but it’s up to the customer to use them effectively. “Relying on a single region or skipping redundancy isn’t a provider failure,” Balynsky said. “It’s an architectural oversight.”
He emphasized the role of Business Continuity Planning (BCP) and Site Reliability Engineering (SRE) practices in bridging this gap. “BCP and SRE teams plan for failure, spread risk, and keep critical systems running during outages,” he said. “The AWS outage illustrates perfectly why those teams are essential.”
Designing for Resilience
David Stone, director in the Office of the CISO at Google Cloud, noted that enterprises have more control over reliability than they may think. “Customers can absolutely design in resiliency,” he said. “By using different data centers in other regions, deploying workloads into different zones, and building frameworks that span multicloud environments, enterprises can protect themselves from most localized outages.”
Srini Srinivasan, founder and CTO of Aerospike, added that today’s hyperscale platforms already offer the tools needed to achieve exceptional uptime. “There’s no reason that, using existing cloud provider features and capabilities, an enterprise cannot achieve four nines of availability,” he said. “The fallacy people have is that the cloud provider will solve everything for them.”
Scale Is Not Invulnerability
Aykut Duman, a partner at global consultancy Kearney, warned that even the most redundant systems can fail in unexpected ways. During the AWS outage, some organizations running workloads across multiple availability zones still went dark due to a DNS resolution failure.
“This incident revealed that reliability depends as much on workload architecture and distribution as it does on provider infrastructure,” Duman said. “Enterprises often assume redundancy at the provider level guarantees uptime, but resilience must be deliberately engineered at the application level.”
He summarized the industry’s core problem succinctly: “Enterprises overestimate cloud reliability because they equate cloud scale with invulnerability. Reliability is high, but not absolute.”
Rethinking Cloud Strategy
As the cloud matures into critical national and corporate infrastructure, the definition of reliability must evolve. Outages will continue, and hyperscale providers will continue to invest billions to minimize them, but enterprises must take greater responsibility for their own uptime.
The future of cloud strategy will likely combine elements of multicloud, edge computing, and application-level resilience, using automation and observability to detect and recover from failures quickly.
True reliability, experts agree, is not about preventing every outage. It’s about ensuring that when failure happens — and it will — it becomes a minor incident rather than a major crisis.
Conclusion
The modern cloud is both powerful and fragile — a distributed marvel that delivers unprecedented scalability, but one that depends on careful architecture, continuous testing, and organizational readiness to stay resilient.
Enterprises that continue to treat cloud reliability as a given are setting themselves up for failure. Those that treat it as a discipline — one that combines technology, process, and culture — will be the ones that stay online when the next major outage inevitably hits.


