The AWS incident: it was not an outage, it was a reality check
Venmo, Fortnite, Snapchat, Zoom, Coinbase... everything started to crash. Not because of ransomware, a large-scale attack or a geopolitical conflict.

Venmo, Fortnite, Snapchat, Zoom, Coinbase... the entire AWS ecosystem started to crash. Not because of ransomware, a large-scale attack or a geopolitical conflict. No. Because a component of AWS US-EAST-1, in Northern Virginia, went off the rails.
People call it a "technical incident". I call it a symptom. And an admission of systemic fragility.
It was not a hack. It was worse.
When there is an attacker, you can identify an adversary, a vector, an intent. Here, there was none of that. Just a small internal glitch, with massive consequences.
The official summary published by Amazon leaves no room for doubt: during the night of 19 to 20 October 2025, a latent race condition in DynamoDB's DNS management system produced an empty DNS record for the service's regional endpoint, dynamodb.us-east-1.amazonaws.com. In plain terms, the address of the service simply vanished.
The mechanism is almost mundane. Two automated systems share the job of updating DNS: a DNS Planner that computes the addressing plan, and a DNS Enactor that applies it in Route 53. That night, one enactor fell behind on an older plan while a second was applying a more recent one, then launched the clean-up. The result: the stale plan overwrote the correct one, and every IP address of the regional endpoint evaporated. Not an intrusion, not sabotage: two internal components that stepped on each other at the wrong moment.
No noise. No alert. Just a brutal silence. And an entire global ecosystem grinding to a halt. Because a single geographic zone concentrates too many dependencies.
That is the real danger. Not the hacker. The silent technical monopoly.

How did a single DNS bug bring down half the web?
Because nothing in this architecture is truly isolated. Once DynamoDB's address had evaporated, the domino effect was immediate. EC2 servers could no longer be launched, load balancers failed their health checks, and dozens of services followed in cascade: Lambda, ECS, EKS, Fargate, Redshift, all the way to the IAM and STS authentication building blocks, which serve the entire planet.
On the consumer side, the list of casualties reads like a directory of everyday digital life: Alexa, Snapchat, Fortnite, Venmo, Coinbase, Zoom, Reddit, Disney+, Roblox, Bank of America, Lyft, DoorDash, the New York Times, and even Amazon.com itself. The DNS fault may have been patched around 2:40 a.m., but restoring the dependent services stretched across the whole day: EC2 instance launches only returned to normal around 1:50 p.m., load balancers around 2:09 p.m., Pacific time. Close to fifteen hours of cascading silence, from the middle of the night to the early afternoon.
One technical detail, in one region, at one provider that accounts for around 30% of the global cloud market. That is the real surface of our dependency.
An isolated outage, or a deeper signal?
No, this is not an accident specific to Amazon. Less than a month later, on 18 November 2025, it was Cloudflare's turn to collapse: a simple database configuration change generated a corrupted file, propagated it across the whole network, and part of the web became unreachable, from ChatGPT to X, Spotify, Uber or Canva. A different provider, a different cause, exactly the same scenario: one central link gives way, and thousands of services fall with it.
Two giants, two internal glitches, two planet-wide domino effects in four weeks. That is not bad luck. It is the signature of a global digital infrastructure resting on a handful of players, whose slightest cough spreads to the entire planet.
The cloud is not magic. It is mechanical.
Many organisations adopted the cloud the way you take a sleeping pill: to sleep easy. Less infrastructure, less complexity, less cost.
But what we gained in comfort, we lost in control.
Today, a large majority of companies do not know what their services really depend on. They use "off-the-shelf" digital tools, interconnected, stacked, often of external origin. But deep down, everything rests on a few very localised zones, always the same ones.
And those zones, too, can go down.
A governance crisis before a technical one
What this incident reveals is our lack of rigour on critical matters:
- We outsource without a fallback strategy.
- We pile up tools without any real contingency plan.
- We trust blindly without asking the right questions.
In a crisis, it is not the tool that makes the difference. It is the ability to understand fast, to pivot, to decide.
And for that, you need readable maps. Not obsolete diagrams in a folder no one can find.
The three blind spots no one likes to look at.
- Geographic concentration: too much data, too many services, too many customers rest on a single critical point. That is not a "hosting strategy", it is a bottleneck. US-EAST-1 is one of the oldest AWS regions, the one where so-called global services such as IAM and STS authentication are anchored, precisely the ones that gave way that day.
- The illusion of redundancy: believing you have an alternative when, in fact, everything runs through the same invisible structures (connections, authentications, flows).
- The autonomy deficit: we do not train teams to operate without their favourite tools. The day the tool fails, everyone just waits for it to come back.
Should you leave AWS after an outage like this?
No. Reproducing the same pattern at another provider solves nothing.
Amazon is not the problem. Our collective passivity is.
Amazon does what Amazon must do: optimise, industrialise, deliver. And the company fixed the fault, documented the timeline and published its analysis in the open. It is even one of the few players to do so that transparently.
The problem is us. It is up to us to keep an architecture designed for failure. It is up to us to spread the load across several regions, or even several providers, rather than stacking everything on a single point. It is up to us to reduce the dependent surface and to know what happens if our favourite provider has a bad Monday.
How do you prepare for the next cloud outage?
The good news is that an outage like this one is in no way unpredictable. It is a design assumption, not a surprise.
In practice: map your real dependencies, test failover regularly, train teams to operate in degraded mode, and check that your "redundancy" is not an illusion routing everything through the same pipes.
The question is not "when will AWS go down again", but "will you be able to cope without it for a few hours?"
Learn more about Christophe Mazzola
Sources
- Technical cause (DNS race condition, DNS Planner and DNS Enactor), detailed timeline and affected AWS services: official incident summary, Amazon Web Services, 20 October 2025.
- Consumer companies affected and AWS share of the cloud market (around 30%): Engadget, 20 October 2025.
- Timeline and scope of the outage: "2025 Amazon Web Services outage", Wikipedia.
- Cloudflare outage of 18 November 2025 (configuration change, domino effect): "2025 Cloudflare outage", Wikipedia.
Frequently asked questions
What caused the AWS outage of October 2025?
According to Amazon's official summary, a latent race condition in DynamoDB's DNS management system produced an empty DNS record for the service's regional endpoint in the US-EAST-1 zone. Two DNS automated systems, a planner and an enactor, fell out of sync: a stale plan overwrote the correct one and wiped out every IP address. No ransomware, no attack: an internal glitch with global consequences.
How long did the AWS outage last?
Close to fifteen hours. The first errors on the DynamoDB API appeared at 11:48 p.m. Pacific time on 19 October 2025. The DNS fault was fixed around 2:40 a.m., but restoring the dependent services (EC2 instance launches, load balancers) stretched into the early afternoon. Amazon announced a return to normal at 6:53 p.m. New York time on 20 October.
Which services and companies were affected?
On the AWS side, EC2, Lambda, ECS, EKS, Redshift and IAM and STS authentication. On the consumer side, Alexa, Snapchat, Fortnite, Venmo, Coinbase, Zoom, Reddit, Disney+, Roblox, Bank of America, Lyft, DoorDash, the New York Times, and Amazon.com itself.
Why is this incident more worrying than a hack?
Facing an attack, you can identify an adversary, a vector and an intent. Here there was nothing to analyse: just a brutal silence revealing a systemic fragility and a silent technical monopoly.
Is this a problem specific to AWS?
No. Less than a month later, on 18 November 2025, a simple configuration change brought down a large part of the Cloudflare network (ChatGPT, X, Spotify, Uber, Canva). A different provider, a different cause, the same planet-wide domino effect: the concentration of global infrastructure on a handful of players is the real subject.
Should you leave AWS after this outage?
No. Reproducing the same pattern at another provider solves nothing. The point is to spread the load across several regions, to keep an architecture designed for failure and to reduce the dependent surface.
What question should we ask instead of "when will AWS go down?"
The right question is: "will you be able to cope without it for a few hours?" It is about mapping your dependencies, testing failover and keeping an architecture designed for failure.

Être en cybersécurité
A cyber roadmap in plain language, for everyone, not just the experts.
