Skip to content

The Race Condition That Emptied the Cloud's Address Book

In October 2025 two AWS automation components raced, and the DNS record for a foundational database endpoint came out empty. Much of the consumer internet limped, and the automation was locked out of its own repair path.

Episode 194 minute read

A blank where an address should be

A blank where an address should be

Overnight on 19 October 2025, AWS's oldest region stopped being able to tell anything where one of its databases lived. In AWS's own account the disruption ran from 11:48 PM Pacific on 19 October to 2:20 PM Pacific on 20 October, in three distinct periods rather than one continuous outage, and the root cause was a latent race condition in the DynamoDB DNS management system that left an incorrect empty DNS record for the service's regional endpoint, which the automation then failed to repair.

DynamoDB matters here beyond its own customers, because a large share of AWS's internal machinery is built on it. The failure was not in the database. It was in resolving the address of the database, which for anything trying to reach it amounts to the same thing.

One planner, three enactors

The design is worth stating exactly, because the lesson lives in it. A DNS Planner watches the health and capacity of the load balancers behind an endpoint and periodically produces a new plan: a set of load balancers and weights. A separate component, the DNS Enactor, applies that plan through Route 53. The Enactor is deliberately built with minimal dependencies so it can work during a recovery, and for resilience it runs redundantly and independently in three availability zones.

That independence is what raced. AWS describes an unlikely interaction between two Enactors, after which an active plan was deleted. All the IP addresses for the regional endpoint were removed at once, and the system was left in an inconsistent state that prevented any Enactor from applying later plans. The automation had produced a state it could not leave, and restoring it required manual operator intervention.

The cascade

Because so much sits on DynamoDB in that region, the blank record spread outward, including into AWS's own systems. Instance launches were impaired as a process that depends on DynamoDB failed its state checks, and EC2 began returning insufficient capacity errors for new launch requests. As recovery proceeded, Network Load Balancer health checks began alternating between failing and healthy, so checks failed against nodes and targets that were themselves healthy.

Outside, it looked like an ordinary bad day on the internet. Lloyds Banking Group confirmed that some of its services were affected. Reports gathered against consumer apps, smart home devices, games and government websites, though those rest on user reports rather than statements from the companies named, and AWS's own summary names no customer at all.

What AWS has and has not done

AWS published the mechanism in unusual detail, and the account has not changed in substance since late October 2025: same window, same cause, same architecture, same remediation list, no retraction and no revised cause. It is worth being exact about the tense of that list, because it is easy to read as work already finished. AWS states that it has disabled the DynamoDB DNS Planner and DNS Enactor automation worldwide. The fix for the race condition and the additional protections against applying incorrect DNS plans are what it says it will do before re-enabling that automation. The velocity control for Network Load Balancer failover is described as being added.

What it teaches

Automation failures are increasingly orderings rather than events. Nothing here was a broken component. A planner produced valid plans and enactors applied them, and the defect lived in an interleaving that no test had exercised. Testing concurrent automation means testing the orderings, not only the operations.

Ask of any automated system whether it can recover every state it can produce. The empty record was inside this system's possibility space, because the system produced it, and outside its recovery space. That asymmetry is findable before the fact: if the automation can write a state, either prove it can unwind it, or prevent it from being written.

And concentration turns a regional defect into a global one. Northern Virginia was the first region AWS opened and remains its main one, which means an extraordinary number of critical paths quietly assume it. The useful question after a day like this is not whose fault it was, but which of your own paths would have noticed.

Sources

4 sources

Every figure in this article traces to one of the following: the same record the episode cites.

01Zof Console

One surface for posture, operations, and what needs attention next.

The authenticated home that engineering, QA, and SRE teams open every day: quality posture, in-flight runs, coverage by module, and what needs attention next.

OPERATIONAL KPIs

  • Runs
  • Coverage
  • Risk

Live across every environment you ship to.

WORK SPINE

  • Specs
  • Tests
  • Schedules

From specification to scheduled regression.

GUARDRAILS

  • RBAC
  • SSO
  • audit

Every action attributable to a named human.

LIVE/console
Zof AI home command center showing 12 runs at 94% pass, 3 open critical issues, 84% coverage, four module traceability bars, the specification pipeline, upcoming schedules, and recommended next actions with an active-runs sidebar.
Console home · Checkout Service · Staging · captured live from the product.
The Race Condition That Emptied the Cloud's Address Book | Zof AI