AWS's Oregon Outage: How 80 Minutes Broke Half the Internet
At 10:55 UTC on July 24, a piece of networking hardware at the edge of Amazon's Oregon data center failed quietly. It handled the routing link between AWS us-west-2 and the Seattle metro — the connection between the region's internal fabric and the rest of the internet. Compute inside the region kept running. Databases kept running. Only one thing stopped working: traffic crossing that boundary.
That was enough to knock Apple Pay, Reddit, Hulu, DoorDash, and PlayStation Network offline for roughly 80 minutes and to send cascades through at least nine other providers' services for the rest of the day. This was the third distinct AWS reliability incident in approximately eleven weeks, and it followed Azure going down the day before.
What Actually Failed — and Why Internal Redundancy Didn't Help
AWS engineers automatically engaged at 11:01 UTC. Initial mitigation completed at 11:15 UTC — a 20-minute core impact window for most customers. For businesses using AWS Direct Connect through the EqSe2 facility at Seattle's Westin Building Exchange, the outage lasted 1 hour and 17 minutes.
The technical nuance here matters: AWS's multi-AZ architecture was not the problem, and it wasn't the solution either. Multiple availability zones protect against failures within a region — hardware failures, power failures, cooling events. This failure was at the boundary between the region and the public internet. No amount of zone-level redundancy helps when the network exit itself breaks.
Route reconvergence — the process where routers re-establish agreement on network paths after a failure — caused a second round of intermittent connectivity from 11:47 to 11:59 UTC, even after the primary mitigation was declared complete. AWS posted its first public Health Dashboard update at 11:40 UTC and its closing summary at 13:01 UTC. That closing update felt premature to a lot of customers still seeing errors.
The Cascade Tail Nobody Talks About
"AWS is back" and "your service is back" are two very different statements. The cascade tail from this outage stretched for most of the working day:
- NinjaOne (remote monitoring and management) remained affected for 9 hours and 29 minutes after AWS restored connectivity. Its agents used backoff-and-jitter reconnection logic to avoid overwhelming recovering infrastructure while 150,000+ devices came back online.
- SendGrid (email delivery) needed 7 hours and 32 minutes to clear its backlog.
- SparkPost took 6 hours and 45 minutes to drain queued mail.
IncidentHub tracked nine confirmed cascade incidents across seven providers, with three additional possible cascades. The root mechanic is straightforward: a 20-minute disruption to mail delivery queues far more messages than 20 minutes of processing capacity can clear. Queue management becomes a capacity problem at scale, not a connectivity problem.
If your business depends on SaaS tools running on us-west-2 — email delivery, agent-based monitoring, payment processors — plan your recovery expectations around realistic tail times, not the timestamp on the AWS closure notice.
July 2026 Has Been Hard on Cloud Infrastructure
This wasn't an isolated bad day. July stacked outages across all three major hyperscalers in a way that deserves attention:
- July 15–16: Google Cloud's europe-west4-a zone (Netherlands) lost a cooling system and went dark for 12.5 hours, affecting VMware Engine and NetApp Volumes.
- July 16: AWS CloudFront's VPC Origins fleet hit an internal connection-management limit, disrupting HTTP traffic for roughly 3.5 hours.
- July 17: AWS US-EAST-1 saw disruptions to Lambda, CloudFormation, and Amazon Connect.
- July 23: Microsoft Azure's West US region lost wide-area network connectivity for nearly five hours after a maintenance operation pulled routes off a wider set of devices than intended.
- July 24: AWS us-west-2, the incident above.
Two of the three AWS incidents this month broke at region-to-network boundaries — not inside the data center, but at the edge where the region connects to the wider internet. Internal redundancy is irrelevant there.
What AWS Compensates — and What It Doesn't
Under the standard EC2 SLA, AWS credits affected customers roughly 10 percent of monthly compute spend on impacted instances, once you file a claim. There is no compensation for lost revenue, damaged customer relationships, or regulatory exposure from unavailability. A 20-minute payment-processing outage at 10:55 UTC on a Friday morning costs real money that a compute credit doesn't touch.
Three Things Worth Reviewing Now
Geographic redundancy at the network level, not just the zone level. AWS's own postmortem noted that customers with Direct Connect paths through a location other than EqSe2 Seattle were unaffected by the extended window. Multi-region architecture means more than distributing VMs; it means distributing your network exit points. If all your Direct Connect circuits land at the same physical exchange, you have one of those points.
Map your SaaS dependencies to cloud regions. If your CRM, email delivery, monitoring stack, and payment processor all run on us-west-2, an Oregon boundary failure takes them all down simultaneously — regardless of where your own servers live. Most businesses have never built this map. It takes an afternoon and is immediately useful.
Build recovery SLAs around the tail, not the headline duration. For outage communications to your customers, the realistic timeline is the cascade tail — hours, not minutes. Knowing that NinjaOne historically needs time to drain its reconnect queue after an AWS event is operational intelligence that changes how you staff and communicate during incidents.
This is why at Falcon Internet, 24x7x365 NOC monitoring tracks upstream infrastructure health, not just our own — so we know about a boundary failure in Oregon before our customers have to wonder why their email isn't delivering.