Systems

  • Reliability
  • real estate
  • recovered
  • Principal — stayed through recovery
  • weeks

The same timeout cascade should not take the product down twice.

What was on fire

The product was down. AWS was green. A shared MongoDB cluster was saturated; timeouts cascaded through ECS/Fargate services.

Constraint set

No maintenance outage. The store was owned elsewhere; a redesign of that plane was not available in the window. The product still had to survive it.

System diagram

ECS / Fargateproduct servicescache layersappsharededgestampede stops hereMongoDBowned elsewhere
Fig. 01 — The app stopped stampeding Mongo when Mongo was already hurting.

Architecture decisions

  1. Cache in front of a store we didn’t own, rather than wait for that team’s redesign. The pager does not care who holds the ticket.
  2. Three layers on purpose — in-process, shared cache, HTTP/edge. A single cache is another single domain.
  3. Under traffic, not a flag day. There was no outage window to “do it right.”

Not again: don’t leave a product on an uncached shared store you don’t own.

Recovery / operate path

Weeks in it. That page class did not recur for months. No game-day claim. No MTTR published.

What changed after

The product took responsibility for surviving the shared store. Ownership of Mongo stayed elsewhere.

Anonymization notes

Unnamed. No cluster names, no org chart, no ticket IDs. Stack is a footnote: ECS/Fargate, MongoDB, in-process / shared / HTTP cache.

Who this is for

  • A product sitting on a shared store another team owns.
  • AWS green, users down, the same saturation cascade coming back.
  • No maintenance window, and the “real” database fix is not this quarter.

Start

Email me what’s breaking

If production is down, put LIVE in the subject. I’ll tell you the same day if I can come in.