Alerts fired. The dashboard looked fine.
Error rate ticked up on checkout. CPU and memory were flat. The team jumped to rollback the last deploy.
Rollback did nothing. The problem was not the release we feared.
Checkout failing in one region
Payments were failing for one region. Pressure was high. Chat filled with theories.
Rollback without proof
We wanted speed without proof. The first assumption felt plausible because the timing matched a deploy.
Evidence was thin. We had not compared failing and healthy requests.
Observe, then change
Observe, isolate, verify, then change. Reproduce the symptom. Narrow scope. Use logs and traces to confirm.
A postmortem note matters only if it names the detection gap.
Incident timeline
14:02 spike on POST /pay in eu-west
14:08 compared traces: upstream 403, not app CPU
14:15 rotated API key for payment provider
14:22 error rate normal; kept enhanced loggingEvidence beats speed
Evidence beats speed without proof. The fix was boring once we looked at the right signal.