Writing1 min read

Field Notes

Debugging Production Without Guessing

Alerts fired. The dashboard looked fine.

Alerts fired. The dashboard looked fine.

Error rate ticked up on checkout. CPU and memory were flat. The team jumped to rollback the last deploy.

Rollback did nothing. The problem was not the release we feared.


Checkout failing in one region

Payments were failing for one region. Pressure was high. Chat filled with theories.


Rollback without proof

We wanted speed without proof. The first assumption felt plausible because the timing matched a deploy.

Evidence was thin. We had not compared failing and healthy requests.


Observe, then change

Observe, isolate, verify, then change. Reproduce the symptom. Narrow scope. Use logs and traces to confirm.

A postmortem note matters only if it names the detection gap.


Incident timeline

Response flow
Short timeline
14:02  spike on POST /pay in eu-west
14:08  compared traces: upstream 403, not app CPU
14:15  rotated API key for payment provider
14:22  error rate normal; kept enhanced logging

Evidence beats speed

Evidence beats speed without proof. The fix was boring once we looked at the right signal.