Runbook · CHK-LAT · Page severity 2
Checkout latency above 800ms
You have been paged because p99 checkout latency has been above 800ms for five minutes. This page is written to be read while that is still happening. Do the triage, then find your symptom. Everything below the first two sections is reference.
On call: payments Escalate after 20 min Last reviewed: 2026-09
Do this first
Triage, in order
-
Is it us, or is it the processor?
chk dash --window 30m --split upstreamIf
processor_waitis most of the latency, this is not our outage. Go to The processor is slow and stop following these steps. -
Did something ship?
chk deploys --since 45mCorrelate the start of the graph with the deploy list. If a checkout-path service deployed within ten minutes of onset, roll it back now and diagnose afterwards. A rollback that turns out to be unnecessary costs four minutes; the alternative costs the incident.
-
Is one dependency doing it?
chk dash --window 30m --split dependencyOne dependency above its own p99 baseline is the usual answer, and it tells you which section below to read. Several at once usually means the database or the network, not the dependencies.
-
Declare, if it is still going
Twenty minutes from page to either recovery or a declared incident. If you are still reading at minute twenty, declare — a declared incident that resolves in a minute costs nothing, and an undeclared one that runs an hour costs a lot.
Find your symptom
What the split is telling you
processor_wait dominates
The card processor is slow and we are waiting on them. We cannot fix it; we can stop making it worse.
Do: raise the circuit breaker threshold so retries stop piling on, and check the processor's status page before escalating to them.
db_wait dominates
Almost always lock contention on orders, and almost always a long-running analytical query that escaped to the primary.
Do: find it with chk db --slow, and kill it. Killing a read query is safe; confirm it is a read first.
inventory_rpc dominates
The inventory service is slow or a pod is unhealthy. Checkout blocks on it, which is a known design problem, not a mystery.
Do: check inventory's own dashboard. If one pod is bad, drain it; if all are, page inventory's on-call.
Nothing dominates
Latency is up everywhere by a similar proportion. That is the platform, not the application: a node pool, the service mesh, or DNS.
Do: stop looking at checkout and page platform on-call with the evidence that it is uniform.
Procedures
The three things you might have to do
chk deploys --since 45m
chk rollback <service> --to <previous>
chk dash --window 10m
Rollback takes about 90 seconds to take effect. Watch for two full minutes before deciding it did not work — the graph lags the fix.
chk breaker show checkout->processor
chk breaker set checkout->processor --failure-ratio 0.5 --window 60s
This makes us more tolerant of a slow processor, which is right when they are degraded and wrong the rest of the time. Put it back when the incident closes, and note it in the incident channel so the next person knows it is not the normal value.
chk db --slow --min 30s
chk db --kill <pid>
Confirm the query is a SELECT before killing it. Killing a write leaves the transaction to roll back, which can take as long as the query has been running and will make things worse for several minutes.
Do not scale checkout up while db_wait is the problem. More replicas means more connections competing for the same locks, and the graph gets worse in a way that looks like the scaling has not taken effect yet.
Do not flush the cache to "start clean". A cold cache under an already-degraded checkout has caused two of the last four incidents on this service, both times during the attempted fix rather than the original fault.
Escalation
Who to wake, and when
| Symptom | Escalate to | When |
|---|---|---|
processor_wait, processor status page is green |
Payments lead | 10 min, they have the vendor contact |
db_wait, no slow query found |
Database on-call | Immediately, this is not a normal cause |
inventory_rpc, all pods unhealthy |
Inventory on-call | Immediately |
| Uniform across dependencies | Platform on-call | Immediately |
| Anything, still going at 20 minutes | Incident commander | 20 min, no exceptions |
What to put in the incident channel
Four things, in this order: what the alert said, what the dependency split shows, what you have already tried, and what you are about to try next. Do not write a narrative — the next person to arrive reads the last message first and needs to know what is in flight, not how you got here.
Checkout, the chk tool, every service name and every threshold on this page are invented. It exists as a Markset example of the shortest and densest genre there is: a page read by someone under time pressure who needs the first screen to be the answer. Source: examples/runbook.md, rendered with examples/runbook.css.
Examples: all · Showcase · Analysis document · Strategy memo · Configuration reference · Architecture overview · Incident review · Capacity review · source