write up this morning's checkout incident for the team

be clear about impact, cause, and what changes next

Checkout recovered at 10:17 UTC after 31 minutes of elevated failures; 18.6% of purchase attempts failed, with no evidence of incorrect charges.1

Impact

MeasureResultConfidence
Incident window09:46-10:17 UTCconfirmed from deploy and alert logs4
Failed attempts1,284confirmed from checkout events1
Affected customers973deduplicated by customer ID1
Incorrect charges0 observedpayment ledger reconciled5

Timeline

UTCEventDecision
09:46Provider latency crosses 4 secondsAlert fires after error rate exceeds 5%
09:51On-call confirms retry amplificationIncident declared SEV-2
10:02Recent application deploy ruled outFocus shifts to provider and retry path
10:11Synchronous retries disabledError rate begins falling immediately
10:17Checkout returns below 1% failuresIncident moves to monitoring

Root cause

The provider slowdown was the trigger, not the complete cause. Checkout retried timed-out authorization requests three times inside the original request. Each attempt consumed another worker slot, so a partial provider slowdown became local queue saturation.2

The feature flag was intended to improve recovery from brief network failures. It had no total time budget and was enabled globally without a saturation guard.3

What changes

OwnerChangeDueVerification
PaymentsMove retries to the background queueJul 14Load test holds p95 below 2s during injected timeouts
PlatformAdd provider-specific circuit breakerJul 16Breaker opens before worker utilization reaches 75%
DataAdd failed-checkout recovery cohortJul 12Support can identify and contact affected customers
On-callAdd retry-amplification runbookJul 11Game day completed by a non-author

Customer follow-up

Send a recovery email only to customers whose final attempt failed and who did not complete a purchase within 24 hours. Do not message customers who recovered on retry or imply that they were charged.5