Replay an incident
Replay 1
The deploy wasn't the problem
A deploy eight minutes before the failures, and the evidence ends up at a slow query instead.
You're watching this oneReplay 2
Two deploys, and the wrong one looked guilty
Two deploys minutes apart, part of the telemetry missing, and the most recent deploy isn't the one that did it.
Replay it →Your own
An incident your team has already been through
Upload the errors from around it, say what you found, and compare. No account, nothing saved.
Replay yours →The deploy wasn't the problem
Checkout starts failing eight minutes after a deploy. The deploy is the obvious suspect, the database telemetry is incomplete, and the first answer is wrong. Step through it and watch the evidence change the answer.
This is a recording, run through the same rules a live project gets: every confidence and evidence line below is what they say about this data at that minute, with only the telemetry that had arrived by then.
10:42 AM · ForgeOps detects the incident
What just came in
- Incident detected: ActiveRecord::ConnectionTimeoutError: could not obtain a connection from the pool within 5.000 seconds; all pooled connections were in use. Errors up 360% (23 in five minutes, after 5 in the five before)
- Deploy 4.8.2 went out 8 minutes before the incident began
- Likely cause: Deployment 4.8.2 (68%)
Likely cause at this point: Deployment 4.8.2
This incident appears related to PR #812 by maya, deployed 8 minutes before the failure.
Evidence
- Deployed 4.8.2 to production 8 minutes before this incident began
- PR #812: Refresh the promo banner copy
- 100% of errors during the spike referenced 4.8.2
- Error rate up 360%
- No sudden slowdown in the database or outside services when this started
10:43 AM · 1 minute in
What just came in
- Read what deploy 4.8.2 changed: 3 files
- Confidence in Deployment 4.8.2 dropped from 68% to 41%
Likely cause at this point: Deployment 4.8.2
This incident appears related to PR #812 by maya, deployed 8 minutes before the failure.
Evidence
- Deployed 4.8.2 to production 8 minutes before this incident began
- PR #812: Refresh the promo banner copy
- 100% of errors during the spike referenced 4.8.2
- Error rate up 360%
- No sudden slowdown in the database or outside services when this started
- Against: None of the 3 files this deploy changed are in the failing stack traces
10:44 AM · 2 minutes in
What just came in
- CheckoutController#create latency increased (+900%)
- LineItem Load query latency increased (+900%)
- Likely cause changed from Deployment 4.8.2 (41%) to LineItem Load query latency (90%)
Likely cause at this point: LineItem Load query latency
Database query latency spiked 3 minutes before this incident began.
Evidence
- LineItem Load query latency +900%
- No PostgreSQL connection-pool data available
- No sudden slowdown in outside services when this started
Also considered
- Deployment 4.8.2 41% None of the 3 files this deploy changed are in the failing stack traces
10:53 AM · 11 minutes in
What just came in
- Deploy 4.8.3 went out 11 minutes after the incident began
Likely cause at this point: LineItem Load query latency
Database query latency spiked 3 minutes before this incident began.
Evidence
- LineItem Load query latency +900%
- No PostgreSQL connection-pool data available
- No sudden slowdown in outside services when this started
Also considered
- Deployment 4.8.2 41% None of the 3 files this deploy changed are in the failing stack traces
10:54 AM · 12 minutes in
What just came in
- Read what deploy 4.8.3 changed: 2 files
Likely cause at this point: LineItem Load query latency
Database query latency spiked 3 minutes before this incident began.
Evidence
- LineItem Load query latency +900%
- No PostgreSQL connection-pool data available
- No sudden slowdown in outside services when this started
Also considered
- Deployment 4.8.2 39% None of the 3 files this deploy changed are in the failing stack traces
11:04 AM · 22 minutes in
What just came in
- Error rate back to normal: incident resolved
Likely cause at this point: LineItem Load query latency
Database query latency spiked 3 minutes before this incident began.
Evidence
- LineItem Load query latency +900%
- No PostgreSQL connection-pool data available
- No sudden slowdown in outside services when this started
Also considered
- Deployment 4.8.2 38% None of the 3 files this deploy changed are in the failing stack traces
What ForgeOps concluded
It wasn't Deployment 4.8.2
ForgeOps first pointed at Deployment 4.8.2 (68%). 2 minutes later, with more of the telemetry in, the evidence pointed at LineItem Load query latency (90%) instead. Against Deployment 4.8.2: none of the 3 files this deploy changed are in the failing stack traces.
What the team found
A wholesale customer started checking out carts of 47 line items, where 3 is typical. The price calculator loaded each line item with a query of its own, as it had for months, and those requests held database connections long enough to starve everyone else's. The 10:34 deploy only changed the promo banner. The fix, 4.8.3, loads a cart's line items in one query.
Inside one request
CheckoutController#create at 10:40 AM took 3.2s, 2.8s of it in 48 database queries and 0.2s in outside services.
LineItem Load ran 47 times in that one request: the shape of an N+1 query.
What wasn't there
- No database connection-pool metric was included.
Replayed from 229 errors, 2 deploys, 72 performance samples, and 50 spans, a minute at a time, through the same detection and likely-cause rules a live project gets. The rules are a small set of named, deterministic checks, not a confirmed diagnosis or AI inference.
Another replay
Two deploys, and the wrong one looked guilty
Two deploys minutes apart, part of the telemetry missing, and the most recent deploy isn't the one that did it.
Have an incident of your own you're curious about?
Upload the errors from around it (a Sentry export, your logs, or a CSV), say what your team found, and see how ForgeOps would have read it. No account needed, and nothing you upload is saved. Download this incident's file to see the format.