Replay an incident
Replay 1
The deploy wasn't the problem
A deploy eight minutes before the failures, and the evidence ends up at a slow query instead.
Replay it →Replay 2
Two deploys, and the wrong one looked guilty
Two deploys minutes apart, part of the telemetry missing, and the most recent deploy isn't the one that did it.
You're watching this oneYour own
An incident your team has already been through
Upload the errors from around it, say what you found, and compare. No account, nothing saved.
Replay yours →Two deploys, and the wrong one looked guilty
Two deploys go out four minutes apart. Carts start timing out, and the performance metrics that would show why stop arriving just as it begins. The most recent deploy is the obvious suspect. Step through it and watch what each deploy changed settle which one it was.
This is a recording, run through the same rules a live project gets: every confidence and evidence line below is what they say about this data at that minute, with only the telemetry that had arrived by then.
10:11 AM · ForgeOps detects the incident
What just came in
- Incident detected: Rack::Timeout::RequestTimeoutError: Request ran for longer than 15000ms. Errors up 283% (23 in five minutes, after 6 in the five before)
- Deploy 5.3.0 went out 7 minutes before the incident began
- Deploy 5.3.1 went out 3 minutes before the incident began
- Likely cause: Deployment 5.3.1 (81%)
Likely cause at this point: Deployment 5.3.1
This incident appears related to PR #1207 by tom, deployed 3 minutes before the failure.
Evidence
- Deployed 5.3.1 to production 3 minutes before this incident began
- PR #1207: Update the order confirmation email footer
- 83% of errors during the spike referenced 5.3.1
- Error rate up 283%
Also considered
- Deployment 5.3.0 72%
10:12 AM · 1 minute in
What just came in
- Read what deploy 5.3.0 changed: 2 files
- Read what deploy 5.3.1 changed: 2 files
- Likely cause changed from Deployment 5.3.1 (81%) to Deployment 5.3.0 (86%)
Likely cause at this point: Deployment 5.3.0
This incident appears related to PR #1204 by priya, deployed 7 minutes before the failure.
Evidence
- Deployed 5.3.0 to production 7 minutes before this incident began
- Changed summary_builder.rb, which is in every failing stack trace
- PR #1204: Show per-item discounts in the cart summary
- 100% of errors during the spike referenced 5.3.0 or a release after it
- Not the most recent deploy: 5.3.1 went out 4 minutes after it
- Error rate up 283%
Also considered
- Deployment 5.3.1 49% None of the 2 files this deploy changed are in the failing stack traces
10:25 AM · 14 minutes in
What just came in
- Deploy 5.3.2 went out 14 minutes after the incident began
Likely cause at this point: Deployment 5.3.0
This incident appears related to PR #1204 by priya, deployed 7 minutes before the failure.
Evidence
- Deployed 5.3.0 to production 7 minutes before this incident began
- Changed summary_builder.rb, which is in every failing stack trace
- PR #1204: Show per-item discounts in the cart summary
- 100% of errors during the spike referenced 5.3.0 or a release after it
- Not the most recent deploy: 5.3.1 went out 4 minutes after it
- Error rate up 283%
Also considered
- Deployment 5.3.1 52% None of the 2 files this deploy changed are in the failing stack traces
10:26 AM · 15 minutes in
What just came in
- Read what deploy 5.3.2 changed: 2 files
Likely cause at this point: Deployment 5.3.0
This incident appears related to PR #1204 by priya, deployed 7 minutes before the failure.
Evidence
- Deployed 5.3.0 to production 7 minutes before this incident began
- Changed summary_builder.rb, which is in every failing stack trace
- PR #1204: Show per-item discounts in the cart summary
- 95% of errors during the spike referenced 5.3.0 or a release after it
- Not the most recent deploy: 5.3.1 went out 4 minutes after it
- Error rate up 283%
Also considered
- Deployment 5.3.1 51% None of the 2 files this deploy changed are in the failing stack traces
10:34 AM · 23 minutes in
What just came in
- Error rate back to normal: incident resolved
Likely cause at this point: Deployment 5.3.0
This incident appears related to PR #1204 by priya, deployed 7 minutes before the failure.
Evidence
- Deployed 5.3.0 to production 7 minutes before this incident began
- Changed summary_builder.rb, which is in every failing stack trace
- PR #1204: Show per-item discounts in the cart summary
- 88% of errors during the spike referenced 5.3.0 or a release after it
- Not the most recent deploy: 5.3.1 went out 4 minutes after it
- Error rate up 283%
Also considered
- Deployment 5.3.1 49% None of the 2 files this deploy changed are in the failing stack traces
What ForgeOps concluded
It wasn't Deployment 5.3.1
ForgeOps first pointed at Deployment 5.3.1 (81%). 1 minute later, with more of the telemetry in, the evidence pointed at Deployment 5.3.0 (84%) instead. Against Deployment 5.3.1: none of the 2 files this deploy changed are in the failing stack traces.
What the team found
The 10:04 deploy added per-item discounts to the cart summary and loaded each item's discount with a query of its own, so carts with many items ran past the request timeout. The 10:08 deploy only changed an email template. The metrics agent was restarting at the time, so the query timings that would have shown it were never reported. The fix, 5.3.2, loads a cart's discounts in one query.
Inside one request
CartsController#show at 10:09 AM took 3.4s, 3.2s of it in 53 database queries.
Discount Load ran 52 times in that one request: the shape of an N+1 query.
What wasn't there
- No database connection-pool metric was included.
- No performance data was reported from 10:03 AM to 10:28 AM UTC, across the start of the incident, so a slowdown then couldn't be measured.
Replayed from 195 errors, 3 deploys, 64 performance samples, and 54 spans, a minute at a time, through the same detection and likely-cause rules a live project gets. The rules are a small set of named, deterministic checks, not a confirmed diagnosis or AI inference.
Another replay
The deploy wasn't the problem
A deploy eight minutes before the failures, and the evidence ends up at a slow query instead.
Have an incident of your own you're curious about?
Upload the errors from around it (a Sentry export, your logs, or a CSV), say what your team found, and see how ForgeOps would have read it. No account needed, and nothing you upload is saved. Download this incident's file to see the format.