Replay an incident

Replay 1

The deploy wasn't the problem

A deploy eight minutes before the failures, and the evidence ends up at a slow query instead.

Replay it →

Replay 2

Two deploys, and the wrong one looked guilty

Two deploys minutes apart, part of the telemetry missing, and the most recent deploy isn't the one that did it.

You're watching this one

Your own

An incident your team has already been through

Upload the errors from around it, say what you found, and compare. No account, nothing saved.

Replay yours →

Two deploys, and the wrong one looked guilty

Two deploys go out four minutes apart. Carts start timing out, and the performance metrics that would show why stop arriving just as it begins. The most recent deploy is the obvious suspect. Step through it and watch what each deploy changed settle which one it was.

This is a recording, run through the same rules a live project gets: every confidence and evidence line below is what they say about this data at that minute, with only the telemetry that had arrived by then.

Errors
+283%
Customers affected
31
Detected
10:11 AM
Resolved
10:34 AM (23 minutes)

10:11 AM · ForgeOps detects the incident

What just came in

  • Incident detected: Rack::Timeout::RequestTimeoutError: Request ran for longer than 15000ms. Errors up 283% (23 in five minutes, after 6 in the five before)
  • Deploy 5.3.0 went out 7 minutes before the incident began
  • Deploy 5.3.1 went out 3 minutes before the incident began
  • Likely cause: Deployment 5.3.1 (81%)

Likely cause at this point: Deployment 5.3.1

This incident appears related to PR #1207 by tom, deployed 3 minutes before the failure.

81%
Match confidence

Evidence

  • Deployed 5.3.1 to production 3 minutes before this incident began
  • PR #1207: Update the order confirmation email footer
  • 83% of errors during the spike referenced 5.3.1
  • Error rate up 283%

Also considered

  • Deployment 5.3.0 72%

10:12 AM · 1 minute in

What just came in

  • Read what deploy 5.3.0 changed: 2 files
  • Read what deploy 5.3.1 changed: 2 files
  • Likely cause changed from Deployment 5.3.1 (81%) to Deployment 5.3.0 (86%)

Likely cause at this point: Deployment 5.3.0

This incident appears related to PR #1204 by priya, deployed 7 minutes before the failure.

86%
Match confidence

Evidence

  • Deployed 5.3.0 to production 7 minutes before this incident began
  • Changed summary_builder.rb, which is in every failing stack trace
  • PR #1204: Show per-item discounts in the cart summary
  • 100% of errors during the spike referenced 5.3.0 or a release after it
  • Not the most recent deploy: 5.3.1 went out 4 minutes after it
  • Error rate up 283%

Also considered

  • Deployment 5.3.1 49% None of the 2 files this deploy changed are in the failing stack traces

10:25 AM · 14 minutes in

What just came in

  • Deploy 5.3.2 went out 14 minutes after the incident began

Likely cause at this point: Deployment 5.3.0

This incident appears related to PR #1204 by priya, deployed 7 minutes before the failure.

86%
Match confidence

Evidence

  • Deployed 5.3.0 to production 7 minutes before this incident began
  • Changed summary_builder.rb, which is in every failing stack trace
  • PR #1204: Show per-item discounts in the cart summary
  • 100% of errors during the spike referenced 5.3.0 or a release after it
  • Not the most recent deploy: 5.3.1 went out 4 minutes after it
  • Error rate up 283%

Also considered

  • Deployment 5.3.1 52% None of the 2 files this deploy changed are in the failing stack traces

10:26 AM · 15 minutes in

What just came in

  • Read what deploy 5.3.2 changed: 2 files

Likely cause at this point: Deployment 5.3.0

This incident appears related to PR #1204 by priya, deployed 7 minutes before the failure.

85%
Match confidence

Evidence

  • Deployed 5.3.0 to production 7 minutes before this incident began
  • Changed summary_builder.rb, which is in every failing stack trace
  • PR #1204: Show per-item discounts in the cart summary
  • 95% of errors during the spike referenced 5.3.0 or a release after it
  • Not the most recent deploy: 5.3.1 went out 4 minutes after it
  • Error rate up 283%

Also considered

  • Deployment 5.3.1 51% None of the 2 files this deploy changed are in the failing stack traces

10:34 AM · 23 minutes in

What just came in

  • Error rate back to normal: incident resolved

Likely cause at this point: Deployment 5.3.0

This incident appears related to PR #1204 by priya, deployed 7 minutes before the failure.

84%
Match confidence

Evidence

  • Deployed 5.3.0 to production 7 minutes before this incident began
  • Changed summary_builder.rb, which is in every failing stack trace
  • PR #1204: Show per-item discounts in the cart summary
  • 88% of errors during the spike referenced 5.3.0 or a release after it
  • Not the most recent deploy: 5.3.1 went out 4 minutes after it
  • Error rate up 283%

Also considered

  • Deployment 5.3.1 49% None of the 2 files this deploy changed are in the failing stack traces

What ForgeOps concluded

It wasn't Deployment 5.3.1

ForgeOps first pointed at Deployment 5.3.1 (81%). 1 minute later, with more of the telemetry in, the evidence pointed at Deployment 5.3.0 (84%) instead. Against Deployment 5.3.1: none of the 2 files this deploy changed are in the failing stack traces.

What the team found

The 10:04 deploy added per-item discounts to the cart summary and loaded each item's discount with a query of its own, so carts with many items ran past the request timeout. The 10:08 deploy only changed an email template. The metrics agent was restarting at the time, so the query timings that would have shown it were never reported. The fix, 5.3.2, loads a cart's discounts in one query.

Inside one request

CartsController#show at 10:09 AM took 3.4s, 3.2s of it in 53 database queries.

Discount Load ran 52 times in that one request: the shape of an N+1 query.

What wasn't there

  • No database connection-pool metric was included.
  • No performance data was reported from 10:03 AM to 10:28 AM UTC, across the start of the incident, so a slowdown then couldn't be measured.

Replayed from 195 errors, 3 deploys, 64 performance samples, and 54 spans, a minute at a time, through the same detection and likely-cause rules a live project gets. The rules are a small set of named, deterministic checks, not a confirmed diagnosis or AI inference.

Another replay

The deploy wasn't the problem

A deploy eight minutes before the failures, and the evidence ends up at a slow query instead.

Replay it

Have an incident of your own you're curious about?

Upload the errors from around it (a Sentry export, your logs, or a CSV), say what your team found, and see how ForgeOps would have read it. No account needed, and nothing you upload is saved. Download this incident's file to see the format.