Replay an incident

Replay 1

The deploy wasn't the problem

A deploy eight minutes before the failures, and the evidence ends up at a slow query instead.

You're watching this one

Replay 2

Two deploys, and the wrong one looked guilty

Two deploys minutes apart, part of the telemetry missing, and the most recent deploy isn't the one that did it.

Replay it →

Your own

An incident your team has already been through

Upload the errors from around it, say what you found, and compare. No account, nothing saved.

Replay yours →

The deploy wasn't the problem

Checkout starts failing eight minutes after a deploy. The deploy is the obvious suspect, the database telemetry is incomplete, and the first answer is wrong. Step through it and watch the evidence change the answer.

This is a recording, run through the same rules a live project gets: every confidence and evidence line below is what they say about this data at that minute, with only the telemetry that had arrived by then.

Errors
+360%
Customers affected
37
Detected
10:42 AM
Resolved
11:04 AM (22 minutes)

10:42 AM · ForgeOps detects the incident

What just came in

  • Incident detected: ActiveRecord::ConnectionTimeoutError: could not obtain a connection from the pool within 5.000 seconds; all pooled connections were in use. Errors up 360% (23 in five minutes, after 5 in the five before)
  • Deploy 4.8.2 went out 8 minutes before the incident began
  • Likely cause: Deployment 4.8.2 (68%)

Likely cause at this point: Deployment 4.8.2

This incident appears related to PR #812 by maya, deployed 8 minutes before the failure.

68%
Match confidence

Evidence

  • Deployed 4.8.2 to production 8 minutes before this incident began
  • PR #812: Refresh the promo banner copy
  • 100% of errors during the spike referenced 4.8.2
  • Error rate up 360%
  • No sudden slowdown in the database or outside services when this started

10:43 AM · 1 minute in

What just came in

  • Read what deploy 4.8.2 changed: 3 files
  • Confidence in Deployment 4.8.2 dropped from 68% to 41%

Likely cause at this point: Deployment 4.8.2

This incident appears related to PR #812 by maya, deployed 8 minutes before the failure.

41%
Match confidence

Evidence

  • Deployed 4.8.2 to production 8 minutes before this incident began
  • PR #812: Refresh the promo banner copy
  • 100% of errors during the spike referenced 4.8.2
  • Error rate up 360%
  • No sudden slowdown in the database or outside services when this started
  • Against: None of the 3 files this deploy changed are in the failing stack traces

10:44 AM · 2 minutes in

What just came in

  • CheckoutController#create latency increased (+900%)
  • LineItem Load query latency increased (+900%)
  • Likely cause changed from Deployment 4.8.2 (41%) to LineItem Load query latency (90%)

Likely cause at this point: LineItem Load query latency

Database query latency spiked 3 minutes before this incident began.

90%
Match confidence

Evidence

  • LineItem Load query latency +900%
  • No PostgreSQL connection-pool data available
  • No sudden slowdown in outside services when this started

Also considered

  • Deployment 4.8.2 41% None of the 3 files this deploy changed are in the failing stack traces

10:53 AM · 11 minutes in

What just came in

  • Deploy 4.8.3 went out 11 minutes after the incident began

Likely cause at this point: LineItem Load query latency

Database query latency spiked 3 minutes before this incident began.

90%
Match confidence

Evidence

  • LineItem Load query latency +900%
  • No PostgreSQL connection-pool data available
  • No sudden slowdown in outside services when this started

Also considered

  • Deployment 4.8.2 41% None of the 3 files this deploy changed are in the failing stack traces

10:54 AM · 12 minutes in

What just came in

  • Read what deploy 4.8.3 changed: 2 files

Likely cause at this point: LineItem Load query latency

Database query latency spiked 3 minutes before this incident began.

90%
Match confidence

Evidence

  • LineItem Load query latency +900%
  • No PostgreSQL connection-pool data available
  • No sudden slowdown in outside services when this started

Also considered

  • Deployment 4.8.2 39% None of the 3 files this deploy changed are in the failing stack traces

11:04 AM · 22 minutes in

What just came in

  • Error rate back to normal: incident resolved

Likely cause at this point: LineItem Load query latency

Database query latency spiked 3 minutes before this incident began.

90%
Match confidence

Evidence

  • LineItem Load query latency +900%
  • No PostgreSQL connection-pool data available
  • No sudden slowdown in outside services when this started

Also considered

  • Deployment 4.8.2 38% None of the 3 files this deploy changed are in the failing stack traces

What ForgeOps concluded

It wasn't Deployment 4.8.2

ForgeOps first pointed at Deployment 4.8.2 (68%). 2 minutes later, with more of the telemetry in, the evidence pointed at LineItem Load query latency (90%) instead. Against Deployment 4.8.2: none of the 3 files this deploy changed are in the failing stack traces.

What the team found

A wholesale customer started checking out carts of 47 line items, where 3 is typical. The price calculator loaded each line item with a query of its own, as it had for months, and those requests held database connections long enough to starve everyone else's. The 10:34 deploy only changed the promo banner. The fix, 4.8.3, loads a cart's line items in one query.

Inside one request

CheckoutController#create at 10:40 AM took 3.2s, 2.8s of it in 48 database queries and 0.2s in outside services.

LineItem Load ran 47 times in that one request: the shape of an N+1 query.

What wasn't there

  • No database connection-pool metric was included.

Replayed from 229 errors, 2 deploys, 72 performance samples, and 50 spans, a minute at a time, through the same detection and likely-cause rules a live project gets. The rules are a small set of named, deterministic checks, not a confirmed diagnosis or AI inference.

Another replay

Two deploys, and the wrong one looked guilty

Two deploys minutes apart, part of the telemetry missing, and the most recent deploy isn't the one that did it.

Replay it

Have an incident of your own you're curious about?

Upload the errors from around it (a Sentry export, your logs, or a CSV), say what your team found, and see how ForgeOps would have read it. No account needed, and nothing you upload is saved. Download this incident's file to see the format.