Python error tracking
Know when your Python app breaks, before your users do
ForgeOps connects the errors your Python app reports to the deploys, traces, and changes around them, so when production breaks, the incident already says who's affected, what changed, and where to look.
Django and Flask integrations report every unhandled request exception automatically. Running on ForgeOps' own dedicated infrastructure, so nothing about a Python app's errors or its users' data gets routed through a third-party error-tracking vendor. See an incident investigated →
Install
pip install forge-ops-tracker import forge_ops_tracker forge_ops_tracker.init(dsn="https://<api_key>@getforgeops.net/api/v1/events")
Django and Flask integrations report every unhandled request exception automatically. Python 3.9 or later.
Every client reports the same shape
Wherever it comes from, an event arrives with an exception class, a message, and a backtrace, plus whatever environment/release/server context that client can gather on its own. ForgeOps fingerprints and groups on the first three, so a Python app's issues list works exactly the same way every other language's does: report the same bug a thousand times and it's still one row, not a thousand.
No agent, no sidecar: it's a plain HTTPS POST to ForgeOps' own ingestion API, so nothing about a Python app's traffic gets routed through a third-party pipeline first.
Works with your framework
The Python client hooks in at the framework level, not by wrapping individual calls, so nothing about how you already write Python code has to change.
Django
Middleware hooks Django's own exception handling; no changes to individual views needed.
Flask
Wraps the WSGI app itself, so every unhandled request exception reports automatically.
Celery
Reports a task that raises, with the task's name, the same as a request.
Scripts and workers
The environment defaults to production, so FORGE_OPS_DSN alone sends; set FORGE_OPS_ENVIRONMENT=development where you don't want it to (needs Python 3.9 or later, and forge-ops-tracker 0.17.0).
What it looks like
The issue list
Every Python issue reported into a project shows up here: title, status, assignee, event count, and when it was last seen, the same columns as any other language's project.
The issue detail
Every occurrence keeps its own timestamp, environment, release, and server, plus the backtrace exactly as Python reported it, expandable per occurrence rather than flattened into one generic stack.
Alerting
A new Python issue, a regression, or a spike past a threshold you set can reach Slack, Microsoft Teams, email, PagerDuty, Opsgenie, or a generic webhook, configured per project, per trigger.
Heartbeats
For the failure mode error tracking can't see on its own: a Python cron job or recurring task that's supposed to run but silently stopped. Each heartbeat gets its own ping URL and its own grace period before it alerts.
Background jobs
A script or worker's own crash needs no wiring either: init() also installs a sys.excepthook that reports anything that crashes the interpreter outright, and anything still queued, the crash included, is sent before the process exits (waiting at most 5 seconds).
Celery and RQ don't let a failing task reach it though: both catch the task's exception themselves, to mark it failed and keep the worker alive, so it never becomes an uncaught interpreter-level crash. Celery's own task_failure signal (RQ has an equivalent exception handler) needs wiring instead:
from celery.signals import task_failure
@task_failure.connect
def report_task_failure(
sender=None, exception=None, **kwargs
):
forge_ops_tracker.capture_exception(
exception,
context={"task": sender.name},
)
Release health
Once your app is reporting into ForgeOps, every request through the Django or Flask integration also counts as a session: crash-free unless an unhandled exception actually affects it. That gives each release's row on a project's Releases page a real crash-free rate, not just an event count. On by default; counted in-process and flushed as one small aggregate report every 60 seconds, not one network call per request. Requires a plan that includes release health; on a plan that doesn't, the periodic reports are accepted but not recorded (the response says why), so a working setup never looks broken.
forge_ops_tracker.init(
dsn="...",
track_sessions=False, # opt out entirely
session_flush_interval=30, # seconds; default 60
)
Performance monitoring
Once the Django middleware below is added (Flask needs no separate step: init_flask(app) already covers it), every request also times its own duration, bucketed by transaction (the matched URL pattern, e.g. "users/<int:id>/", rather than the literal path, so a distinct user id doesn't explode into its own separate transaction) and flushed as a small periodic aggregate report every 60 seconds, the same delivery philosophy as release health above. Requires a plan that includes performance monitoring; on a plan that doesn't, the periodic reports are accepted but not recorded (the response says why), so a working setup never looks broken. Each report also carries a small latency histogram (forge-ops-tracker 0.10.0 and later), so ForgeOps shows an approximate p50/p95/p99 per transaction, not just an average, accurate to the width of the latency bucket a duration falls into.
MIDDLEWARE = [
...,
"forge_ops_tracker.integrations.django.ForgeOpsTrackerPerformanceTrackingMiddleware",
]
forge_ops_tracker.init(
dsn="...",
track_performance=False, # opt out entirely
performance_flush_interval=30, # seconds; default 60
)
Distributed tracing
For one slow request, the Django and Flask integrations (ForgeOpsTrackerTracingMiddleware and init_flask, both covered above) capture its full nested call tree: the view, plus every database query, Redis command, and outbound requests call under it. A Redis span is named after the command alone (Redis GET), never its keys or values, and a pipeline is one span. Only a slow request's trace is ever sent: whether it crossed the threshold (one second by default, configurable) is decided entirely inside your process, so a fast request costs nothing over the wire. Redis spans need forge-ops-tracker 0.15.0. Requires a plan that includes distributed tracing.
forge_ops_tracker.init(
dsn="...",
track_tracing=False, # opt out entirely
trace_capture_threshold_ms=500, # milliseconds; default 1000
)
from forge_ops_tracker.integrations.requests import init_requests
init_requests() # outbound requests-library calls nest in too
from forge_ops_tracker.integrations.redis import init_redis
init_redis() # redis-py commands too; pip install "forge-ops-tracker[redis]"
Wrap your own service-layer code by hand to give it its own span. It also works as a decorator, and it does nothing outside a traced request.
with forge_ops_tracker.span("PaymentService.charge"):
charge_card(order)
Errors now carry the request they happened in: its endpoint (POST /checkout) and its trace id. An issue names its affected endpoint, each occurrence links to its exact trace, and a request that fails is always traced, however fast. The trace follows the request across services in the standard traceparent header, picked up from whatever called your Django or Flask app and passed on to whatever it calls through requests, so once the projects are linked in ForgeOps an error shows the errors another project raised in the same request. Needs forge-ops-tracker 0.12.0.
from forge_ops_tracker.integrations.requests import init_requests
forge_ops_tracker.init(
dsn="...",
propagate_traces=True, # the default
trace_propagation_targets=["api.example.com"], # only these hosts; None (the default) means all
)
init_requests() # outbound requests calls carry the header
SQL in a request
When the request an issue failed in was traced, the issue opens with its slowest query: how long it took, its share of the request's time, how many queries the request ran, and a likely N+1 warning when the same statement ran five or more times. Every query the Django and SQLAlchemy integrations record already carries its SQL, masked so every string and number becomes a question mark and bind values are never sent. Turn on explain_slow_queries and a slow SELECT on PostgreSQL gets its plan too, shown beside the query with any sequential scan over a large table pointed out; it runs plain EXPLAIN, never EXPLAIN ANALYZE, and only for a single plain SELECT, off the request on its own connection inside a read-only transaction with a two second timeout, at most once per statement every ten minutes. Needs forge-ops-tracker 0.14.0; plans need a plan that includes performance monitoring.
forge_ops_tracker.init(
dsn="...",
explain_slow_queries=True, # default False; PostgreSQL only
explain_threshold_ms=500, # milliseconds; default 500
)
# Django's tracing middleware already records every query; a SQLAlchemy engine opts in:
from forge_ops_tracker.integrations.sqlalchemy import track_sqlalchemy_queries
track_sqlalchemy_queries(engine)
What changed
A feature flag flip or a config edit can break things with no deploy at all. Record one with record_change and it shows under What changed on any issue that starts in the next two hours, and on the project's Changes page beside your deploys, labeled potentially relevant, never as the cause. Changes between deploys are detected for you: once per process, init() sends the Python version and every installed package version from a background thread, so a package upgrade shows up with nothing to call and startup never waits on it. Environment variable names (never values) are opt-in. Needs forge-ops-tracker 0.13.0, on a plan that includes change tracking.
forge_ops_tracker.init(
dsn="...",
detect_changes=True, # the default; False sends no startup snapshot
track_env_var_names=True, # default False; names only, never values
)
forge_ops_tracker.record_change(
"feature_flag",
"gateway_retry_v2 turned on for 100% of checkouts",
details={"flag": "gateway_retry_v2", "from": "10%", "to": "100%"},
actor="priya@example.com",
)
Custom metrics and infrastructure monitoring
Track a business event you name yourself (a signup, a payment) or a reading from one of your own hosts, with one explicit call each; nothing is automatic. The value defaults to 1, so a bare call is a counter; pass a real amount for anything else (it can be negative, for a refund). A captured value is stored as its own row, so counts and sums you compute later are exact. Calls are buffered and flushed as one batch in the background, so they are safe to make inside a request, and flush_metrics sends whatever is buffered right now. The hostname of an infrastructure reading defaults to the configured server name. Requires a plan that includes custom metrics and infrastructure monitoring.
forge_ops_tracker.capture_metric("signup") # value defaults to 1.0: a bare counter
forge_ops_tracker.capture_metric("payment", 49.0) # a real magnitude; it may be negative (a refund)
forge_ops_tracker.capture_infrastructure_metric("cpu", 0.42) # hostname defaults to server_name
forge_ops_tracker.capture_infrastructure_metric("disk", 0.81, hostname="db-1")
forge_ops_tracker.flush_metrics() # optional: send right now
Database errors
When an error comes from a database call, the issue shows which stored procedure, table or view its SQL touched, so you know where to start looking. SQLAlchemy errors carry the statement, so it's read off them, or off the exception they were raised from, with nothing to add. Django's database errors carry none, so a Django app sees only the names it can find in the error text. The names are sent by default and are identifiers, never values. The SQL statement itself is opt-in, with every string and number replaced by a question mark before it leaves your app, and a project setting can stop ForgeOps storing the statement at all.
forge_ops_tracker.init(dsn="...", capture_sql_statement=True) # default False # capture_sql_objects=False (default True) stops even the names
Questions
Do I need to change how I already handle errors?
Almost never. Every client hooks the framework's own exception handling directly, the same way the install snippet above shows, so your own logs, error pages, and rescue blocks keep working exactly as they did before, and ForgeOps just also finds out about it. The one common exception is a background-job runner that doesn't forward through that same mechanism on its own; that needs one small hook added instead, not a change to anything already there.
Where does the data actually go?
Straight to ForgeOps' own ingestion API over plain HTTPS. No agent, no sidecar, and no third-party ingestion service in between, so nothing about your app's traffic or its users' data gets routed through anyone else's pipeline first.
What happens if the same bug fires a thousand times?
It's fingerprinted from its exception class, message, and backtrace, so a thousand reports of the same bug still show up as one issue with an event count of 1,000, not a thousand separate rows to dig through. And a runaway loop can't use up your monthly event quota: once one issue repeats faster than your plan's hourly limit (50 an hour on Free), further repeats are still counted, and still count toward spike alerts, but they aren't stored or charged to your quota.
Is anything scrubbed before it's stored?
Yes. Emails, credit card numbers, and known API key/token formats are redacted out of every event before it's even written, on every plan, with room for per-project custom field names on top of the defaults.
Try it with your own Python app
Free plan included, no credit card required. New organizations start with 14 days of Business.
Get started freeGet