Android error tracking

Know when your Android app breaks, before your users do

ForgeOps connects the errors your Android app reports to the deploys, traces, and changes around them, so when production breaks, the incident already says who's affected, what changed, and where to look.

Installs an uncaught-exception handler and an ANR watchdog automatically; every report is written to disk and uploaded on the next launch. Running on ForgeOps' own dedicated infrastructure, so nothing about a Android app's errors or its users' data gets routed through a third-party error-tracking vendor. See an incident investigated →

Install

// module build.gradle.kts (Maven Central is already in a default Android project's repositories)
dependencies { implementation("io.github.luke-popwell.tracker.android:forgeopstracker:0.12.0") }

ForgeOpsTracker.init(this) { config ->
    config.dsn = "https://<api_key>@getforgeops.net/api/v1/events"
}
ForgeOpsTracker.installHandlers()

Installs an uncaught-exception handler and an ANR watchdog automatically; every report is written to disk and uploaded on the next launch.

Every client reports the same shape

Wherever it comes from, an event arrives with an exception class, a message, and a backtrace, plus whatever environment/release/server context that client can gather on its own. ForgeOps fingerprints and groups on the first three, so a Android app's issues list works exactly the same way every other language's does: report the same bug a thousand times and it's still one row, not a thousand.

No agent, no sidecar: it's a plain HTTPS POST to ForgeOps' own ingestion API, so nothing about a Android app's traffic gets routed through a third-party pipeline first.


Works with your framework


The Android client hooks in at the framework level, not by wrapping individual calls, so nothing about how you already write Android code has to change.


Uncaught exceptions

A Thread.UncaughtExceptionHandler chains onto whatever handler, if any, was already installed.

ANR watchdog

A background thread pings the main Looper; a blocked main thread reports as an ANR, the standard technique real Android crash reporters use.

Caught exceptions

captureException reports directly from a catch block, with optional context.

What it looks like

The issue list

Every Android issue reported into a project shows up here: title, status, assignee, event count, and when it was last seen, the same columns as any other language's project.

The issue list for a project reporting from an Android app

The issue detail

Every occurrence keeps its own timestamp, environment, release, and server, plus the backtrace exactly as Android reported it, expandable per occurrence rather than flattened into one generic stack.

An Android issue's page in ForgeOps: its title, the exception class and the code it came from beneath it, then its verdict with its failure count, customers affected, and when it was first seen, above the triage buttons

Alerting

A new Android issue, a regression, or a spike past a threshold you set can reach Slack, Microsoft Teams, email, PagerDuty, Opsgenie, or a generic webhook, configured per project, per trigger.

Notification rules configured per trigger and channel, e.g. a new issue to Slack, a regression to email

Heartbeats

For the failure mode error tracking can't see on its own: a Android cron job or recurring task that's supposed to run but silently stopped. Each heartbeat gets its own ping URL and its own grace period before it alerts.

A heartbeat's ping URL and status, showing its expected interval, grace period, and last check-in

Performance monitoring

Times whatever you wrap and reports one small aggregate per transaction, flushed on Configuration.performanceFlushIntervalMillis, never one network call per timed call. This client has no web framework integration, so nothing is timed automatically: you choose what to wrap. Keep transaction names low-cardinality ("HomeScreen.load", not one per item id), since every distinct name is its own row. Safe to call from any thread; a call before init is a silent no-op. Turn it off with trackPerformance = false. An Android process is killed by the OS rather than exited, so nothing is flushed at shutdown: call ForgeOpsTracker.flushPerformance() when your app goes to the background (from Activity.onStop, say). It delivers on a background thread, never the caller's; flushPerformanceSync() blocks until delivery finishes, for a host app that wants its own scheduling. Each report also carries a small latency histogram, so ForgeOps shows an approximate p50/p95/p99 per transaction, not just an average, accurate to the width of the latency bucket a duration falls into. Requires a plan that includes performance monitoring; on a plan that doesn't, the periodic reports are accepted but not recorded (the response says why), so a working setup never looks broken.

The Performance page, showing each transaction's request count and p50, p95 and p99 latency, plus slow queries
// Wrap a block; recorded even if it throws (the error propagates unchanged), and its value comes back:
val user = ForgeOpsTracker.timeTransaction("HomeScreen.load") { repository.loadUser() }

// Or record a duration you measured yourself, in milliseconds:
ForgeOpsTracker.recordPerformance("Checkout.pay", elapsedMs.toDouble())

Distributed tracing

One flow's own call tree, such as a screen load, a sign-in, or a network round trip and what it triggered. An Android flow hops between the main thread, background threads, and coroutines, so a trace is an explicit object you pass around or capture in a lambda, safe to use from any thread. Only a slow flow's trace is ever sent: whether it crossed the threshold (one second by default, configurable) is decided entirely inside your app, so a fast flow costs nothing over the wire. Requires a plan that includes distributed tracing.

A trace's waterfall for one slow request, with its controller, service, database, cache and HTTP spans
ForgeOpsTracker.trace("HomeScreen.load") { trace ->
    val feed = trace.measureSpan("fetch feed", kind = "http") { api.fetchFeed() }
    trace.measureSpan("decode", data = mapOf("items" to feed.size)) { decode(feed) }
}

// Or hold the trace across threads and finish it when the flow ends:
val trace = ForgeOpsTracker.startTrace("Checkout.pay")
thread { trace.recordSpan("charge", kind = "http", startedAtMillis = started, durationMs = ms) }
trace.finish()

Make a request inside a trace with measureHttpSpan (or startHttpSpan for a callback client) and it's recorded as its own http span and handed the traceparent header to send, so your backend's request nests under it. An error captured while the trace is open carries its trace id: with your backend on the current ForgeOps server SDK and the two projects linked in ForgeOps, the failed payment shows the API error from the very same request. Needs forgeopstracker 0.8.0.

A mobile app's checkout failure listing the API's gateway timeout as the same request, matched by trace ID
ForgeOpsTracker.trace("Checkout.pay") { trace ->
    try {
        val status = trace.measureHttpSpan("POST", payUrl) { headers ->
            val connection = URL(payUrl).openConnection() as HttpURLConnection
            connection.requestMethod = "POST"
            headers.forEach { (name, value) -> connection.setRequestProperty(name, value) }
            connection.responseCode
        }
        if (status >= 400) throw PaymentFailedException("payment returned $status")
    } catch (e: Exception) {
        ForgeOpsTracker.captureException(e, mapOf("orderId" to order.id)) // carries this trace's id
    }
}

SQL in a request

When an error is captured inside a trace that ran database queries, its issue opens with the slowest of them: how long it took, its share of the trace's time, how many queries ran, and a likely N+1 warning when the same statement ran five or more times. Pass a query's SQL as statement on a database span (a query against the app's own SQLite database, say). The statement is masked before it leaves your device, every string and number replaced with a question mark, and the same SQL shows under that span's bar in the trace's waterfall. Needs forgeopstracker 0.10.0, on a plan that includes distributed tracing.

An issue's Slowest query in this request card: a 2.7 second order search that took 86% of a 3.2 second request, a likely N+1 warning for a line_items query run 24 times, and the query plan calling out a sequential scan on orders
val sql = "SELECT * FROM messages WHERE thread_id = 42 AND read = 0"

ForgeOpsTracker.trace("Inbox.load") { trace ->
    val cursor = trace.measureSpan("Load messages", kind = "database", statement = sql, dbSystem = "sqlite") {
        database.rawQuery(sql, null)
    }
    show(cursor)
}
// Sent as "SELECT * FROM messages WHERE thread_id = ? AND read = ?"

What changed

When a feature flag flips or a remote config value changes, record it from your flag client's callback, and it shows under What changed on any issue that starts in the next two hours, and on the project's Changes page beside your releases, labeled potentially relevant, never as the cause. "Crashes started right after new_checkout turned on" becomes one glance. recordChange returns immediately and sends on a private daemon thread, so it's safe from the main thread, and it never throws. An app has no deploy of its own to compare, so every change is one you record. Needs forgeopstracker 0.9.0, on a plan that includes change tracking.

An issue's verdict listing what changed just before it started, labeled potentially relevant: the gateway_retry_v2 flag turned on for 100% of checkouts 3 minutes before, and a stripe upgrade and a new environment variable detected at startup after the deploy 7 minutes before
import io.github.lukepopwell.tracker.android.ForgeOpsTracker

// Your flag client's change listener, with the key and the old and new values
flagClient.addFlagChangeListener { key: String, oldValue: Boolean, newValue: Boolean ->
    ForgeOpsTracker.recordChange(
        "feature_flag",
        "$key turned ${if (newValue) "on" else "off"}",
        details = mapOf("key" to key, "from" to oldValue, "to" to newValue),
        actor = "flag-service",
    )
}

Custom metrics and infrastructure monitoring

Track a business event you name yourself (a signup, a payment) or a reading from one of your own hosts, with one explicit call each; nothing is automatic. The value defaults to 1, so a bare call is a counter; pass a real amount for anything else (it can be negative, for a refund). A captured value is stored as its own row, so counts and sums you compute later are exact. Calls are buffered and flushed as one batch in the background, so they are safe to make inside a request, and flushMetrics sends whatever is buffered right now. The hostname of a reading defaults to one derived from the device. Requires a plan that includes custom metrics and infrastructure monitoring.

A dashboard with custom-metric widgets (orders placed, checkout conversion, orders over time) and infrastructure widgets (average readings by host and by metric)
ForgeOpsTracker.captureMetric("signup")                 // value defaults to 1.0: a bare counter
ForgeOpsTracker.captureMetric("payment", 49.0)          // a real magnitude; it may be negative (a refund)

ForgeOpsTracker.captureInfrastructureMetric("battery", 0.42)                    // hostname defaults to serverName
ForgeOpsTracker.captureInfrastructureMetric("disk", 0.81, hostname = "db-1")
ForgeOpsTracker.flushMetrics()                          // send now, on the upload thread

Database errors

When an error comes from a database call, the issue shows which stored procedure, table or view its SQL touched, so you know where to start looking. SQLite's own error text, including Room's, ends with the statement, and this client reads it from there, so a local database error needs nothing added. The names are sent by default and are identifiers, never values. The SQL statement itself is opt-in, with every string and number replaced by a question mark before it leaves your app, and a project setting can stop ForgeOps storing the statement at all.

An issue occurrence's Database section, naming the stored function and the view the failing query touched, with the statement's values replaced by question marks
ForgeOpsTracker.init(context) { config ->
    config.captureSqlStatement = true // default false
    // config.captureSqlObjects = false // default true; false stops even the names
}

Questions

Do I need to change how I already handle errors?

Almost never. Every client hooks the framework's own exception handling directly, the same way the install snippet above shows, so your own logs, error pages, and rescue blocks keep working exactly as they did before, and ForgeOps just also finds out about it. The one common exception is a background-job runner that doesn't forward through that same mechanism on its own; that needs one small hook added instead, not a change to anything already there.

Where does the data actually go?

Straight to ForgeOps' own ingestion API over plain HTTPS. No agent, no sidecar, and no third-party ingestion service in between, so nothing about your app's traffic or its users' data gets routed through anyone else's pipeline first.

What happens if the same bug fires a thousand times?

It's fingerprinted from its exception class, message, and backtrace, so a thousand reports of the same bug still show up as one issue with an event count of 1,000, not a thousand separate rows to dig through. And a runaway loop can't use up your monthly event quota: once one issue repeats faster than your plan's hourly limit (50 an hour on Free), further repeats are still counted, and still count toward spike alerts, but they aren't stored or charged to your quota.

Is anything scrubbed before it's stored?

Yes. Emails, credit card numbers, and known API key/token formats are redacted out of every event before it's even written, on every plan, with room for per-project custom field names on top of the defaults.

Try it with your own Android app

Free plan included, no credit card required. New organizations start with 14 days of Business.

Get started free