Group Errors by Behavioral Signature Instead of Stack Trace Text
Add six lines to the top of a file and ship it. Every stack trace through that file now carries different line numbers, so a tracker that hashes the trace files a bug it has seen a hundred times as a brand-new issue. Meanwhile an older issue titled TimeoutError has collected comments from people who each fixed a different timeout and wondered why it never closed.
Grouping fails in both directions at once. One bug turns into hundreds of issues because the message embeds an order ID, the line numbers moved, or a framework upgrade changed the frames in the middle. Two bugs turn into one because every timeout in the codebase raises the same exception type from the same retry helper.
The version worth building groups by behavior: what failed, in which operation, against which dependency, with the variable parts of the message masked out. It keeps two levels of buckets, so you can read forty precise ones or five coarse ones depending on how bad the day is. It sits between raw logs and a full observability stack: more structure than grep, far less machinery than a platform.
The Idea Is Older Than Your Tracker
Grouping crashes by signature was proven at scale long ago. Windows Error Reporting bucketed crash reports from huge numbers of machines, and Microsoft wrote up ten years of running it in “Debugging in the (Very) Large”, published at SOSP in 2009. Mozilla’s crash-stats service, Socorro, files crashes under a signature computed from the crashing thread’s stack. Sentry groups by stack trace by default and lets you override that with custom fingerprints and stack trace grouping rules.
Those overrides are the right escape hatches, but they’re still escape hatches. You write a rule after you’ve watched a bug split in half, one bug family at a time. What’s missing is a default signature that already includes the operation and the dependency, plus a coarser level above it so the dashboard has something short to show.
What Goes Into the Signature
A signature is six normalized fields, hashed: exception type, top in-app frames, operation, dependency, masked message and an optional input-shape class. The normalization is where the work is.
Frames lose their line numbers and build-specific paths. Consecutive repeats collapse into one. Anything under site-packages, node_modules or a vendor directory folds into a single placeholder, so a framework upgrade that adds two frames changes nothing. The operation is the route template (POST /checkout/{id}/pay, never the URL), the job name or the CLI command. The dependency is the outermost call that left the process: the folded library frames say whether it was HTTP or a database driver, and the host or database name usually comes from the message, or from tracing data when you have it. The message gets the masking that log template miners such as Drain apply to log lines (the post on reducing logs before shipping them covers the mechanism), so numbers, hex, UUIDs, IPs, emails and long quoted strings become placeholders. The input-shape class records which field was null or which argument was out of range, when the message names it.
Hash the canonical string with a real hash such as SHA-1 or BLAKE2, never the language’s built-in hash(), which Python salts per process. Store the canonical string beside the hash, because the first thing anyone asks about a bucket is why these events are in it.
The second level is the super-bucket: exception type, the frame where the failure originated and the dependency, with the operation, caller frames, message and input shape dropped. A bad cache entry that breaks three routes makes three precise buckets and one super-bucket. So does a flaky payment host.
raw events
ReadTimeout: HTTPSConnectionPool(host='api.payments.example', port=443): Read timed out. (read timeout=10)
gateway.py:88 charge_card < checkout.py:41 submit < [flask x6] POST /checkout/8812/pay
ReadTimeout: HTTPSConnectionPool(host='api.payments.example', port=443): Read timed out. (read timeout=15)
gateway.py:91 charge_card < checkout.py:44 submit < [flask x6] POST /checkout/9047/pay
ReadTimeout: HTTPSConnectionPool(host='api.payments.example', port=443): Read timed out. (read timeout=10)
gateway.py:91 charge_card < checkout.py:44 submit < [flask x8] POST /checkout/9113/pay
bucket 9f3c1a07 (3 events)
type ReadTimeout
frames charge_card < submit
operation POST /checkout/{id}/pay
dependency http:api.payments.example
message HTTPSConnectionPool(host=<str>, port=<num>): Read timed out. (read timeout=<num>)
super-bucket 4be20d11 = type + origin frame + dependency
Four things differ across those events (line numbers, timeout value, order ID, framework depth) and none of them reaches the signature. Point the same retry helper at Postgres instead and the type stays the same, but the dependency field changes, so the bucket splits where it should.
Test It Against Your Own Fix Commits
A signature is only good if it matches how you fix bugs, so replay a month of your own errors through it and check the buckets against the work you did. One fix should close one bucket. Two numbers fall out: buckets per fix (above 1 means you’re splitting a bug) and fixes per bucket (above 1 means you’re lumping two).
The labels are mostly sitting in your tracker already. Every time someone merged two issues by hand, they told you two buckets were one bug. Every ticket reopened with a different cause told you one bucket held two. Join those to the commits that reference the tickets and you have a labeled set nobody had to annotate. Tune the masks and frame rules against that set and ignore how the dashboard looks.
Where Signatures Go Wrong
Minified and obfuscated stacks break everything downstream. A release-build JavaScript stack is a list of one-letter function names, and an Android release stack looks the same until the ProGuard or R8 mapping is applied. Sign before symbolication and every release gets its own buckets, because the names change per build. So symbolicate first, and when the source map or mapping file is missing, fall back to a coarse signature (type, operation, dependency) and mark the bucket as degraded instead of pretending it’s precise.
Async stacks are the next problem. In an async runtime the stack at the throw site is often just the event loop, and the real caller is gone unless the runtime keeps async frames. There the operation field carries the weight, which is one reason it’s in the signature at all. Recursion is easier: one request fails 50 frames deep, another 5,000, so collapse consecutive repeats before taking the top few.
Renames are the slow poison. Refactor charge_card into charge and every bucket through it gets a new hash on deploy day, with its history left behind. You can follow renames through git (file renames reliably, function renames only partly), or tolerate the drift: when one bucket stops at a release and a new one starts at that same release with the same operation, dependency and message template, link the second as the first’s successor and show one history.
Masking is tuning with consequences. Mask too little and one bug is hundreds of buckets; mask too much and column "email" does not exist merges with column "phone" does not exist, which may be two missing migrations. Keep the mask list short, per language, and visible in the bucket’s tuple. It has a second job too. Messages carry emails and tokens, and a bucket’s sample events outlive the incident, so mask before you store.
A First Version That Refuses Most Things
Version 0.1 ingests JSON error events from whatever SDK or log shipper you already run: exception type, message, a list of frames, and tags for the route and the dependency. It computes signatures for two languages, Python and Java say, because both give you structured frames without a build step in between, so the first release tests the grouping idea without the symbolication pipeline. It prints buckets with counts and first and last seen, as a CLI table and as a webhook that fires once per new bucket.
Each bucket should keep one sample you can act on. For an API failure that means a failed request packaged into one replayable file. The bucket says what keeps breaking; a timeline of state changes can say what the data looked like beforehand. Both are other tools. And it won’t join five different exceptions caused by one bad deploy, because that’s a job for release markers and time.
It refuses alerting suites, performance monitoring and dashboards. A signature function with honest buckets is small enough to finish.
One fix should close one bucket. Everything else is tuning.