Debug From a Timeline of State Changes Instead of Ten Thousand Log Lines
A ticket comes in: order 8812 shows as cancelled, but the customer’s card was charged. You open log search. The checkout service wrote a couple of hundred lines for that request; the payment webhook handler wrote a few dozen; the job that expires unpaid orders wrote a line for every order it looked at that minute. The answer is in there somewhere. Forty minutes later you find it. The expiry job cancelled the order thirty seconds after checkout began, and the payment provider’s webhook landed half a second after that.
Three lines would have told you the same thing:
$ timeline order:8812 --since 14:00
14:02:11.204 status pending -> awaiting_payment req 7f3a POST /orders/8812/pay
14:02:41.380 status awaiting_payment -> cancelled job e19c expire_unpaid_orders
14:02:41.902 payment_state authorized -> captured req b220 POST /webhooks/payments
Logs record what code ran. Most debugging questions are about what changed, and who changed it. A tool that captures state transitions as their own record type and draws them as a timeline per entity answers those questions without the archaeology. The version worth building starts in the database, because that’s where most backend state already lives, and capturing it there costs the application almost nothing.
The Missing Record Type
Frontend developers have debugged this way for years. Redux DevTools lists every dispatched action next to the state it produced and lets you step backward through them, which is the time travel everybody demos. The Stately inspector does the same for XState machines: you watch the machine move between named states and see which event moved it. The standard case for Redux in React Native state management is that its explicit action flow makes state changes auditable, and the DevTools are where that case pays out.
Backends mostly lack an equivalent. They have logs, which are prose written by whoever wrote the function, and traces, which say which calls happened and how long each one took. A span for POST /orders/8812/pay tells you the request took 180 ms and hit the database twice. It doesn’t tell you the order went from pending to awaiting_payment. The usual signals (logs, metrics, traces) all describe execution, and none of them is a record of state.
The raw material is lying around. Debezium and the wal2json plugin turn Postgres logical decoding into a stream of row changes. Event-sourced systems store transitions by design, and trigger-based audit tables are decades old. What nobody hands you is the debugging layer on top: a narrow record that links each change to the request behind it, plus a screen that shows one entity’s story.
That record is the primitive hiding in this idea. It fits on one line:
entity, entity_id, field, old, new, cause, trace_id, at
cause is the request or job that made the change, as a name plus an ID; trace_id ties it to whatever tracing you already run. Everything else falls out of the shape. A timeline is a filtered scan sorted by time. Swimlanes for one request are a group-by on cause. “Which jobs have ever moved an order from paid to cancelled?” becomes one query, where today it’s an afternoon of grep. The record is cheap, too: eight short fields, against a log line that restates the arguments in a sentence of English. Keep records in a local append-only file for the CLI, or forward them as OpenTelemetry span events if you want them next to your traces. Both destinations take the same eight fields.
Three Capture Paths
Database state comes first, through change data capture. Set wal_level = logical, create a replication slot, and consume it with Debezium or wal2json. Every committed change to a watched table becomes a transition, and the application code doesn’t change. The catch is linking each change to its request, since logical decoding doesn’t carry session state. Postgres accepts custom settings with a dotted name, so the app stamps each transaction:
BEGIN;
SELECT set_config('app.request_id', 'req_7f3a', true); -- true = SET LOCAL
UPDATE orders SET status = 'awaiting_payment' WHERE id = 8812;
COMMIT;
set_config with true is SET LOCAL in a form that takes bind parameters. It’s also safe behind a transaction-mode connection pooler, because the value dies at commit; a plain SET would leak into whichever client gets that connection next. A small trigger on each watched table copies current_setting('app.request_id', true) into a context table keyed by txid_current(), once per transaction. That insert rides the change stream like any other row, so the consumer can stamp every change in the transaction with the request ID. (Postgres can also write a transactional message straight into the WAL with pg_logical_emit_message, if your decoding plugin passes messages through.)
In-process state is the second path. A circuit breaker or an in-memory rate limiter never touches a table. Here you declare the fields and wrap them: a JavaScript Proxy with a set trap sees every assignment to an object, and a Python data descriptor does the same for a class attribute. The wrapper compares old and new and emits only real changes, which is what stops it from sliding back into logging.
Explicit state machines give the best data of the three, because every transition already has a name. If the code models an order as a machine with a transition table, emitting a record from the transition function is a single line, and the record carries the event (payment_captured) alongside the before and after.
Deciding What Counts
Capture is the easy half. The hard half is taste, and the tool has to make you write it down. Watch every column of every table and you’ve rebuilt noisy logging in a new file format, so it should refuse to start without a watch list:
watch:
orders:
key: id
fields: [status, payment_state, shipping_address_id]
users:
key: id
fields: [plan, status, email]
redact: { email: hash }
fold:
orders.view_count: 1m
High-frequency fields need folding. A view_count that changes four hundred times a minute becomes one record per minute (“0 -> 412, 412 changes”) or it drowns the timeline. Personal data needs a per-field policy, because old and new values are exactly what you’d rather not copy into a debugging store. A hashed email still shows that the address changed, and that it changed back, without storing it.
Old values cost more than people expect. With Postgres defaults, a decoded UPDATE gives you the new row and, at most, the old primary key. REPLICA IDENTITY FULL writes the whole old row into the WAL, which is fine for a small table and expensive for a hot one. The cheaper design keeps a shadow copy of just the watched fields inside the consumer, seeded from an initial snapshot, and computes old values itself.
Ordering is solid inside one database, where commit order in the WAL is the truth. Across services it falls apart. Clocks drift by milliseconds, and races live in exactly that gap: the job and the webhook in the ticket above were half a second apart, which is easy, while two services writing within a millisecond of each other is not. Order by trace context where calls are causally linked, and draw everything else as concurrent instead of guessing.
Overhead is mostly an operations problem. Reading the WAL costs the application nothing, but a replication slot holds WAL until its consumer confirms it, so a dead consumer can fill the primary’s disk. Set max_slot_wal_keep_size and alert on slot lag before this goes anywhere near production.
What v0.1 Refuses
The first version does one capture path well: Postgres CDC with request-ID linking through set_config, a declared watch list with redaction and folding, a local append-only file, and a CLI that prints a timeline per entity and swimlanes per request. A small web view comes second. It refuses generic log ingestion and it refuses metrics. Both have good homes already, and either would bury the one thing this tool says.
In-process observers for one language and OpenTelemetry export come next. Storage that keeps history as the primary object is a separate design problem, which is the subject of an embedded database built around changes. Errors are the natural neighbor: a timeline helps most when you open it from an error bucket grouped by behavior, and the same idea of one file you can attach to a ticket drives the flight recorder for AI agents.
Order 8812 took forty minutes in the logs. In a timeline, it’s three lines.