Most Log Volume Is a Few Templates Repeated, So Reduce Logs Locally Before Shipping Them
Take a day of production logs from one service, mask the numbers, IDs and quoted strings in every line, and sort the resulting patterns by how often they occur. The top of that list is short and dull: health checks, cache hits, a retry notice from a client library nobody owns. Each of those lines was serialized, shipped, parsed and indexed millions of times to say the same thing, and you paid for every copy. Run it on your own logs. The claim is easy to check, so don’t take it on faith.
Most backends bill on how much you ingest, how much you index, or both, but the information in a stream tracks the number of distinct templates much more closely than its size. So reduce at the source. A small local process fingerprints each line’s template, forwards a count and a few representative samples for the repetitive ones, and forwards everything whole for the lines that matter. The decision is deterministic, with no model on the hot path, so it’s cheap enough to run next to every service, and its output is a number a finance person understands: bytes in, bytes out.
The template mining is well studied. Drain, a 2017 algorithm, parses logs online with a fixed-depth tree; IBM’s open-source Drain3 packages it for streaming, and the Loghub collections give you public datasets to test against. The commercial idea is proven too. Cribl built a large business routing and reducing telemetry before it reaches expensive backends, Grafana Cloud’s Adaptive Logs finds low-value patterns so they can be dropped or sampled, Edge Delta processes telemetry at the edge, and Datadog’s Logging without Limits separates ingesting a line from indexing it. Shippers such as Vector and Fluent Bit can drop and sample by rules you write by hand, which you then maintain forever.
So the space is crowded, and a newcomer won’t win on the concept. The opening is the small, per-host, deterministic version that runs in front of whatever backend you already pay for and prints its savings before anyone commits.
What the Reducer Does With Each Line
Tokenize the line, mask the variable tokens (numbers, IP addresses, UUIDs, hex strings, quoted strings), and walk a Drain-style tree: first by token count, then by the leading tokens, down to a leaf holding a short list of templates. If the masked line is similar enough to one of them (a threshold on the share of matching tokens), it joins that template and any differing token becomes a wildcard. Otherwise it starts a new template. Cost per line depends on the tree’s depth and the few templates at a leaf, so it stays flat as the template count grows, which is why this works as a streaming step.
For each template and each interval, keep a count, first and last timestamps, bytes in, and a few samples: the first line, the last, several picked by reservoir sampling, and any line with a rare variable value. At the end of the interval, emit one summary record where there were thousands of lines.
{"type":"summary","template":731,"level":"ERROR",
"pattern":"payment gateway timeout after <num>ms host=<str>",
"window":"2026-10-05T08:00:00Z/60s","count":43282,"bytes_in":5918204,
"samples":["payment gateway timeout after 10012ms host=pg-3","..."],
"kept_values":{"tenant_id":{"t_204":41970,"t_17":1100,"other":212}}}
Some lines are never summarized. New templates go out whole until they’ve been seen a set number of times, rare templates go out whole, errors go out whole while their rate stays under a threshold, and anything matching a keep rule goes out untouched, audit and security logs first. The rules fit in a short file.
window: 60s
samples: {first: 1, last: 1, reservoir: 5, rare_value: 3}
raw_retention: 24h # raw lines stay on local disk
forward_whole:
- match: {source: [audit, auth]} # checked before the reducer sees the line
- match: {level: [ERROR, FATAL]}
while_rate_below: 5/min
- new_template: {first_n: 100}
keep_fields:
- template: "payment gateway timeout*"
fields: [tenant_id]
Two more pieces make the output trustworthy. Raw lines stay on local disk for a window, so an engineer in the middle of an incident can pull the real lines behind a summary. And every summary carries bytes in and bytes out per template, so the savings are a report instead of a claim. The Logs part of Precomputing follows the same outline: it learns templates with the Drain algorithm, keeps raw lines locally for 48 hours and sends reduced answers upstream.
A model has a place here, but off the hot path. When a new template appears, ask one question once (what does this line mean, and what should it be called?) and cache the answer by template ID. The model never decides what gets dropped. Deterministic code does that, so the cost stays flat as volume grows.
Where Reduction Goes Wrong
Template drift is the first problem. A deploy rewrites payment timeout as payment gateway timeout, and the reducer sees a template with no history, which the rules forward whole. Multiply that by every reworded message in a release and the day after a deploy becomes your most expensive day. Cap whole-forwarding at a fixed number of lines per new template, and merge a new template into an old one when the old one stops at the release where the new one starts and the tokens mostly overlap (add the logger name or call site to the match when the line carries one). The same successor trick helps when grouping errors by behavioral signature, which relies on the same masking step.
Variables that are the signal come next. Masking assumes the numbers are noise. In a fraud log the user ID is the whole point, and in a noisy-neighbor hunt the tenant ID is. So templates can declare keep-fields, and the reducer holds bounded counts of the values it saw per interval (the Space-Saving algorithm is the usual choice for top-k counting in fixed memory), which is what the kept_values field above carries. Lines with rare values go out as samples, so one odd tenant doesn’t get averaged away.
Compliance is simpler to state than to guarantee. Audit and security logs have to arrive whole, and the rule that ensures it runs before the reducer sees the line, so a bug in template matching can’t touch them. The reduction itself needs an audit trail: what was dropped, how many lines, under which rule.
Anything that counts raw lines breaks when lines become summaries. Emit a counter per template and level as a metric (a telemetry database that forgets on purpose can keep those cheaply long after the raw lines are gone), and move every alert that matches on a log line either into the reducer, where it still sees the raw line, or onto that counter. Counting belongs in metrics anyway, per the standard split between logging, metrics and tracing, and the reducer simply forces the move. It’s dull work, and skipping it is how a pager goes quiet in the middle of an outage.
Then the hardest question: did you drop the line that mattered? Make it checkable. Per interval, lines in must equal lines summarized plus lines forwarded whole, and the reducer reports any gap. Replay old incidents through it and confirm the first line an engineer cared about was forwarded whole, appears among a summary’s samples, or sits in the local raw window. Now the safety number sits next to the savings number in the same report.
Who Pays, and What Version 0.1 Skips
Anyone with a logging bill of five or six figures a month. The pitch is a measured percentage of a number they already see on an invoice, which is why reduction sells when many infrastructure tools have to argue their value: run it in shadow mode for a week, read the report, then turn forwarding on. The algorithm is the free part. The keep rules, the audit trail and the proof that nothing important went missing are what people pay for. The same logic runs through the case that small infrastructure tools should reduce data near the source.
Version 0.1 reads stdin or a followed file, runs the tree with masks for the five variable types, writes JSON-lines summaries, honors a keep-rules file, keeps raw lines in rotating local files, and prints the in/out report with the gap check. It refuses to be a shipper (your existing agent moves the bytes), a search UI or a backend. Model-written template names wait for 0.2. The other half of the local story, a queryable file you can hand to someone, is a single-file log database.
Print the bytes saved. That number is the product.