A Telemetry Database That Forgets on Purpose Keeps Rollups and Anomalies Instead of Raw Rows
At 300 requests a second, one latency metric adds about 26 million rows a day. Nobody will read most of them again. What people ask for a month later is the p99 per route for last Tuesday, and the one request that took nine seconds. A rule that deletes rows after 30 days throws away both answers. Keeping every row forever pays storage bills for data nobody reads.
The telemetry store worth building decides at write time what has to survive. It folds everything else into coarser buckets on a schedule and keeps the few raw events that carry information, whole. Forgetting becomes part of the schema, declared in a policy, instead of a cron job that deletes by age. This is the argument for reducing data near the source, applied to time series.
Half of this is old and proven. RRDtool (Tobias Oetiker, 1999) keeps fixed-size round-robin archives and consolidates old samples with AVERAGE, MIN, MAX or LAST. Graphite’s Whisper does it with retention strings like 10s:6h,1m:7d,10m:5y. Prometheus has recording rules for precomputed aggregates, and Thanos adds downsampling to 5-minute and 1-hour resolution. TimescaleDB has continuous aggregates, and Druid can roll data up as it’s ingested. Each answers one question well, which is how coarse old data can get.
The other half comes from industrial process historians, which face the same flood from sensors. They use deadband filtering, which stores a sample only when it moves more than a tolerance from the last stored one, and swinging-door compression, which stores a point only when a straight line from the last stored point can no longer stay within tolerance of the samples in between. OSIsoft’s PI popularized both. These answer a different one: which raw points the signal needs.
None of the parts is new, which is what makes it buildable. Combine both halves with mergeable sketches and you get a database that names what must survive, folds the rest into coarser tiers, and keeps interesting raw events whole. In a field this crowded a newcomer has to be small: one file, no server, a policy a person can read. Precomputing, which we covered here, is one attempt at that for dashboard answers. You write a short policy naming the answers you want and how long each level of detail lives, old detail fades on schedule, and unusual events are kept whole. It’s version 0.1, tested on simulated data.
A Survival Policy Instead of a Retention Rule
The unit of configuration is a policy per metric: its dimensions, how long each resolution lives, which aggregates each bucket carries, and which raw events outlive their tier. Something like this, in invented syntax for a tool that doesn’t exist:
metric: http_latency_ms
dimensions: [service, route, status_class]
max_series: 5000 # new series beyond this fold into "_other"
tiers:
- { resolution: raw, keep: 1h }
- { resolution: 1m, keep: 7d }
- { resolution: 1h, keep: 400d }
- { resolution: 1d, keep: forever }
aggregates: [count, sum, min, max, first, last, variance]
quantiles: tdigest
distinct: { column: client_id, sketch: hll }
top: { column: route, k: 20 }
distill:
deadband: 2%
anomaly: { above: p99.9, window: 1h, budget: 50, context: 3 }
sample: { per: 1h, size: 20 }
keep: 90d
late: { grace: 10m, then: merge_into_coarser }
Aggregates update as rows arrive, in the finest bucket, the way Druid rolls up at ingestion. Coarser tiers are merges of finer ones once those close. The distiller watches the raw tier as it fills. It keeps a raw point when the value moves more than the deadband from the last kept one, or when the point lands above the previous hour’s p99.9 (read from the sketch the engine already holds). A small reservoir sample keeps some ordinary points too, so you can see what normal traffic looked like. Each anomaly keeps the three points on either side for context. Everything unmarked is dropped when its tier expires, and the kept events stay whole for 90 days.
The size comes from the tiers, the number of series and the sketch sizes. One series sampled every second is 86,400 points a day raw, 1,440 at one-minute buckets, 24 at one-hour buckets, and one a day after that. The plain aggregates cost on the order of a hundred bytes per bucket. For regular metrics that’s a reduction of orders of magnitude per series, but the total depends on cardinality, and a quantile sketch in a one-minute bucket can cost as much as the sixty samples it replaces. Put sketches where they pay: coarser tiers, or the few metrics where tail percentiles matter.
Only Mergeable Aggregates Can Roll Up
Rolling a minute into an hour means combining summaries without the raw rows, and not every summary can be combined. The median of 60 one-minute medians is not the hourly median. An average of averages is wrong unless every bucket had the same count, which is why the policy stores sum and count and divides at read time. Variance merges if you keep the count, the mean and the sum of squared deviations from Welford’s algorithm, because two such triples combine with a closed-form update. First and last values merge only if they carry their timestamps.
Percentiles need a sketch (t-digest or KLL) or a histogram with fixed bucket boundaries, which merges by adding counts but is only as accurate as its boundaries. Distinct counts need HyperLogLog, which merges by taking the maximum of each register. Top-K needs a Space-Saving summary, which merges with a modest loss of accuracy. The sketch database post covers those error bounds.
Deciding What Matters Before You Know
The distiller is the risky part. It has to decide what’s interesting before anyone knows what the incident will be. Thresholds taken from a signal’s own recent history adapt better than fixed limits, but they misbehave on a regime change: after a deploy that makes every request slow, everything looks anomalous and the kept set floods. Hence the budget in the policy, 50 an hour, with a counter of how many were dropped so you can tell the budget was hit. Anomalies that only show across series (three services slowing together) are invisible to a per-series rule, and a first version should say so plainly.
Loss is irreversible, which makes policy mistakes permanent. Two mitigations help. Keep raw rows locally for a short window, so a new question can still be answered exactly for the recent past; Precomputing’s Logs part does this for log lines, keeping raw lines locally for 48 hours and sending reduced answers upstream. And let the tool test a policy before you trust it: replay the raw window through the proposed tiers and compare each rollup answer to the exact one. Logs fit this shape too, with a template and a count standing in for a metric and a value, which is what reducing logs before shipping them is about.
Two more things bite in production. The first is late data: a device that was offline uploads yesterday’s readings after yesterday’s minute buckets were folded into the hour. Give each tier a grace period, then merge a late point straight into the coarser bucket (which is why every aggregate has to be mergeable) and mark the finer tier incomplete. The second is cardinality. Every distinct label combination is a series, and each tier multiplies the storage, so the policy needs a series cap and a rule for the overflow, such as folding it into one _other series and counting what was folded.
What v0.1 Does
A first version is one SQLite file. A series table maps label sets to ids, one table per tier holds buckets, and a kept table holds the raw events the distiller saved. It supports count, sum, min, max, first, last and variance, plus one quantile sketch, and it takes lines from a pipe or a library call. Compaction runs as small transactions, one coarse bucket at a time: write the merged bucket and delete its fine buckets together, so a crash can’t leave a gap. Distinct counts and top-K wait until the sketch layer exists. Alerting is out, and so is any aggregate you decide you want after the fact.
If one tier is all you need, the hourly rollup tables in the cookieless analytics idea from Nine Apps Worth Coding are the smallest version of this.
Forgetting will happen anyway. Choose what goes.