Recent Posts
A Background Job Queue in One SQLite File Covers What Most Apps Deploy Redis For
A signup handler has to send a welcome email, and sending it inline puts a third-party API call inside your request. So you move it into a background job, and the usual next step is Redis, a client library, a worker process, a dashboard, and a note in the deploy docs about what happens to queued jobs when Redis restarts without persistence. If the app runs on one machine, you’ve added a second stateful service to look after a table’s worth of data.
A Database Where Every Value Remembers Its Source, Confidence and Extraction Time
Somebody pastes a population figure into a shared spreadsheet: 1,234,567. It might have come from a census table, a Wikipedia infobox, a city press release or a model’s summary of one of those, and three weeks later nobody can say which. The paste kept the digits. Everything that made them believable stayed in a browser tab that closed long ago.
A value without provenance is a rumor with a data type.
A Feature Store Small Enough to Embed: Point-in-Time Reads for Local and Edge ML
A churn model looks excellent in the notebook and mediocre in production. The model isn’t broken. The training table computed orders_7d with a rolling window over the full history, so some rows counted orders placed after the label date. The serving code, written later in another language, counts a calendar week instead of seven days. One feature, two definitions, and a quiet look into the future. Those are the two bugs a feature store exists to prevent: training-serving skew and label leakage.
A Schema Diff That Warns Which Migration Will Lock a 40-Million-Row Table or Lose Data
The diff in the pull request is one line: ALTER TABLE orders ALTER COLUMN total TYPE bigint;. It reads like a rounding fix, so the reviewer approves it, and the SQL is correct. On a table with 41 million rows it’s also a full rewrite under a lock that stops every read and write until it finishes. A diff shows what changes, never what the change costs.
A schema diff worth using would print consequences instead of SQL, in sentences like “this rewrites a 41-million-row table and blocks every read and write” or “the rollback is impossible without a backup”. Those sentences depend on three things: the statement, the engine and its version, and the size of the table it lands on. With all three you can post a risk report on every migration pull request. With only the first you have a linter.
A Single-File Log Database: Pipe Logs In, Query Them With SQL, Hand the File to Anyone
The incident is over, the postmortem is Thursday, and someone wants the logs from the bad three hours. You can attach a gzip nobody will open, paste a screenshot of a grep, or stand up a search cluster to answer one question. The question is usually “which paths returned 500, and when did it start”, which has a GROUP BY hiding in it, and a GROUP BY is where grep runs out. A full observability stack answers it, but that’s a heavy answer for a team that doesn’t already run one.
A Telemetry Database That Forgets on Purpose Keeps Rollups and Anomalies Instead of Raw Rows
At 300 requests a second, one latency metric adds about 26 million rows a day. Nobody will read most of them again. What people ask for a month later is the p99 per route for last Tuesday, and the one request that took nine seconds. A rule that deletes rows after 30 days throws away both answers. Keeping every row forever pays storage bills for data nobody reads.
The telemetry store worth building decides at write time what has to survive. It folds everything else into coarser buckets on a schedule and keeps the few raw events that carry information, whole. Forgetting becomes part of the schema, declared in a policy, instead of a cron job that deletes by age. This is the argument for reducing data near the source, applied to time series.
A Tiny SQL Engine for JSON Streams Fits Between jq and a Stream Processing Cluster
Your service writes one JSON object per line to stdout, and you want the average cpu for host web-3 over the last five minutes. jq can’t answer that, because it forgets each object once it has printed it and has no idea what “five minutes” means. You can fake both with jq -n 'reduce inputs as $e (...)' plus a timestamp comparison, and you’ll have a script nobody wants to touch twice. DuckDB answers the question in one line, but only against a file as it exists right now. ksqlDB, Materialize, RisingWave, Arroyo and Flink SQL answer it continuously, once you’ve deployed a service and agreed to operate it.
An Embedded Sketch Database: Billions of Events, Megabytes of Storage, Error Bars on Every Answer
Counting unique visitors exactly means remembering every visitor. Say 6 million distinct IPv4 addresses hit a checkout service today. Packed at 4 bytes each, that set is 24 MB before any hash table overhead, and you need one per service, per region, per day you want to compare. Redis answers the same question with a HyperLogLog in about 12 KB per key, at a standard error of about 0.81%.
The thin part is the database around it: a SQLite-shaped file of sketches with time buckets, rollups and a SQL front door (check before you start; this space moves). The version worth building stores nothing but mergeable sketches, grouped by dimension and time bucket, and every aggregate returns its error bound next to the value. You never keep the events. It’s the extreme case of reducing data near the source: keep the answer, drop the rows.
An Embedded Workflow Engine With Five Primitives: Trigger, Condition, Action, Wait and Retry
A user signs up on Monday. If they haven’t activated by Wednesday, you send a nudge email, and if the email provider times out, you try a few more times before flagging the account. That’s the whole spec. Most teams build it as a cron job that scans for users in the wrong state, a nullable nudge_sent_at column and a retry loop written late on a Friday. It holds until the cron overlaps itself, or a deploy lands mid-loop, or the provider times out after actually sending the message and one customer gets three emails.
An Evidence Graph Built From Extracted Claims Keeps Contradictions Attached to Their Sources
Feed forty articles about a rumored plant restart to a summarizer and you get one smooth paragraph: several outlets report that production will resume in March, though the ministry has pushed back. It reads well. It also throws away what you’d need to act on it. You can’t tell which outlets reported it, whether each had its own source or all copied one wire story, or what the ministry actually said, a denial or a polite refusal to comment.