Screening 20 Infrastructure Primitives Money First: Why Distill and Precomputing Survive
Run the 20 primitive project ideas through the four-question screen strictly and the ranking changes a lot. Several technically attractive ideas turn into weak projects. The economic buyer is unclear, free infrastructure already solves the problem well enough, or the moat amounts to “ours is smaller.”
So Statefile, Lease, BareQueue, TTL, Count, Remember, Freeze and probably ConfigLog go out at this stage. They all solve real engineering problems, but the money question is poor. Redis, SQLite, existing Go libraries, message queues, object storage, Git and cloud infrastructure already give acceptable answers. A tiny implementation might make a respectable GitHub project. “Simpler and smaller” doesn’t tell you who eventually writes a meaningful check.
Gate and Pulse last a bit longer but still land below the build line. Rate limiting and health monitoring matter, but they’re features already built into proxies, API gateways, observability platforms and cloud infrastructure. They’re exactly the kind of thing a platform kills by adding a checkbox.
That leaves a more interesting group: Distill, Provenance, Once, Precomputing, Invariant, TemporalKV, Delta and TinyGraph. Narrow it again and five deserve serious investigation: Distill, Precomputing, Provenance, Once and Invariant. They don’t deserve equal weight.
Distill: the strongest money answer
Distill is the number-one candidate. For a primitive, the money answer is unusually plausible. Infrastructure and observability teams already spend real money collecting, moving, indexing and retaining telemetry. The eventual buyer is a VP Engineering, head of infrastructure, SRE or observability lead, or a FinOps lead whose organization pays large observability or cloud-storage bills.
The problem is measurable. Huge volumes of repetitive logs, metrics, network observations and machine events get retained even though only a fraction carries durable information. The workarounds are just as visible: sampling, retention tiers, filtering rules, aggregation, compression, and deleting old data. Distill attacks the bill before the expensive storage and indexing stages, instead of being one more place to store the data.
The plain-language pitch is excellent: Distill watches a continuous stream and keeps what’s informative (changes, anomalies, extremes, distributions, representative examples) while throwing away repetition.
The weakness is moat. A generic reducer can be copied. So the moat has to come from the distillation format and algorithms: deterministic reduction rules, application-specific reducers, compatibility with existing telemetry pipelines, accumulated knowledge of what’s safe to discard, and eventually a standard representation downstream tools understand. The commercial product may not be Distill itself. It could be an observability cost-reduction product built around it. That’s the Linux and Red Hat split: the primitive stays open, while deployment, fleet management, policies, integrations, auditability and enterprise operation become the business.
Precomputing: the deepest technology bet
Precomputing comes second. It probably has more technological depth than Distill, with a harder first sales story. The money sits in expensive computation and queries repeated far more often than their answers change. Buyers are engineering organizations running high-query-volume applications, analytics systems, APIs or AI and data pipelines. Today they cope with caches, materialized views, scheduled jobs, Redis, hand-rolled precomputation and ever more tangled invalidation logic. Real salaries and real infrastructure spend are attached to that.
The explanation is simple: tell Precomputing which answers matter, and it keeps them updated as the underlying data changes, so asking for them later is almost free.
The moat could end up stronger than Distill’s, if the engine gets genuinely good at dependency tracking and incremental recomputation. Storing a cached answer is easy. Working out exactly what must be recomputed when something changes is the hard part. With its own compact dependency representation, incremental algorithms and eventually query-aware optimization, the IP gets much harder to reproduce than another cache. This one stays very much alive.
Provenance, Once and Invariant
Provenance ranks third, largely because AI has made its problem far more economically relevant. Information now passes through models, retrieval systems, transformation pipelines and agents, which raises a basic question: where did this answer actually come from? Buyers include AI platform teams, data governance groups, regulated enterprises, intelligence organizations and, eventually, compliance departments. Today’s workaround is metadata scattered across databases, logs, vector stores and application code. That’s precisely the ugly workaround worth looking for.
The pitch is clean. Every piece of information gets a record of where it came from and what happened to it, and when something produces new information, Provenance links the result back to its sources. Powerful for AI, and especially interesting for OSINT. But hashing objects and recording parent-child links is trivial to copy. The moat has to come from an interoperable provenance format, integrations, evidence-chain semantics and cryptographic verification. It’s a standards play as much as a software play. If other tools start producing or consuming the format, the format becomes the moat.
Once ranks fourth, and it passes the money test better than its size suggests. Duplicate execution does real financial and operational damage in payments, provisioning, APIs, job systems and distributed workflows. Developers patch it with idempotency keys, database constraints, Redis, queues and custom logic. The buyer starts in engineering: a CTO, a platform engineering lead, an infrastructure lead. The demo writes itself. A network call times out after the remote operation succeeded, the client retries, and the operation happens twice.
Its pitch may be the best of the lot: give an operation a name, and Once guarantees retries won’t perform it twice. The problem is moat again. The mechanism isn’t hard enough to protect a company. Once makes plenty of sense as a primitive and an open source project, much less as the final product. The commercial descendant would have to become reliable execution infrastructure (distributed Once, durable functions, execution history, recovery, cross-service guarantees), and that market is valuable and crowded. Build Once to earn a reputation for elegant infrastructure primitives. Don’t make it the main company bet.
Invariant is fifth, and the speculative pick. Its money case is the weakest of the five, because customers spend heavily on monitoring without seeing “invariants” as a budget line. Yet the problem is real. Monitoring systems pump out measurements and alerts when operators really want one answer: is the system still in an acceptable state? Teams express that today through dashboards, alert rules, queries and runbooks.
The pitch is strong: tell Invariant what must always be true about your system, and it tells you when one of those truths stops being true. That’s a product sentence. The primitive can stay tiny: assertions, changing values, state transitions. Prometheus tells you latency is 312 ms. Invariant tells you the system has entered a state you declared unacceptable. The long-term moat would have to come from accumulated operational knowledge, reusable invariant packs, integrations and dependencies between invariants. Riskier commercially, but different enough not to kill.
TemporalKV, Delta and TinyGraph go in a keep-but-don’t-build drawer. TemporalKV solves a real problem, but temporal and versioned databases make differentiation hard. Delta is foundational technology with a poor standalone money answer: easy to explain, hard to name the person who buys Delta rather than a product with Delta inside. TinyGraph has the same problem in another form. Relationships matter enormously, especially for OSINT, but graph databases exist and “a smaller graph database” fails the differentiation test. Both are likely worth more as internal parts of something else than as projects promoted on their own.
Two halves of the same bill
The strongest result of the screen isn’t the ranking. Distill and Precomputing belong together conceptually while solving opposite halves of the same economic problem. Distill asks what, of everything arriving, is worth retaining. Precomputing asks what, of everything users might ask later, is worth answering now. Conventional infrastructure keeps the inputs and calculates answers afterward. These two challenge both habits: don’t keep redundant input, and don’t keep recalculating important answers.
Together they pass the four tests better than almost anything else in the original list. The money comes from cutting storage, observability, database and compute spend. The problem is excessive machine-generated data plus repeated computation. The technology explains itself without jargon: keep the useful information instead of all the raw data, and keep important answers ready instead of recalculating them. And there’s a plausible route to a moat through reduction algorithms, incremental computation, dependency tracking, integrations and, eventually, a shared data representation.
So the build decisions today look like this. Distill is the strongest new primitive to prototype. Precomputing stays the strongest larger technology bet. Provenance earns a prototype because AI is building a commercial context around it fast. Once is a fine compact open source and reputation project. Invariant needs more design work before code. Everything else stays in the notebook until one of the five fails its prototype or customer test, or gives a reason to bring another idea back.
One more step before writing Distill code, under the money-first rule. Find five real observability bills, pricing models or engineering accounts showing what organizations pay to ingest and retain data. Define exactly where Distill sits in that pipeline. Then do a brutally concrete calculation: 1 TB a day goes in, 180 GB a day comes out, and at this vendor’s ingestion and retention prices that saves $X a year.
If that number looks good, it’s no longer just an interesting primitive. It’s the start of a business case.