A Feature Store Small Enough to Embed: Point-in-Time Reads for Local and Edge ML
A churn model looks excellent in the notebook and mediocre in production. The model isn’t broken. The training table computed orders_7d with a rolling window over the full history, so some rows counted orders placed after the label date. The serving code, written later in another language, counts a calendar week instead of seven days. One feature, two definitions, and a quiet look into the future. Those are the two bugs a feature store exists to prevent: training-serving skew and label leakage.
The products that solve this are built for platforms, and most of the platform is dead weight when the model runs on a phone, a gateway or a laptop. The idea underneath is small. Keyed values with timestamps, a rule for when a value expires, one read for serving that returns the latest, one read for training that returns what was true at a given moment, and one definition of each feature used by both. That fits in a library over one file.
Feast is the open source reference, and it can use SQLite as a local online store, so the serving half already runs on a laptop. Around it sit a registry, an offline store and a materialization step, which is a lot of configuration for a model that fits on a device. Tecton, Hopsworks, Featureform and Chalk are built around streaming pipelines and teams. Plenty of small projects skip all of it and keep a Postgres table of entity, feature, time and value, with a convention. That table is the honest competitor. The version worth building is that table plus the parts people get wrong: the as-of query, the expiry rule and the parity guarantee. The user is someone whose model runs on a device (the on-device ML that works is narrower than the hype, but real) and who has no Redis, no Kafka and sometimes no network.
Two Timestamps and One Index
The core is an append-only table: entity, feature, event_time, ingest_time, value, with an index on (entity, feature, event_time DESC). Event time is when something happened in the world. Ingest time is when the store learned about it. You need both because data arrives late. A device offline for two days uploads its events on Thursday, and a correction is a second row for the same event time with a later ingest time. That makes the table bitemporal, the shape a change-history database has, narrowed to features. Rows are never updated in place.
An online read is one index seek per feature: the newest row whose event time isn’t in the future and falls inside the TTL, with the later ingest time winning ties. It comes back with the value, its age and an expired flag, so the app can log what the model was actually fed. A training read asks a different question: what would serving have returned at time t? Walk the same rows with one more filter, ingest_time <= t. DuckDB’s ASOF JOIN shows the operation, and kdb+ had as-of joins long before. In plain SQL it’s a correlated subquery.
-- times are epoch seconds; TTL is 2 days (172800)
SELECT l.entity, l.ts, l.label,
(SELECT f.value FROM feature_values AS f
WHERE f.entity = l.entity AND f.feature = 'orders_7d'
AND f.event_time <= l.ts
AND f.event_time > l.ts - 172800
AND f.ingest_time <= l.ts
ORDER BY f.event_time DESC, f.ingest_time DESC
LIMIT 1) AS orders_7d
FROM labels AS l;
The TTL line matters as much as the leakage line. Training has to apply the same expiry as serving, or the model learns from values serving would never have shown it. The query is one index seek per label and feature, fine for thousands of labels. A million labels want a sorted merge walk per entity, which is the optimization to add when a profile asks for it.
One Definition, Two Paths
Skew comes from two code paths, so the design removes one. A feature is declared once, with its source events, operation, window, refresh interval and expiry, and the store computes it. Training never recomputes anything in a notebook; it reads rows the same code wrote. That limits transformations to a short list of windowed operations such as count, sum, mean, min, max and last, which suits a design where arbitrary Python can’t run inside a phone app anyway. A C core with bindings means the training machine and the device run literally the same function. A feature is a derived value, so tracking what depends on what applies here too.
The API below is a sketch, not a shipped library.
fs = FeatureStore("features.db")
fs.define("orders_7d", entity="customer", source="orders",
op="count", window="7d", refresh="1h", ttl="2d", on_expired="null")
fs.ingest("orders", entity="customer:42", at="2026-10-04T09:00Z")
fs.get_online("customer:42", ["orders_7d"])
fs.training_set(labels, ["orders_7d"], known="then")
A rolling count also decays with no new events, because old orders fall out of the window. So the definition carries a refresh interval and the store rewrites the value on that schedule. The interval is the staleness you’ve agreed to, and training sees exactly the same staleness because it reads the same rows. The known argument picks the mode: "then" adds the ingest filter from the query above, "now" drops it.
Where It Gets Hard
Late data comes first. “As known then” and “as now known” give different training sets, and the default should be the first, because it matches what serving saw. A backfill sets a trap here: replaying history today stamps every row with today’s ingest time, so a known-then read at last March’s label time finds nothing. Backfills must stamp ingest time with when the value would have been available (event time plus the normal pipeline lag) and flag the rows as backfilled.
Expiry semantics come next. When a value is past its TTL, serving can return null, a default or an error. Null is usually right, because “no recent orders known” is information and a default of zero claims knowledge you don’t have; an error suits cases where a wrong guess costs real money. Make it a per-feature setting stored with the definition, and apply it in training too. Facts that expire is the general form of this question.
Definitions change. Make the window 14 days instead of 7 and you have a different feature, so every row carries the definition’s version and training_set refuses to mix versions silently. Storage grows with entities times features times refreshes, because every refresh is a row. Thin old rows only if you keep the value that was current at every boundary you still promise to answer for, and keep history at least as long as your longest label lookback.
Then there’s the ceiling. One file means one writer at a time, and file locking over a network filesystem is unreliable (SQLite is the well-known example), so two machines can’t share it. You’ve outgrown it when a second service needs the same online store, or when the training join no longer fits one disk and one patient afternoon. At that point, move to Feast or a platform; the definitions and the as-of semantics carry over, which is a good reason to keep them boring.
What Version 0.1 Does
A single file, one writer, a C core with a Python binding. Declarative definitions with count, sum, mean, min, max and last. ingest, get_online, get_as_of and training_set, with known-then as the default. Null on expiry, definition versions, and backfill with honest ingest times. It refuses arbitrary user code, streaming sources and anything multi-node. Embeddings are one more value type, a blob, and the nearest-neighbour lookups beside them belong to a one-file vector index rather than here.
Ship the parity test with it. Replay a month of events, build training rows at a few hundred random timestamps, then run the serving read with its clock set to each of those moments. Any mismatch is skew, and the library should fail loudly on it.
Training should never see what serving couldn’t.