Context Window Management as a Database: Ask for 12,000 Tokens, Get the Most Useful Ones
An agent is ninety tool calls into a refactor. Step 12 was a grep that returned 52 KB, and all of it is still in the prompt. The user’s instruction from step 1, “never touch db/migrations”, fell out of the window ten steps ago, because the framework drops the oldest messages first. The agent edits a migration. Nothing malfunctioned; the truncation rule did exactly what it was written to do.
Most agent codebases grow this logic somewhere: a pile of rules about what to drop, when to summarize, which messages to keep. It lives in application code, gets rewritten per project, and can’t be tested without running the whole agent. The version worth building is a database call. Hand a store everything the agent has seen or produced, then ask for a window: context_get(budget=12000) returns the most useful items that fit in 12,000 tokens, in an order that’s safe to send. Selection, compaction and token counting sit behind that call, where you can test them without a model.
What Exists, and Where the Gap Is
Memory products are plentiful. Letta grew out of the MemGPT research and keeps memory in tiers, with the model itself deciding what to page in and out. Zep builds a temporal knowledge graph out of conversations, and mem0 stores extracted memories for later retrieval. LangChain and LlamaIndex ship memory modules too, mostly last-N messages, running summaries and similarity retrieval. Model platforms now ship compaction and context editing of their own.
They’re good at long-term memory, and Letta goes after the window problem head on. What they leave open is the per-request question: given 12,000 tokens, which things go into this prompt, in what form, in what order? The answer comes from a model making choices you can’t unit-test, or from a framework that wants to own the agent loop. Platform features belong to one provider, so your policy changes the day you switch.
What’s thinner is the storage layer underneath: typed items with token counts, a deterministic selector, and originals you can always get back. That’s a smaller thing than a memory platform. It could sit under one.
Items, Links and a Ladder of Forms
Model the store as two tables in one SQLite file. Items are messages, tool results, file reads and summaries. Each has a body, a content hash, a timestamp, a source (the tool call ID, or the file path plus the file’s hash), a token count per tokenizer, a pin flag, and an importance the application can set. Links relate items: summarizes, derived_from, replies_to, cites. An FTS5 index over the bodies gives BM25 relevance with no embedding model. Here’s a sketch of the call:
store = open_store("run-4411.db")
store.add("message", "Never touch db/migrations.", pin=True)
store.add("tool_result", grep_output, source="call_12")
window = store.context_get(
budget=12000,
model="primary",
query="login test fails after refactor",
)
window.items # ordered, ready to send
window.manifest # what went in, in which form, and why
store.expand(8841) # the original behind a stub or summary
Each item can enter the window in one of four forms: nothing, a stub (one line with the ID, the size and what it was), a summary, or the full body. Picking a form per item to maximize total score under a token limit is a multiple-choice knapsack, and a greedy pass is enough. Put pins in at full size, start everything else at its stub or at nothing, then spend what’s left on the upgrades with the best score per extra token. That loop is deterministic and can print why each item landed where it did. You want that at 2 a.m., when the agent has forgotten something.
$ ctx get --budget 12000 --explain
pinned 1840 system prompt, 3 user constraints
recent 4100 last 14 messages, verbatim
relevant 3900 9 older items, ranked against the query
summaries 1650 4 summaries, each expandable
stubs 310 17 references, e.g. "grep output, 52 KB, id 8841"
used 11800 of 12000
dropped 61 items, listed below with scores
A stub only helps if the model can follow it, so expand is a tool the agent can call. The 52 KB grep result costs one line until the agent asks for a slice.
Plain rules can cut a lot of context before any model gets involved. Exact duplicates collapse by content hash, so the same file read three times costs its tokens once. Large JSON shrinks to its shape (keys, array lengths, the fields the query matched) with the rest stored behind a reference. A finished subtask, such as a test run that ended green, collapses to its result plus a link to the steps. It’s the same idea as a deterministic reducer between an MCP server and the model, moved from the wire into storage, where it covers every source of context at once.
Summaries come second, through a callback the application supplies, so the store never needs a model key. Results are cached by the hash of the originals plus the prompt version. A summary always links to those originals, and a rolled-forward one is regenerated from them plus the new material, never from the previous summary. That costs more tokens, but facts don’t get rounded off a little further each round.
The Hard Parts Start With Scoring
Everything above is plumbing. The score decides what the model will need next, and nobody can compute that exactly. A workable v0.1 score is a weighted sum of recency (decay by step distance), inbound references (a file the agent edited twice matters more than one it opened once), relevance to a query, and an importance the application sets for things it knows matter. An agent loop rarely has a clean query, so the current subtask stands in. The weights live in config, not code. The defaults will be wrong for your agent, so make them easy to change and easy to measure.
Items also go stale. A file read at 12:04 carries its path and hash, so a changed hash flags the item. That’s a small case of facts that expire, and it only works because every item records where it came from and when. Embeddings can wait; at a few thousand items per run, exact vector search is enough, as in the post on one-file vector search.
Prompt caching fights the optimizer. Providers cache prompt prefixes and bill hits at a reduced rate, so a window that reshuffles on every call pays full price every call. The store has to give up some relevance for a stable prefix: pins first in a fixed order, then history in time order, new items appended at the end. The default call is sticky: it returns the previous window plus whatever was added since, and only re-solves when that would overflow the budget. When it does, it compacts the oldest region in one batch, so the cache breaks once per compaction instead of once per turn. An identical window also means an identical request, which is what response caching in agent retries and CI depends on.
Token counts differ per model, and some tokenizers sit behind a counting endpoint. The store keeps a count per tokenizer, computes it lazily, and treats the budget as a ceiling with a margin. Chat formats add per-message overhead, and the caller has to subtract what the store can’t see: tool definitions and the reply allowance.
The last problem is knowing whether any of it helped. A policy that saves tokens but quietly drops the one constraint the agent needed is worse than truncation. Every window the store serves gets recorded (item IDs, forms, scores, a hash of the text), the same record you’d want from a flight recorder for agents, and it explains failures after the fact. Judging the policy takes a task set: say fifty tasks, each ending in a pass/fail check. Run each at full context, then half, then a quarter, under your policy and under the dumbest baseline there is, keep the last N messages. A policy that can’t beat that at the smaller budgets is complexity for nothing.
What v0.1 Does and Refuses
The first version is a SQLite file, a small library and a CLI: items, links, pins, FTS5 relevance, the greedy selector, sticky windows, the rule-based compaction, expand, and a manifest for every window served.
It refuses to extract entities or facts, to keep memory across users and sessions, to run the agent loop, or to pick your model. Those jobs belong to the memory platforms, which could sit on top of the store.
Who pays: teams whose agents run long enough that tokens show up as a line item. The library is easy to copy, so it can’t be the product. The part worth money is the evaluation suite: re-run tasks at smaller budgets, score each policy, show which items a failed run was missing.
Keep every original. Make everything else a view.