Embedded Vector Search Is Crowded: What a One-File Vector Database Needs to Stand Out
You have 40,000 support articles and an app that should offer “similar articles”. Embedded at 768 dimensions in float32, that’s 40,000 × 768 × 4 bytes, about 123 MB of vectors. The standard advice is to stand up a vector database: a service, a port, credentials, a client library and a backup story for 123 MB. A file would do.
Below a few million vectors the server is the overhead, which is why the embedded field filled up fast. So the useful question is what a newcomer brings that the free libraries don’t. My answer is one sharp edge you can state in a sentence, plus measurements the others leave you to make yourself.
The field is busy. sqlite-vec, Alex Garcia’s SQLite extension and the successor to sqlite-vss, starts from fast brute-force search in vec0 virtual tables, which is the right place to start at this scale and comes with SQL and a file you may already ship. LanceDB is embedded and stores data in the columnar Lance format. USearch is a compact HNSW library, hnswlib and FAISS are the long-standing libraries, Chroma has an embedded mode, and DuckDB has a vss extension that adds HNSW indexes. If you already ship SQLite, use sqlite-vec first and measure. A new tool has to beat that, and it has to beat it for someone.
Size Arithmetic Settles the First Design Question
One million 768-dimension float32 vectors take 1,000,000 × 768 × 4 bytes, about 3.07 GB. Int8 quantization cuts that to 768 MB, and binary quantization (one bit per dimension) to 96 MB. Those numbers decide more than any index choice, because exact search over a flat array is memory-bound. Take 50,000 vectors: about 38 million multiply-adds, which SIMD handles easily, but also 154 MB of float32 to stream per query, roughly 15 ms at an assumed 10 GB/s on one core. The same scan over int8 reads 38 MB, and over binary vectors under 5 MB. So quantization is a speed feature before it’s a size feature, and below some count flat search wins on simplicity and perfect recall. At a million vectors the float32 scan reads 3 GB per query, and an index starts to earn its place somewhere in the hundreds of thousands, depending on your latency budget.
Above that, HNSW is the usual answer, and its graph isn’t free. With 16 links per node (a common default), the bottom layer stores up to 32 neighbor ids at 4 bytes each, about 128 bytes per vector. For a million vectors that’s 128 MB of links: under 5% of the float32 data, but more than the binary vectors themselves. Quantize hard and the graph becomes the biggest thing in the file.
Quantization costs recall, and how much depends on your embeddings; a published benchmark won’t tell you. Binary vectors rank candidates roughly, so the standard fix is to rescore an oversampled shortlist (say 10 times k) from the full-precision copy. Whatever you choose, measure recall@10 against exact search on a sample of your own queries. A build step that prints that number is worth more than a tuning guide.
One File With Five Sections
The file holds a header, quantized vector blocks, full-precision vectors for rescoring, metadata columns with simple indexes, and the graph. Reads are mmap, so opening a 100 MB file costs page faults instead of a load step, and the OS decides what stays hot. Search runs on the quantized blocks, then touches only the shortlist’s full-precision rows. It’s the SQLite and nginx pattern applied to vectors: one file you can copy, no daemon.
The header matters more than it looks. It pins the embedding model’s ID, the dimension, the metric, whether vectors are normalized, and the quantizer’s parameters. Vectors from different models aren’t comparable even when the dimension matches, and nothing in the math complains: two different 768-dimension models give similarity scores that look plausible and mean nothing. So a query that declares a different model ID than the header should fail. An embedded feature store has the same problem with feature definitions and the same fix, which is to put the version in the file.
The API is five calls. These names are a proposal, not a shipped library.
vdb *db;
vdb_open("docs.vdb", &db); // header pins model id, dimension, metric
vdb_meta m[] = { vdb_int("tenant", 42), vdb_str("lang", "en") };
vdb_add(db, 1001, vec, m, 2); // id, float vector, metadata
enum { K = 10 };
vdb_filter f = vdb_eq("tenant", 42);
vdb_hit hits[K];
int n = vdb_search(db, query, K, &f, hits); // exact if the filter is selective
vdb_delete(db, 1001); // tombstone; space returns on compact
vdb_close(db);
Deletes in HNSW are tombstones. Unlinking a node breaks paths other nodes rely on, so you mark it dead, skip it in results, and rebuild once the dead fraction passes a threshold. The simplest design that survives updates is segments: an immutable graph built in bulk, a small mutable tail searched flat, a tombstone bitmap, and a compaction that rebuilds the lot. Build memory is the other bill. A million-vector build wants the vectors in RAM, about 3 GB of float32 plus the graph, on a machine whose app needs far less to serve. Build on a big machine, ship the finished file, and let the device append to the tail.
Filtered Search Is the Trap
Real queries carry filters: this tenant, this language, documents newer than last month. HNSW degrades under them. When a filter excludes most nodes, the walk keeps stepping onto vertices it can’t return, the allowed nodes are no longer well connected to each other, and the search ends early with too few results or poor ones. The cheap fix is searching without the filter and discarding hits afterward. Through a filter that keeps 1% of rows, finding 10 hits means retrieving about 1,000 candidates first, and it still comes back short when matches cluster.
A newcomer can compete by treating this as a planner problem. Estimate the filter’s selectivity from the metadata index. If the filter keeps fewer vectors than a graph search would touch anyway (often a few thousand distance computations), scan those vectors exactly: it’s faster and recall is perfect. If it keeps most of the data, traverse the graph with the filter as a predicate. In between, widen the search and traverse with the predicate. Research such as ACORN (SIGMOD 2024) builds filter awareness into the graph itself, and a first version doesn’t need it. Make the choice visible, though: an explain call that prints the plan, the way a SQL database would.
What Version 0.1 Picks
Pick one edge. The candidates are a binary small enough to embed where SQLite isn’t already present, with a clean C API and no dependencies; a stable single-file format with quantization and rescoring built in; or filtered search that stays fast and exact. Doing all three is how you become the eighth tool nobody tried. I’d pick filtered search with a printed recall report, because per-user and per-tenant filters are the ordinary case in applications, and a build-time measurement changes decisions.
Version 0.1 does flat search over int8 with full-precision rescoring, one HNSW segment built offline, a flat tail, tombstones, equality and range filters on typed metadata, the three-plan planner, a model-ID header, and a build report with recall@10 against exact search. It refuses multi-vector documents, hybrid keyword scoring, network access and concurrent writers.
Who would switch? Someone with an app that carries per-tenant filters, a corpus under a few million vectors and no appetite for a server, who tried sqlite-vec or hnswlib and hit the filter wall. That includes on-device features, the kind of AI in apps that works, and an agent’s context database that needs similarity search over its own stored items. Everyone else should keep what they have.
Pick one edge, then print the recall.