Verified Absence Is Evidence: A Database for Recording That You Checked and Nothing Was There
An auditor asks you to show that customer 1187 was screened against a sanctions list on March 3. The screening job writes a row when it finds a match, and customer 1187 has no rows. That fits two histories: the job ran and found nothing, or the job never ran. The database can’t tell them apart, so you hold no evidence either way.
This is the closed-world assumption doing what SQL has always done: a row that doesn’t exist is false. RDF and OWL went the other way with an open-world assumption, where a missing statement means unknown. Neither has a place for the sentence you actually want to store: I looked here, in this way, at this time, and found nothing.
An absence observation deserves to be a record with the same standing as a sighting. It has a subject, a scope, a method, a sensitivity (how likely that method is to notice the subject if it’s there), a time window and a result. “Port checked at 14:00, vessel wasn’t there,” “sensor checked, no motion” and “company queried, not on the list” all share that shape, and the differences between them (how good the check was, how wide, how long ago) decide how much each one is worth.
One line to avoid a mix-up: security research uses “negative database” for something else, storing the complement of a dataset so its contents stay private (Esponda, Forrest and colleagues). Different problem. This post moves on.
The Idea Exists Everywhere Except in the Schema
Monitoring tools understand silence. Prometheus has absent() and absent_over_time(), and heartbeat services alert when an expected ping doesn’t arrive. They answer whether something is missing right now, then forget, so there’s nothing to hand an auditor next year. In compliance the records often live in a log table or a folder of PDF exports, which works until somebody has to query it. Ecologists have the math: a survey that finds no sign of a species doesn’t show it’s absent, so they use occupancy models with repeat visits to separate “not there” from “there but missed.” Negative results in most other fields never get published at all, the file-drawer problem. The knowledge exists in every one of these. What’s missing is a small store that treats it as data.
A Record Per Check
Here’s a sketch of the core table:
CREATE TABLE observation (
id INTEGER PRIMARY KEY,
subject TEXT NOT NULL, -- 'vessel:example_1', 'party:acme_ltd'
scope TEXT NOT NULL, -- 'port:rotterdam', 'list:sanctions@2026-10-01'
window_start TEXT NOT NULL,
window_end TEXT NOT NULL,
method TEXT NOT NULL, -- 'port_camera', 'list_api', 'manual'
basis TEXT NOT NULL, -- upstream feed or sensor; same basis, not independent
sensitivity REAL, -- P(method reports it | it is there)
result TEXT NOT NULL
CHECK (result IN ('present', 'absent', 'inconclusive')),
observer TEXT NOT NULL,
evidence_id INTEGER -- raw response or image, stored by hash
);
Three columns carry the idea. scope says where and against what: a list name plus its version, a port, a bounding box. sensitivity is the chance the method would have reported the subject if it were there, which is what separates a check from a glance (an optical satellite pass over a cloudy port has a low one, because clouds block it). And result takes three values instead of two, because a check that timed out or covered half the scope is inconclusive, and filing it as absent is how false assurance gets into a database. The evidence_id points at the same kind of hashed snapshot a provenance store keeps, so the raw API response can be re-read years later.
Questions arrive in four parts: subject, scope, window and a minimum sensitivity. The answer is never an empty set. An outer join against the list of required checks means a query always returns one of four states (present, verified absent, inconclusive or unchecked) with the checks behind it. A sketch of the command line, with invented names:
$ absence ask --subject vessel:example_1 --scope port:rotterdam \
--from 2026-10-04T13:30Z --to 2026-10-04T14:30Z --min-sensitivity 0.5
verified absent: 2 independent checks, both would miss it 16% of the time if present
obs_8812 port_camera sens 0.60 13:58 absent
obs_8813 ais_query sens 0.60 14:02 absent
The arithmetic is worth seeing with small numbers. Say you think a vessel is as likely to be in port as not, and your camera spots a vessel that’s there 60% of the time and never sees one that isn’t. One “absent” report multiplies the odds of presence by the miss rate, 0.4, which takes you from 50% to about 29%. A second independent check gets you to about 14%, a third to about 6%. One check from a sensor with 99% sensitivity gets you to about 1%. One good check beats three poor ones.
The word that matters there is independent. Three checks that read the same upstream feed, or the same camera from the same angle, share a blind spot, and multiplying their miss rates overstates what you know. That’s what basis is for: checks with the same basis count once. The store also keeps the likelihoods and never the posterior. Combining happens at query time with a prior the caller states, so nobody mistakes “6%” for a measurement.
The Hard Part Is What Absent Means
Scope comes first. “Not on the sanctions list” hides a dozen decisions: which list, which version, which spellings of the name, how fuzzy a match was allowed to be. An exact-match search for a transliterated name is a weak check that looks identical in a log to a strong one. So the scope stores the list version and the search parameters, and the evidence row keeps the raw response.
Calibrating sensitivity is next, and nobody hands you the numbers. A vendor’s accuracy figure describes a benchmark, not your port on a foggy morning. Start with coarse bands a person sets, then measure with known positives wherever you can. Seed a listed name into every screening batch, and if the check doesn’t find its own canary, every absent in that batch becomes inconclusive. Record which calibration produced each number.
Record explosion comes after that. The set of non-events is infinite, but the set of checks isn’t: it’s schedule times subjects times methods, a number you can plan for. Store the checks, never the absences themselves. High-frequency sensors need compaction. A motion sensor sampled every second shouldn’t write 86,400 rows a day, so consecutive absent results from one method extend a single record’s window, and a present result closes it. Don’t bridge gaps wider than the method can cover, though: a camera frame at 14:00 and another at 15:00 say nothing about 14:30.
Absence also decays. “Absent at 14:00” says little about 15:00 when the subject moves and a lot when it doesn’t, so a record is true about its own window only. A query for “absent now” needs a fresh check or a freshness rule, which is the policy side of facts that expire.
The worst absence is the check that never ran. The store needs the schedule it expected, a table of required checks (subject, scope, method, cadence), and a gaps query that lists required checks with no record in their window. A silent job then shows up as a gap and can’t pass for a clean result. The schedule itself comes from whatever runs the jobs; cron for AI agents covers the timeouts and overlap rules that decide whether a check happened at all.
Language matters too. “Not found” means three different things on a screen: never checked, checked and absent, checked and unsure. A UI that prints “No results” for all three has rebuilt the original bug. The container and vessel tracking app in Nine Apps Worth Coding is the friendly version: when a carrier’s feed goes quiet, a small importer needs to know whether the ship wasn’t seen or wasn’t looked for. Claim graphs have the same need, since “no outlet has denied this” only means something if the search behind it is on record.
Regulated settings add retention rules. Screening records may need keeping for years, while the customer names inside them fall under rules to minimize or delete personal data. Keep raw evidence in its own retention class, store minimal identifiers in the observation row, and put retention periods in config per method and scope, because regulators differ and the rules change.
Scope for a First Version
Version 0.1 is a SQLite file with the observation table, a table of required checks and three commands: record, ask and gaps. The combining math ships behind a flag that makes the caller supply a prior. It has no network access of its own, so checks run elsewhere and report in. It refuses to declare anything compliant, refuses to pick a prior, and refuses to turn four months of inconclusive into absent.
Write down that you looked.