40 Projects Worth Building: SQLite for X, Tiny Infrastructure, MCP Tools and Edge Systems
The most promising territory for a small team isn’t another large framework. It’s the opposite: infrastructure software you can explain in one sentence, ship as one binary or embed as one library, run with almost no operational burden, and that solves one irritating problem unusually well. SQLite and nginx are the right mental models because their appeal is architectural, not just functional: small surface area, predictable behavior, few dependencies, easy deployment, and usefulness far beyond the original use case. (More on that in the SQLite and nginx pattern as a product strategy.)
There’s a practical advantage for a one-person, AI-assisted project too. ChatGPT or Claude can reason about a 10,000–30,000-line C, C++, Rust, Zig or Go codebase far more reliably than about a sprawling distributed system with 40 services.
The current open source landscape backs this up. EmbedDB gets down to a 4 KB minimum memory requirement while combining KV and time-series storage with relational features, and NanoTS is pursuing a compact embedded time-series architecture. Fluent Bit shows that a “tiny data processor sitting in the middle” can grow into an enormous category: it now handles logs, metrics and traces, with SQL stream processing and 70+ plugins.
Data and storage
1. SQLite for API responses. One of the strongest ideas here. A single embeddable library, call it ApiDB, sits between an application and external APIs. Every request and response is persisted automatically, indexed by URL, parameters and time, deduplicated, compressed and queryable with SQL. GET /weather?id=123 becomes a key, the JSON response becomes a value, and selected JSON fields become queryable columns: SELECT temperature FROM api_cache WHERE endpoint='/weather' AND timestamp > .... Add TTLs, stale-while-revalidate, request coalescing and historical snapshots. Suddenly it’s an API cache, an offline API, a debugging database and a tiny data warehouse at once. The first release could genuinely be small.
2. SQLite for JSON streams. Not a general database: a tiny engine that accepts arbitrary JSON objects and makes them queryable right away, no schemas. stream.write(json), then something like SELECT avg(cpu) FROM stream WHERE host='x' WINDOW 5m. It could run entirely in RAM, optionally spill to one file, and keep incremental aggregates instead of retaining everything. Clean positioning between jq, SQLite, stream processors and observability systems.
3. An aggregate database. Worth exploring even more seriously. Call it RollupDB. It deliberately doesn’t keep raw data forever. You feed it millions of observations and say what must survive: count, sum, min, max, average, variance, percentiles, top-K, cardinality, histograms, maybe approximate distinct counts. Old observations fold into ever coarser buckets: seconds, minutes, hours, days, months. A 100 GB telemetry stream might become a 50 MB database that’s still useful for historical analysis. Good for IoT, monitoring, analytics and AI agent telemetry. And its purpose is instantly clear.
4. SQLite for logs. Don’t compete head-on with full observability stacks. Build LogLite instead: one executable that reads stdin, files or syslog, recognizes common formats, extracts fields, deduplicates repeated messages, compresses them and writes one queryable file. cat nginx.log | loglite access.llog, then loglite query "status >= 500 group by path". No Elasticsearch, no daemon, no JVM, no Docker. Fluent Bit proves the demand for lightweight telemetry processing, but its feature set is now broad enough that a deliberately tiny local tool could hold a different niche.
5. Semantic log reduction. A particularly interesting twist. Instead of shipping 10 million nearly identical log lines to an expensive observability backend, a tiny local process fingerprints templates and sends something like “error template #731 occurred 43,282 times; here are 10 representative samples.” Deterministic parsing, not an LLM, keeps it extremely cheap, with optional AI processing only for newly discovered patterns. That’s potentially a company, not just a GitHub utility, because cutting telemetry ingestion has an immediately measurable dollar value.
6. PrecomputeDB. An unusual database category that optimizes answers instead of storage. The developer declares the queries or functions that matter, and the engine keeps their results current as the underlying values change. Materialized views reduced to a tiny embeddable engine: watch("SELECT count(*) FROM events WHERE country='US'"), and reads become effectively O(1). There’s a lot of serious database theory behind incremental view maintenance, but a constrained implementation could be surprisingly small. This is the territory Precomputing works in, and where a learned query optimizer could eventually meet it.
7. DerivedDB. Takes that a step further. It stores raw facts plus deterministic derived values. Put {price: 100, qty: 4} in, define total = price*qty and tax = total*rate, and the engine maintains the dependency graph automatically. SQLite meets spreadsheet recalculation. Useful for configuration systems, dashboards, financial models and reactive applications. The novelty is in making it extremely small, not in building another application framework.
8. StateDB. Tiny persistent application state. Programs constantly need counters, flags, timestamps, leases, queues, locks, configuration, small objects and TTL values, and developers often pull in Redis or SQLite for something embarrassingly simple. A one-file embedded state engine with get, set, incr, CAS, expire, queue, transactions and crash-safe persistence: Redis semantics without Redis. The constraint is part of the product. No networking, no cluster, no administration.
9. QueueDB. SQLite-like embedded durable queues. One library, one file, multiple named queues, visibility timeouts, retries, delayed jobs, priorities and dead-letter queues. Applications often deploy Redis, RabbitMQ or a cloud service just to get reliable background jobs. A boring, crash-safe, zero-admin embeddable queue could win real developer adoption. It also suits AI-assisted implementation well: the API stays tiny, and correctness can be attacked hard with property tests and crash simulations.
10. EventDB. An embedded append-only event store. Not a Kafka competitor; deliberately Kafka for one machine. Append events, subscribe from sequence N, keep consumer offsets, compact streams, query metadata. One executable, one file, maybe memory-mapped. Useful for desktop apps, edge servers, agents, industrial systems and local-first software.
11. SyncDB. More ambitious and more startup-like: two embedded databases that synchronize directly. SQLite-style local storage with built-in change tracking, conflict resolution and peer replication. The first version doesn’t need to solve arbitrary distributed databases. Keep the model KV or document-oriented and constrain the conflict semantics. Local-first software stays attractive because developers want applications that keep working without a permanent cloud connection. A truly simple db_sync(a, b) API would have huge conceptual appeal.
12. The SQL/KV hybrid, sharpened. Not “SQLite plus KV.” The principle is that every record has two personalities: the same bytes can be reached by key in microseconds or read relationally when SQL is needed. No duplicate storage, no translation layer. That’s a stronger proposition than supporting two APIs, and it’s the idea behind AltSql’s native hybrid. EmbedDB shows the KV/relational/embedded intersection is real, down to very constrained hardware, so the differentiator has to be crystal clear: identical storage from MCU to gateway, very simple source distribution, or SQL over the exact same KV records.
13. SQLite for vectors. Room exists, but not for another generic vector database. TinyVector would be one embeddable file holding vectors, metadata and an ANN index, aimed at tens of thousands to a few million vectors. No server, no accounts, no cloud; tinyvector.db travels with the application. Very GitHub-friendly, but the field is crowded, so it needs a memorable edge such as an astonishingly small binary or a particularly clean C API.
14. ContextDB. A better AI database bet: a tiny storage engine built for LLM and agent context. It stores messages, tool results, summaries, embeddings and provenance, but the key feature is automatic compaction. As context grows, older material is summarized, deduplicated or indexed while references to the originals survive. The application asks context_get(token_budget=12000) and gets the most useful context that fits. Context-window management becomes infrastructure instead of application code.
15. PromptCache. A tiny semantic and deterministic cache for LLM calls. It hashes the model, system prompt, relevant parameters and normalized input, stores outputs, optionally detects near-equivalent requests, tracks cost and latency, and supports TTL and invalidation. A single executable could run as an OpenAI- or Anthropic-compatible proxy. The economic pitch is simple: install it and spend less on APIs.
MCP tools
16. Not another MCP gateway. That space is filling fast. Existing projects already aggregate servers, bridge stdio and HTTP transports, lazy-load tool schemas, and handle authentication, logging, rate limiting, caching and observability. Some advertise context reduction by loading tools only when needed. A new project needs a narrower hook.
17. MCP-nginx. One such hook: not an MCP platform but an absurdly small MCP traffic engine. One binary, one config file, maybe in Rust or Zig. Route tools, rename tools, hide tools, rate-limit calls, cache responses, cap response sizes, log traffic and proxy transports. No dashboard, database or orchestration. Conceptually: agent → mcpx → MCP servers. Existing projects are moving this way, so size and speed alone probably won’t be enough. The differentiator would be a genuinely nginx-like configuration and extension architecture.
18. MCP Context Firewall. Much more interesting. Its job isn’t security in the usual sense but controlling what enters the model’s context. A tool returns 8 MB of JSON; the firewall extracts the 15 KB that matter. A server exposes 150 tools; the firewall shows five meta-tools first and reveals real schemas only when they’re relevant. Responses can be filtered, projected, truncated, summarized, cached or turned into references. MCP gateways already recognize schema and context waste as a problem, which validates the pain. A dedicated context-reduction engine could go much deeper.
19. MCP Reducer. One of the strongest ideas on the list. Put it between the MCP server and the model. Rules say things like github.search → keep $.items[*].{name,url,description}, database.query → max 100 rows, filesystem.read → strip binary, or tool response > 20KB → summarize/index. Then measure original tokens, delivered tokens, tokens saved and estimated cost. Developers get the benefit instantly, and the README metric writes itself: “MCP Reducer cut agent context consumption by 73%.” It can start fully deterministic, with optional LLM reduction later.
20. MCP Cache. A standalone component, not yet another gateway. Many tool calls are repeatable: documentation retrieval, repository metadata, schemas, package information, configuration, public API calls. The proxy decides cacheability, normalizes parameters, stores responses and invalidates them by rule. Eventually it could learn which calls are safe to cache. One executable, perhaps backed by a tiny in-house KV engine, which lets the projects reinforce each other.
21. MCP Record/Replay. Excellent GitHub potential. Run an agent once while recording every MCP interaction into one file. Developers can then replay the session deterministically without the original servers, APIs, credentials or network. VCR for MCP. It covers testing, debugging, CI and bug reproduction: mcprec record claude ..., then mcprec replay session.mcp. The implementation is manageable and the usefulness obvious.
22. MCP Fuzzer. Small and potentially popular. Point it at an MCP server and it discovers tools, generates valid and invalid parameter combinations, tests schemas, and checks huge responses, malformed output, timeouts, concurrency and protocol edge cases, then writes a compatibility report. MCP is expanding fast and testing infrastructure tends to lag implementations. A single-purpose quality tool stands out more easily than the twentieth MCP gateway.
23. MCP Bench. The ApacheBench or wrk of MCP: mcpbench https://server/mcp --tool search --concurrency 100. Measure latency distributions, throughput, startup cost, tool discovery latency, payload sizes and context-token overhead. Very small, very achievable, useful to every MCP server developer, and easy to star because people know what it does at a glance.
APIs, pipelines and agents
24. API Bench with state. Goes beyond MCP. HTTP benchmarks are good at hammering endpoints, but modern APIs involve workflows: authenticate, create an object, query it, update it, delete it. A tiny declarative benchmark engine could run stateful scenarios without turning into a giant testing framework. YAML in, latency and throughput report out.
25. API Reducer. Perhaps commercially stronger. Sit between APIs and applications and strip useless payload. Developers routinely fetch 100 KB JSON responses to use five fields. Define projections at the proxy, cache the transformed result, and optionally compress or re-encode it. Later the reducer could watch client access patterns and recommend reduced schemas. Especially interesting at the edge.
26. JSON Accelerator. A tiny daemon or library that receives JSON and precomputes indexes, paths, aggregates or transformed representations. Applications re-parsing enormous JSON documents again and again is surprisingly wasteful. A persistent binary representation that keeps JSON semantics but allows very fast path lookup could be broadly useful. The pitch: SQLite, but instead of querying tables you query persistent JSON at memory-like speed.
27. Tiny ETL. Has potential if it’s aggressively small. One binary running pipelines like HTTP → JSONPath → filter → aggregate → SQLite, S3 or stdout. Not Airflow, not Kafka Connect, not NiFi. A 10 MB executable and a 15-line config. The real competitors are cron, curl, jq, awk and SQLite glued together, and that’s a good sign. When developers keep gluing five Unix tools together, there’s room for one clean tool.
28. Data Distiller. Even more distinctive. It continuously turns high-volume raw streams into compact information. Unlike compression, the output is lossy on purpose but analytically useful: anomalies, first and last samples, changes, extremes, distributions and representative examples survive, redundant observations don’t. Feed it sensor data, network metrics, logs or API observations. It meets log reduction and precomputation in the middle, and it could become a genuinely new infrastructure primitive. (It’s the idea behind Distill in the money-first screen.)
29. ChangeDB. Deceptively simple: store only changes. Feed it JSON snapshots, files, API responses or configuration objects and it records compact structural deltas. Ask “what was this object on August 1?” or “which fields changed most often?” Git semantics for structured data, embedded and tuned for machine-generated snapshots. It could power website monitors, API monitors, price trackers, configuration auditing and AI agent memory.
30. Snapshot API. Builds on ChangeDB. Give it any URL or API and it stores content-addressed snapshots on a schedule, deduplicating identical portions, with a tiny SQL interface over historical versions. SELECT price FROM snapshots WHERE url=... ORDER BY timestamp then works against APIs that have no history of their own. Both open source and SaaS potential.
31. Tiny Feature Store. For ML developers: one embedded database of keyed features with timestamps, TTLs, atomic updates and point-in-time reads. Most feature stores are infrastructure-heavy. A SQLite-like version could target local inference, edge ML and small AI applications instead of huge training clusters.
32. Model Router Lite. Attractive but crowded unless it picks an unusual constraint. One binary implementing the OpenAI-compatible API and routing requests among OpenAI, Anthropic, local models and others by price, context length, latency or availability. The distinctive version has no UI and no external database and runs on a $5 VPS, with local cost accounting and caching. Useful, but it ranks below the reduction and data ideas because AI gateways are already plentiful.
33. Token nginx. More unusual. Instead of routing HTTP by host and path, it routes AI requests by token characteristics: context size, expected output, tool needs, privacy tags, cost ceiling. Configuration along the lines of: under 4,000 tokens goes to a small model, coding goes to a coding model, confidential goes to a local model, otherwise model B. One binary, one config.
34. Agent Flight Recorder. Serious potential. A lightweight daemon records prompts, model responses, tool calls, MCP traffic, timing, token counts, errors and parent-child relationships into one portable file. When an agent behaves strangely, you send someone bug.agentlog and they can replay and examine exactly what happened. Chrome HAR files or Linux strace, for AI agents. Better than another observability dashboard, because a portable standard file format could become valuable in itself.
35. Agent strace. The even smaller version: run agenttrace command... and watch every model, tool, file and network action as a terminal stream. No backend, no account, no telemetry cloud. Filters like --tools, --tokens, --errors, --cost. It could output OpenTelemetry later, but the original appeal is that it works in thirty seconds.
36. Agent top. Fun and very GitHub-friendly: a terminal UI like top, showing running agents, active tool calls, tokens per second, context sizes, cost, cache hits and model latency. Probably not a company on its own, but it could get attention and become the front end to Agent Flight Recorder.
37. Cron for agents. A potentially large category if kept simple. A tiny scheduler runs prompts and tools on cron expressions, keeps state in one local file, and handles retries, timeouts and webhooks. Agent platforms are plentiful, but agentcron.toml plus one binary has real elegance. Don’t become an orchestration platform. Being the boring cron replacement is the selling point.
38. Make for APIs. Define dependencies between HTTP and MCP calls and recompute only the downstream nodes whose inputs changed. DAG execution, caching and precomputation in one. If API A hasn’t changed, don’t rerun transformations B, C and D. Incremental builds applied to data and API workflows, with make as the model, not Airflow.
39. Content-addressed API cache. A technically elegant little project. Every API result is immutable and addressed by hash; requests map to hashes; duplicate payloads are stored once. Add Merkle structures and synchronization, historical snapshots and verification all get cheap. It could end up underneath several of the projects above.
40. Embedded analytical sketch database. Worth serious investigation. A tiny engine built around probabilistic structures: HyperLogLog for cardinality, Count-Min Sketch for frequencies, t-digest or KLL for quantiles, Bloom filters for membership, heavy-hitter algorithms for top-K. Feed it billions of observations and the database stays megabytes in size. Query it with a tiny SQL dialect: SELECT approx_count_distinct(ip), percentile(latency, .99), top(url, 20). It fits the low-footprint, precomputing, reduction direction well, and it has real computer science substance without needing a huge codebase.
Reduction, not storage
Collapse all of this into a shortlist and a cluster stands out around reduction rather than storage. MCP Reducer reduces context. Log reduction reduces telemetry. RollupDB reduces historical data. API Reducer reduces payloads. ContextDB reduces accumulated agent memory. Agent Flight Recorder makes what remains inspectable. The shared philosophy: modern systems produce far more data than the consumer needs, so process it near the source and keep information, not bytes. That’s a stronger, less crowded thesis than “another lightweight database.”
For quick GitHub attention, MCP Record/Replay, MCP Bench, Agent strace, QueueDB and LogLite stand out. A developer understands each in ten seconds and tries it in a minute. For deeper startup potential: MCP Reducer and the context firewall, semantic log reduction, RollupDB and Data Distiller, ContextDB, and API-response storage with precomputation. For technical credibility: the sketch database, PrecomputeDB and the SQL/KV hybrid. And for a project that could start at 3,000–8,000 lines written mostly with ChatGPT or Claude and grow from there, MCP Record/Replay is probably the cleanest.
Whatever gets chosen, impose one pattern on it. One binary or one library. No mandatory daemon where embedding works. No external database, no Redis, no Docker requirement. Human-readable configuration. A stable on-disk format where it applies. A C API for libraries, bindings later. And a README where the complete useful example fits on the first screen.
“Download this one file and it solves X” is a product strategy in itself. SQLite and nginx aren’t loved for having fewer features. They’re loved because they make complicated infrastructure feel disproportionately simple.
A family, not a fresh start each time
None of this means leaving the database direction. A coherent family is visible: AltSql as the embedded SQL/KV storage layer, a tiny incremental rollup engine on top of it, and separate MCP, API and log tools that use the same engine inside.
One substantial piece of engineering, several independent GitHub projects, instead of starting from zero every time.