An Embedded Workflow Engine With Five Primitives: Trigger, Condition, Action, Wait and Retry
A user signs up on Monday. If they haven’t activated by Wednesday, you send a nudge email, and if the email provider times out, you try a few more times before flagging the account. That’s the whole spec. Most teams build it as a cron job that scans for users in the wrong state, a nullable nudge_sent_at column and a retry loop written late on a Friday. It holds until the cron overlaps itself, or a deploy lands mid-loop, or the provider times out after actually sending the message and one customer gets three emails.
The version worth building is a workflow engine small enough to read in an afternoon: state in one SQLite file, a library you embed in your own process, and five primitives. A trigger starts an instance from an event, a condition continues or stops it based on recorded data, an action calls your code, a wait sleeps through restarts, and a retry runs an action again within limits. No server, no cluster, no worker fleet. A tool that can’t express a loop or a parallel branch is a tool you can still reason about when it breaks, and that limit is the product.
The Big Engines Are Big on Purpose
Workflow engines grow into platforms, and the platforms are good. Temporal replays your workflow code against a recorded history, which gives you ordinary-looking code that survives crashes and costs you determinism rules plus a service to run. Airflow schedules DAGs of batch tasks, a different job with a different shape. Inngest, Trigger.dev, Restate and Hatchet offer durable execution as a product, Cloudflare Workflows offers it inside that platform, and n8n and Zapier lean on a visual editor and a catalog of connectors. DBOS is the closest relative: durable workflows as a library on Postgres. If you already run Postgres, start there and build nothing.
The gap is smaller and odder. It’s an app that ships with a SQLite file on one machine (a desktop program, a small SaaS on a single server) and needs three or four durable multi-step jobs. A platform is more machinery than those jobs deserve, and cron scripts are less safety than they deserve.
Five Verbs, One File
A definition is plain data, YAML or a builder call in code. This is the signup nudge as a sketch, for a library that doesn’t exist yet.
workflow: activation_nudge
version: 3
trigger: user.signed_up
steps:
- id: pause
wait: 48h
- id: fetch
action: app.load_user
input: { id: "{{ trigger.user_id }}" }
- id: gate
condition: "fetch.activated == false" # false ends the instance as skipped
- id: nudge
action: mail.send_template
input: { to: "{{ fetch.email }}", template: nudge }
key: "nudge/{{ instance.id }}"
retry: { max: 3, backoff: exponential, base: 30s, jitter: true }
The trigger creates one instance per event. It carries an event ID under a unique constraint, so an event delivered twice can’t start two instances (if events come from an embedded event log, the log offset does that job). pause is a row in a timers table, not a sleeping thread, so a restart on Tuesday changes nothing. gate reads fetch, not the signup event, on purpose: a condition may only look at persisted outputs of earlier steps, so checking whether the user activated means loading the user again after the wait. That looks like ceremony. It means every decision can be replayed, tested and explained from the file alone, which a condition that peeks at live state can never offer.
The executor is one loop with one invariant. It claims a due step from a queue (the kind a SQLite-backed job queue provides, with leases and crash recovery already handled), runs the action, then writes the step result, the new cursor and the next timer in a single transaction. Because the queue and the workflow tables share one file, “job done” and “next step scheduled” commit together or not at all.
Retry looks like the simple primitive and isn’t. A policy needs a maximum, backoff with jitter (so 400 instances that failed together don’t retry together) and a way to tell a failure worth retrying from one that never will. A timeout or a 503 deserves another go; a 400 doesn’t. When attempts run out, the instance moves to failed and stays inspectable.
Exactly Once Is Not on the Menu
Suppose nudge calls the provider, the provider sends the message, and the process dies before the commit. On restart the engine sees an unfinished step and runs it again. No engine can prevent that, because the side effect happens outside its transaction. What it can offer is at-least-once execution plus a key that stays the same across attempts: every try receives nudge/8812, and the action’s job is to make a second call harmless. Some providers accept an idempotency key on the request; if yours doesn’t, the action checks its own record of sends first. The engine can make safety possible. It can’t make it automatic, and the docs should say so on the first page.
Versioning is where small engines usually break. Deploy a new definition while 400 instances sit halfway through a two-day wait, and version 4 may not have a step called pause. So pin each instance to the version it started with, store every definition the file has run, and never edit one in place. Old versions stay loadable until their last instance finishes, and the CLI should list which versions still have live instances. Migrating a live instance means mapping its old cursor to a new one, which is hard to get right, so version 0.1 doesn’t try.
Long waits meet clock changes. A timer has to be stored as an absolute UTC time, because a monotonic clock doesn’t survive a restart. When the system clock steps (an NTP correction, a board that boots with the wrong date), a backward jump delays every timer by the size of the jump and a forward jump fires them early. For a nudge email, shrug. For a billing grace period, decide what you want and write it down.
Observability decides whether anyone trusts the engine. “Why is instance 8812 stuck?” should be one command with a plain answer: waiting on pause until Wednesday 09:14 UTC, or on attempt 2 of 3 with the provider’s error text, or leased by a worker that died. That’s all rows in the file already, so the CLI only reads them. Testing is the cheap part. Put the clock behind an interface and a 48-hour wait costs microseconds. Then run random sequences of events, clock advances and executor kills (stop it between any two statements) and assert that every instance reaches a terminal state, no step records two different outputs, and every repeated action call carries the same key.
What Version 0.1 Refuses
Version 0.1 is a library for one language with the five primitives, SQLite state, persisted timers, retries, idempotency keys and a CLI that lists and inspects instances. It refuses parallel branches, sub-workflows, a visual editor and workers on more than one machine. Each of those gets requested within a month, and each one changes what a step means. Two neighbors stay out of scope as well. A build graph that reruns only the API steps downstream of a change answers “what’s stale”, while a workflow answers “what happens next to this customer”. And cron for agents starts new runs on a calendar, where a wait delays a run that already started.
The honest part of the roadmap is the exit. Someone who needs fan-out, human approvals with a UI or workers on three machines has outgrown the tool, so keep definitions as plain data and actions as ordinary functions with idempotency keys, which any bigger engine can call. As for who pays, probably not the people writing nudge emails. Embedded libraries are free by default, and the money, if there is any, sits in what a file can’t give you: a hosted inspector, history across machines, support for a team halfway to Temporal. The four screening questions, which start with who pays, cost an evening and are worth asking first.
When someone asks for the sixth primitive, they’ve outgrown it. Say so.