Cron for AI Agents: Schedules, Timeouts, Overlap Rules and Cost Ceilings in One Config File
A nightly agent job triages new issues. One night an issue arrives with a 400 KB log attached. The agent reads it, a tool call fails, the agent retries, and the loop runs until the month’s token budget is gone. Meanwhile cron, which has no idea what a token is, starts the next run on schedule. Now two copies are fighting over the same issue tracker.
Cron was built for jobs that cost nothing, finish in seconds and do the same thing every time. Agent jobs cost money per token, run for an unpredictable time and can loop. The useful ones are still boring and periodic: summarize what changed in a repo since yesterday, triage new issues each morning, check a competitor’s pricing page weekly. What they need is a scheduler that stays as small as cron and adds four things: a cost ceiling, a timeout that actually kills, an overlap rule, and state carried between runs. Anything past that is an orchestration platform, a different and much larger product.
What Already Exists
Most of the scheduling problem is solved. Anacron runs jobs a machine missed while it was off. systemd timers with Persistent=true catch up after downtime the same way. A Kubernetes CronJob has a concurrencyPolicy of Allow, Forbid or Replace for overlapping runs, plus startingDeadlineSeconds to limit how late a missed run may still start. An agent scheduler should copy all of that.
None of those tools knows a run costs money. Workflow engines such as Temporal, Airflow and Inngest are the other neighbor. They handle durable multi-step execution, and they come as services you have to run, with concepts to learn. Agent platforms, hosted and self-hosted, mostly bundle scheduling inside something bigger. The gap is small and specific: one binary, one plain config file, and the four additions.
One File, Four Additions
Here’s the shape of that file for a proposed tool (the command names are illustrative):
[defaults]
timezone = "Europe/Berlin"
overlap = "skip" # skip | queue | replace
on_missed = "run_once" # run_once | skip
catch_up_within = "2h" # later than this, a missed run is skipped
[job.repo-digest]
schedule = "0 7 * * 1-5"
command = "agent-cli --model primary --prompt prompts/digest.md"
timeout = "10m"
max_cost = 0.75 # dollars per run
max_calls = 60 # model calls per run, a crude loop breaker
retries = 2
backoff = "2m"
carry = ["last_output", "kv"]
notify = { on = ["failure", "budget_stop"], webhook = "https://hooks.example.com/agents" }
[job.issue-triage]
schedule = "*/30 * * * *"
command = "agent-cli --model cheap --prompt prompts/triage.md"
timeout = "5m"
max_cost = 0.20
The command is any command line, so the scheduler doesn’t care which agent runs. It sets environment variables for the job: the time of the last successful run, the path to its state directory, the budget left. That’s how “since last run” works. The previous run’s output and a small key-value store carry over, and everything lives in one SQLite file next to the config.
Overlap defaults to skip, because a job that overran its slot usually means the next one would too, and two agents editing the same thing is worse than one missed run. Queue holds at most one waiting tick, so three ticks piling up behind a slow run produce one run, not three. Replace kills the old run, which suits jobs where only the latest answer matters. Missed runs get a deadline, borrowed from startingDeadlineSeconds: a morning digest that would start at 3 p.m. because a laptop slept through 7 a.m. is stale, so skip it.
A Ceiling Means Sitting in the Request Path
The scheduler can’t read tokens out of an arbitrary process. What works is a local proxy that the scheduler starts and points the job at through an environment variable, since many SDKs read their base URL from one. The proxy reads usage from each response, multiplies by a price table you maintain in the config (prices change, so they can’t live in code), and keeps a running total per run.
Usage is only known after a response completes, so checking between calls can overshoot by one full response. The proxy refuses a call before forwarding it if the input alone would cross the ceiling, and clamps the output limit to what’s left. At the ceiling it kills the run and records a budget stop. The wall-clock timeout works the same way: terminate, wait out a grace period, then kill the whole process group, because the children an agent spawns (a shell, a test runner, a headless browser) outlive their parent otherwise. A proxy in this position could route by cost ceiling too, but that’s a separate tool’s job, covered in routing LLM requests from one config file.
The Hard Parts Are Time and Side Effects
Daylight saving time is where schedulers go to be wrong. Classic cron implementations differ on the skipped and repeated hours, which is the reason to pick a rule and print it. For fixed-time jobs like 0 7 * * 1-5, fire once per local day at 07:00. If that time falls in a skipped hour, fire at the first instant after the gap; in a repeated hour, fire on the first pass only. Interval jobs fire on real elapsed time, so the change doesn’t touch them. A sched next repo-digest command should list the next five fire times in local time and UTC, so you can check the decision before it bites. Not scheduling anything between 01:00 and 03:00 local helps too.
Retries are the next trap. A retried agent isn’t replaying anything; it makes fresh decisions, so the second attempt may do something the first never did, or do the same thing twice. Say the first run posts a comment on an issue, then dies on a timeout. The retry reads the issue, sees a comment, and posts another. So retries should default to read-only jobs. A job that changes the world needs idempotent tools (a comment that carries a marker the agent checks for) or a human in the loop. The scheduler can’t enforce a tool allowlist inside an agent it didn’t write. It can only pass the list along, which means the real boundary has to sit below the agent.
State has its own failure mode. “Since last run” works only if the cursor moves on success. A run killed at its ceiling after reading half the issues must not mark all of them as seen, so the scheduler commits the last-success time only on a clean exit and gives the job a separate checkpoint key for partial progress. Overlap needs a lock that survives a crash of the scheduler itself. A row holding the PID and the process start time (the start time guards against PID reuse) lets a restarted scheduler tell a live run from a dead one.
An agent running at 3 a.m. with nobody watching should also have a tight network policy and a record of what it did. That layer sits below the scheduler and should stay there. VPN Works is one example: alpha software that seals an agent in a Linux network namespace whose only way out is through its own process, checks every connection against a short policy and writes each one to a record. Because a job’s command is just a command line, a wrapper like that slots in front of any job. Per-run logs (prompt, model calls, exit status) go in one directory per run, and packing a directory into one file you can hand to someone is the job of an agent flight recorder.
What v0.1 Does and Refuses
The first version is one binary and one TOML file: cron expressions with time zones, the metering proxy, timeouts that kill the process group, the three overlap policies, missed-run catch-up with a deadline, retries with backoff, the SQLite state file, webhooks, and a log directory per run. It runs under systemd and loses nothing on restart.
It refuses the rest. No DAGs, no branching, no waiting on human approval, no web UI, no model routing, no tool hosting. A job that needs steps and waits has outgrown a scheduler and belongs in a workflow engine with a handful of primitives. General queueing (priorities, fan-out, many workers) is what a job queue in one SQLite file already covers.
Keep it boring.