Agent OS
An operating system for AI agents, with git as its memory
creator · open source, Apache-2.0 · v0.2.0, October 2026
Why it exists
Every team building agents rebuilds the same plumbing: memory that outlasts a context window, a way back when step 47 goes wrong, checks on the work, a cap on the spend. Agents today run on bare metal.
What I built
Agent OS is the operating system I’m building underneath them. You hand it a goal and what done looks like. A central brain plans the work, each piece runs in its own process, and a monitor that can’t touch anything signs off on the result. Every step lands as a commit in one git repository, so nothing unverified ships and nothing is ever lost.
Where it’s going
A developer writes a few lines of agent logic, plugs it in, and the OS handles planning, memory, monitoring and recovery until the job is done. It’s the Linux moment for agents.
Code & current status (opens in new tab)Architecture notes (opens in new tab)Original design thesis
What it has shown
- Runs real benchmark tasks end to end: 29 trials across 16 Errata-Bench tasks, every one accepted by the benchmark’s admission check.
- Five trials at once, with no interference.
- Survives a crash without repeating itself: a paid model call is never paid for twice.
- Spending is capped per goal and checked before every call.
aControlthe central brain
bExecutionone process per contract
cMemoryone git repository
Follow a goal through the system
You give it a goal. A central brain breaks the goal into contracts, a separate worker process carries out each one, and an observe-only monitor certifies what it finished. Every plan, model call, command, verdict and rollback is a commit in one git repository. Work that was not certified is never delivered, and nothing that happened is lost.
Point at any part, or play the walkthrough, to see what it does, what it leaves on record, and what has been shown of it so far.
The repository’s rollback demo, which CI runs on every push
EVAL id=step-02 passed=False reason=`pytest -q` exited 1 sha=4e88e78 REV id=step-02 to=d59174c remaining_retries=2 EVAL id=step-02 passed=True reason=`pytest -q` exited 0 sha=d407440
Caught and rolled back. A scripted worker breaks the tests on its second step; the retry lands clean. From the single-process kernel this runtime grew out of.
The system, following a goal. The central brain plans and decides (a). Each contract runs in a worker process of its own, watched by a monitor that cannot act (b). What they do becomes commits in one git repository (c): a checkpoint, certified, a failed attempt.
Every part, as text
- Goal (criteria, limits). A task in plain words, the criteria that must hold at the end, a budget and a deadline.
- success criteria: What must hold at the end. The planner has to account for every one, with evidence.
- limits: The budget, the deadline, the models and the call limits are fixed on admission. A plan cannot raise them.
- on record: The goal is recorded before anything else happens.
- Planner (one call per plan). One model call per plan. It is given the goal, its criteria and the world on record: every branch, and the outcome of every worker so far.
- one contract per worker: The task, the outputs to produce, and the criteria its monitor will judge.
- judges every criterion: Each plan assesses every goal criterion, with the evidence for it.
- reads only the record: It never sees a worker’s private context, only what is on record.
- on record: Each plan is a commit on control. A plan that fails to parse changes nothing.
- Runtime (the durable cycle). The loop that makes plans happen. It waits while workers work, and asks the planner to revise when something gives it a reason to: a finished worker, a failure, a refusal.
- collect and deliver: It collects the report of each worker that ended and delivers certified outputs into integration.
- replan or finish: A failure, a refusal or a stopped worker is a reason to replan, not a reason to stop.
- bounds: A run ends when the plan is complete, or at the clock, the budget, or a limit on plans.
- on record: Each effect is recorded before it happens and after. After a crash it is looked up, not repeated.
- Supervisor (owns the processes). It owns the worker processes: it prepares a branch for each contract, starts the worker in its own process group, enforces the deadline, and reaps the whole process tree.
- start: One operating-system process per contract, on its own branch of memory.
- stop: A worker is asked to stop before it is killed, so it can settle a paid call and write its report.
- reap: The whole process tree is reaped at the end. No process escapes.
- on record: A start is recorded before the process exists; after a crash the process is found by its identity.
- Worker (its own process). An operating-system process with one contract, one branch of memory and a set of tools.
- contract in: It starts from integration, with the task, the outputs to produce and the criteria.
- the loop: Model call, tool call, result, again.
- claim out: It ends with a structured claim: completed, blocked, or needing another turn, with its evidence.
- on record: Checkpoints on its own branch: the files, the transcript so far, and the reason.
- Terminal broker (task environment). When workers act on a container they share, every command goes through a broker. Workers share the environment; the brain stays out.
- one at a time: One command runs at a time, whichever worker sent it.
- timeouts: A command past its timeout is killed with the processes it started. Its output is kept and the container stays usable.
- uncertain: A command whose outcome cannot be confirmed is recorded as uncertain and never run again on a guess.
- on record: Each command, its exit status and its output enter a ledger before the worker sees the result.
- Model API (priced, budgeted). Every model call, from the planner, the workers and the monitors, goes through one facade.
- priced: A model with no verified price cannot be called.
- budgeted: Before a call is sent, its worst-case cost is reserved against what is left of the budget.
- paid once: After a crash the stored reply is replayed; the provider is not asked, and not paid, a second time.
- on record: The reply is journaled before anything acts on it.
- Monitor (observe-only). Every worker has a monitor that only observes. It has no tools, so it cannot fix what it finds: it can only say so.
- reads the transcript: As it grows, not only at the end.
- sends verdicts: The worker is shown them and cannot claim completion until it has answered.
- certifies, or refuses: It judges the exact end state against the contract. Only certified work can be delivered.
- on record: A verdict is a note attached to the exact state it judges.
- Memory (one git repository). A session’s memory is one git repository. What other harnesses add as side files is already there, with history, diffs and merges.
- a branch per actor: control holds the plans, integration only certified results, and each worker has its own.
- a commit per checkpoint: Files, the transcript so far, and the reason.
- rollback is an append: A new commit restores an earlier state. The failed attempt stays readable for the next plan.
- typed merges: Handing work over is a typed merge. A conflict is a value to act on; no model invents a resolution.
- Result (reply + record). The final reply, the certified outputs and the complete record.
- the final reply: Written by the planner when every criterion is met.
- a closing reply: If time or budget runs out first, it says what was done and what was not.
- nothing unverified: Only work a monitor certified reaches integration.
- on record: Nothing is deleted: every plan, call, command, verdict and rollback stays in the repository.
Why it is built this way
Git is the memory
A branch is an agent’s context, a commit is a checkpoint, and a rollback is a new commit that restores an earlier state. What other harnesses add as side files (progress notes, summaries, retry logs) is already there, with history, diffs and merges.
Whoever does the work does not grade it
Planning, doing and checking are separate roles with separate contexts. A worker’s claim is not evidence; a monitor’s certificate of the exact end state is.
Benchmarks run as published
The agent runs each task exactly as the benchmark publishes it, with no overrides. It hands the container to the benchmark’s grader whatever its own audit found.
The repository
taste-is-all-you-need
Git memory · multi-process runtime · PythonTwo generations in one repository: the original single-process kernel, where a planner, a worker and a monitor ran as functions in one loop and rolled back on a failed check, and the multi-process runtime that grew out of it, memory layer first.
Memory is one git repository
The central brain
Workers and their monitors
Intent, effect, receipt
Nothing is deleted
Nothing unverified moves forward
The terminal broker and the model facade
A goal ends with an answer
What has been shown
Try it
Python · Git objects & notes · Worktrees · State machines · Claude Agent SDK · Azure OpenAI · Docker · systemd · Harbor · pytest