Skip to content
All projects

Agent OS

An operating system for AI agents, with git as its memory

creator · open source, Apache-2.0 · v0.2.0, October 2026

Why it exists

Every team building agents rebuilds the same plumbing: memory that outlasts a context window, a way back when step 47 goes wrong, checks on the work, a cap on the spend. Agents today run on bare metal.

What I built

Agent OS is the operating system I’m building underneath them. You hand it a goal and what done looks like. A central brain plans the work, each piece runs in its own process, and a monitor that can’t touch anything signs off on the result. Every step lands as a commit in one git repository, so nothing unverified ships and nothing is ever lost.

Where it’s going

A developer writes a few lines of agent logic, plugs it in, and the OS handles planning, memory, monitoring and recovery until the job is done. It’s the Linux moment for agents.

What it has shown

  • Runs real benchmark tasks end to end: 29 trials across 16 Errata-Bench tasks, every one accepted by the benchmark’s admission check.
  • Five trials at once, with no interference.
  • Survives a crash without repeating itself: a paid model call is never paid for twice.
  • Spending is capped per goal and checked before every call.
  1. aControlthe central brain

  2. bExecutionone process per contract

  3. cMemoryone git repository

Follow a goal through the system

You give it a goal. A central brain breaks the goal into contracts, a separate worker process carries out each one, and an observe-only monitor certifies what it finished. Every plan, model call, command, verdict and rollback is a commit in one git repository. Work that was not certified is never delivered, and nothing that happened is lost.

Point at any part, or play the walkthrough, to see what it does, what it leaves on record, and what has been shown of it so far.

The repository’s rollback demo, which CI runs on every push

EVAL id=step-02 passed=False reason=`pytest -q` exited 1  sha=4e88e78
REV  id=step-02 to=d59174c remaining_retries=2
EVAL id=step-02 passed=True  reason=`pytest -q` exited 0  sha=d407440

Caught and rolled back. A scripted worker breaks the tests on its second step; the retry lands clean. From the single-process kernel this runtime grew out of.

10 parts

The system, following a goal. The central brain plans and decides (a). Each contract runs in a worker process of its own, watched by a monitor that cannot act (b). What they do becomes commits in one git repository (c): a checkpoint, certified, a failed attempt.

Every part, as text
  1. Goal (criteria, limits). A task in plain words, the criteria that must hold at the end, a budget and a deadline.
    1. success criteria: What must hold at the end. The planner has to account for every one, with evidence.
    2. limits: The budget, the deadline, the models and the call limits are fixed on admission. A plan cannot raise them.
    3. on record: The goal is recorded before anything else happens.
    Shown so far: In the real-model trials each goal was a benchmark task, run as published in its own container.
  2. Planner (one call per plan). One model call per plan. It is given the goal, its criteria and the world on record: every branch, and the outcome of every worker so far.
    1. one contract per worker: The task, the outputs to produce, and the criteria its monitor will judge.
    2. judges every criterion: Each plan assesses every goal criterion, with the evidence for it.
    3. reads only the record: It never sees a worker’s private context, only what is on record.
    4. on record: Each plan is a commit on control. A plan that fails to parse changes nothing.
    Shown so far: In the first real-model trials, a completed task took two plans.
  3. Runtime (the durable cycle). The loop that makes plans happen. It waits while workers work, and asks the planner to revise when something gives it a reason to: a finished worker, a failure, a refusal.
    1. collect and deliver: It collects the report of each worker that ended and delivers certified outputs into integration.
    2. replan or finish: A failure, a refusal or a stopped worker is a reason to replan, not a reason to stop.
    3. bounds: A run ends when the plan is complete, or at the clock, the budget, or a limit on plans.
    4. on record: Each effect is recorded before it happens and after. After a crash it is looked up, not repeated.
    Shown so far: In a drill where working time ran out, the run still closed with a reply that named what was run and what was not verified.
  4. Supervisor (owns the processes). It owns the worker processes: it prepares a branch for each contract, starts the worker in its own process group, enforces the deadline, and reaps the whole process tree.
    1. start: One operating-system process per contract, on its own branch of memory.
    2. stop: A worker is asked to stop before it is killed, so it can settle a paid call and write its report.
    3. reap: The whole process tree is reaped at the end. No process escapes.
    4. on record: A start is recorded before the process exists; after a crash the process is found by its identity.
    Shown so far: When the benchmark’s own time limit cancelled the agent, the trial was sealed within 5 seconds, with its commands recorded in order.
  5. Worker (its own process). An operating-system process with one contract, one branch of memory and a set of tools.
    1. contract in: It starts from integration, with the task, the outputs to produce and the criteria.
    2. the loop: Model call, tool call, result, again.
    3. claim out: It ends with a structured claim: completed, blocked, or needing another turn, with its evidence.
    4. on record: Checkpoints on its own branch: the files, the transcript so far, and the reason.
    Shown so far: On a benchmark task, one worker at a time works in the task’s container.
  6. Terminal broker (task environment). When workers act on a container they share, every command goes through a broker. Workers share the environment; the brain stays out.
    1. one at a time: One command runs at a time, whichever worker sent it.
    2. timeouts: A command past its timeout is killed with the processes it started. Its output is kept and the container stays usable.
    3. uncertain: A command whose outcome cannot be confirmed is recorded as uncertain and never run again on a guess.
    4. on record: Each command, its exit status and its output enter a ledger before the worker sees the result.
    Shown so far: In a drill, a command that ignored termination and left a detached child was ended alone in 1.1 seconds; the container stayed usable.
  7. Model API (priced, budgeted). Every model call, from the planner, the workers and the monitors, goes through one facade.
    1. priced: A model with no verified price cannot be called.
    2. budgeted: Before a call is sent, its worst-case cost is reserved against what is left of the budget.
    3. paid once: After a crash the stored reply is replayed; the provider is not asked, and not paid, a second time.
    4. on record: The reply is journaled before anything acts on it.
    Shown so far: The trials ran gpt-6-astra on Azure OpenAI in every role. A run stopped when its working time ran out still reported its exact cost.
  8. Monitor (observe-only). Every worker has a monitor that only observes. It has no tools, so it cannot fix what it finds: it can only say so.
    1. reads the transcript: As it grows, not only at the end.
    2. sends verdicts: The worker is shown them and cannot claim completion until it has answered.
    3. certifies, or refuses: It judges the exact end state against the contract. Only certified work can be delivered.
    4. on record: A verdict is a note attached to the exact state it judges.
    Shown so far: With the evidence itself in front of it, the certifier finished a long task in under five minutes for about $4.40, against 17 minutes and $18.58.
  9. Memory (one git repository). A session’s memory is one git repository. What other harnesses add as side files is already there, with history, diffs and merges.
    1. a branch per actor: control holds the plans, integration only certified results, and each worker has its own.
    2. a commit per checkpoint: Files, the transcript so far, and the reason.
    3. rollback is an append: A new commit restores an earlier state. The failed attempt stays readable for the next plan.
    4. typed merges: Handing work over is a typed merge. A conflict is a value to act on; no model invents a resolution.
    Shown so far: The memory layer has its own test suite, with a test for every defect an adversarial review found, and a crash-consistency demo.
  10. Result (reply + record). The final reply, the certified outputs and the complete record.
    1. the final reply: Written by the planner when every criterion is met.
    2. a closing reply: If time or budget runs out first, it says what was done and what was not.
    3. nothing unverified: Only work a monitor certified reaches integration.
    4. on record: Nothing is deleted: every plan, call, command, verdict and rollback stays in the repository.
    Shown so far: 29 real-model trials of 16 Errata-Bench tasks, all accepted by that benchmark’s own admission check.

Why it is built this way

Git is the memory

A branch is an agent’s context, a commit is a checkpoint, and a rollback is a new commit that restores an earlier state. What other harnesses add as side files (progress notes, summaries, retry logs) is already there, with history, diffs and merges.

Whoever does the work does not grade it

Planning, doing and checking are separate roles with separate contexts. A worker’s claim is not evidence; a monitor’s certificate of the exact end state is.

Benchmarks run as published

The agent runs each task exactly as the benchmark publishes it, with no overrides. It hands the container to the benchmark’s grader whatever its own audit found.

The repository

taste-is-all-you-need

Git memory · multi-process runtime · Python

Two generations in one repository: the original single-process kernel, where a planner, a worker and a monitor ran as functions in one loop and rolled back on a failed check, and the multi-process runtime that grew out of it, memory layer first.

Memory is one git repository

An actor’s context is a branch with its own working tree, held under a lease. A checkpoint is a commit: the files, the transcript so far, and the reason. A verdict is a note attached to the exact state it judges, and a message is a typed, acknowledged entry in the recipient branch’s inbox. Three branches matter to the central brain: control holds its plans, decisions and model receipts; integration holds certified results and nothing else; each worker has its own.

The central brain

One process owns the goal from start to finish. The planner makes one model call per plan: given the goal, its criteria and the world on record, it returns one contract per worker, its assessment of every criterion with the evidence, and whether the goal is complete. It never sees a worker’s private context. The runtime is the loop that makes plans happen: it collects the reports of workers that ended, delivers certified outputs, accounts for spending, and asks the planner to revise when a finished worker, a failure or a refusal gives it a reason to. The supervisor owns the worker processes, from the branch it prepares for each contract to the process tree it reaps at the end.

Workers and their monitors

A worker is an operating-system process with one contract, one branch of memory and a set of tools. It loops through model call, tool call and result, and ends with a structured claim: completed, blocked, or needing another turn, with its evidence. Its monitor only observes. It reads the transcript as it grows and records verdicts the worker must answer before it can claim completion, then judges the exact end state against the contract and certifies it or refuses. The monitor has no tools, so it cannot fix what it finds: it can only say so.

Intent, effect, receipt

Anything that cannot be undone (a model call, a process start, a command, a delivery) is recorded as an intent before it happens and as a result after. A crash between the two leaves an intent with no result. On restart that effect is looked up, not repeated: a paid reply is replayed from its journal, a started process is found by its identity, and a command with no confirmed outcome is marked uncertain.

Nothing is deleted

A rollback appends. A failed attempt keeps its files, its transcript and the reason it failed, which is what the next plan is written from.

Nothing unverified moves forward

A worker’s claim is not evidence. Work reaches integration only after a monitor has certified the exact state it came from. Handing work over is a typed merge whose conflicts are values to act on: git computes them exactly, and no model is asked to resolve one, because it would invent a combination that no worker produced.

The terminal broker and the model facade

When workers act on a container they share, one command runs at a time, and each command, its exit status and its output enter a ledger before the worker sees the result. A command that outlives its timeout is killed with the processes it started, and the container stays usable. Every model call goes through one facade: a model with no verified price cannot be called, and a call’s worst-case cost is reserved against what is left of the budget before it is sent.

A goal ends with an answer

The budget, the deadline, the models and the call limits are fixed when the goal is admitted, and a plan cannot raise them. A goal that owes a reply reserves part of its time for it. If the clock, the budget or the planner fails first, the workers are stopped in good order, their reports are collected, and the planner is asked once more for a closing reply that says what was done and what was not.

What has been shown

The memory layer has its own test suite, with a test for every defect an adversarial review found, and a crash-consistency demo. The single-process kernel has recorded runs with Claude, including a feature added in 43 seconds for about $0.10, and a rollback demo that CI runs on every push. The multi-process runtime has run 29 trials of 16 Errata-Bench tasks with real models, all accepted by that benchmark’s own admission check.

Try it

One pip install of the v0.2.0 release, then the rollback demo, which needs no API key. The memory layer also works on its own. The runtime is a Python library; the taste command drives the single-process kernel, with a live dashboard.

Python · Git objects & notes · Worktrees · State machines · Claude Agent SDK · Azure OpenAI · Docker · systemd · Harbor · pytest

zanwenfu/taste-is-all-you-need