agentplane

Durable execution · Deterministic replay · Policy-governed

The journal is the plan of record.

An agent runtime in Rust where one gate decides all four questions about an outward call — may this principal act, may this value go there, can it be repeated, and what does it leave behind. They are usually answered in four different places, by four systems that cannot see each other's answers.

Get started View source

Pre-alpha · Rust 1.94.1+ · #![forbid(unsafe_code)] · MIT OR Apache-2.0

What goes wrong without it

Why the usual answers stop short

Each of these is a real control, and none of them is careless. The difference is structural, and it is a list you can go and check rather than a claim about who is more serious.

A gateway or perimeter

Authorizes the calls that cross it. A memory write, a model call, a commission and a case mutation never do. It holds no label on the value, and its audit trail is a second log beside execution rather than the one execution recovers from.

A framework hook

Runs inside the process it is governing, so it is advisory against that process. And it fires at a boundary where the argument has already lost its history — no interceptor can reconstruct where a value came from once it is just a string in a call.

A durable execution engine

Journals whatever activity boundary the SDK happens to expose, which makes the workflow survivable. It has no declaration to check an argument against, no label on a value, and nothing it can refuse.

Recovery is the commodity now — durable agents are a few lines of integration. What a bound layer cannot add afterwards is the part below it: a reviewable declaration, a field-level flow decision, an erasure that reaches every copy, and an audit somebody runs without the runtime. How the pieces fit.

What one gate buys

Authorization that sees the arguments

Not may this agent call that tool but may these values populate these fields. A recipient must derive from the tool that looked it up; a value a model chose cannot fill an authority-bearing argument, and the refusal names the field.

Labels that travel with the value

Trust, sensitivity and the sources that shaped it, carried per field and joined on every combination. Untrusted data cannot reach a mutating sink without a typed, policy-authorized, journaled release — scoped to one destination, so it launders nothing else.

Authority that only narrows

Delegation is audience-, time- and depth-bounded, bound at admission and read back on replay. A peer's run executes under its own chain, never the plane's authority, and no downstream card or tool annotation can widen what it holds.

Evidence that is also the recovery mechanism

One log, not two. The record a crash resumes from is the record an auditor reads — append-only, hash-chained, optionally signed, entered in a Merkle log. A separate compliance trail would eventually disagree with execution, and the disagreement would be invisible.

Anchored outside the operator

A checkpoint is cosigned only if it provably extends the last one seen; a shrunken log is rejected, and so is a second history of the same size. The auditor reads the anchor back from the witness rather than from the party being audited — which is the only claim here that survives a dishonest operator.

Erasure that keeps the proof

Drop a payload's bytes and the chain still verifies: it only ever committed to a digest. A later read says expired, on this date, for this reason — never missing. This is the case that decides whether an evidence tier is real, because it is where most of them break.

The recovery you would expect is here too and is deliberately not the headline: resume from the last completed effect, exactly-once enforced by the store rather than by a code path someone might forget, and month-long business processes modelled as cases so a deploy never migrates an in-flight workflow.

Why you should believe any of it

Every claim above is the kind that is easy to write and hard to keep, so none of them is offered on trust. Each is broken on purpose and the test written for it must fail — a guarantee whose removal nothing notices is treated here as a defect, not as a feature that happens to be untested.

14
TLA+ specifications, model-checked on every push
71
deliberately broken specs; each must trip its own check
1385
code mutations, each naming the one test that must fail
0
unsafe blocks — forbid(unsafe_code)

A guarantee can be implemented, tested and green while deleting it fails no test. Only removing it shows that, so the sweep removes each one — how this is proven is the long answer.

Thirty seconds

cargo add agentplane
let store: Arc<dyn JournalStore> = Arc::new(RedbStore::open_in_memory()?);

let runtime = Runtime::builder(Arc::clone(&store))
    .owner("my-service")
    .skill(Triage)
    .build();

let out = runtime.run("ticket.triage", Tainted::trusted(json!({ "text": "printer on fire" }))).await?;

// Re-executes the logic and reads every effect back. Nothing is performed again.
let replayed = runtime.replay(out.run_id, Mode::Strict).await?;
assert_eq!(out.output, replayed.output);

Read the guide Cookbook

What it looks like when it works

One command, no account, no key. It runs a pipeline, replays it, crashes it mid-run, resumes it, and then replays it against a changed build.

cargo run --example durable_pipeline
1. live run      → Succeeded
   external calls: 3

2. strict replay → Succeeded
   external calls: 3 (unchanged: true)

3. run crashed   → Failed("simulated crash after stage 0")
   external calls: 1
   resumed        → Succeeded
   external calls: 3 — stage 0 was replayed, not repeated

4. changed build → Quarantined("non-determinism at seq 8: expected ek:cf87…")

5. all journals verify — no record was altered after the fact

Line 2 is the one to read twice: replaying performed nothing. Line 3 is the one that pays for itself. Line 4 is a different build refusing to rewrite history rather than quietly accepting it.

Questions people ask before adopting it

Does this replace Temporal, Restate or DBOS?

No — it answers a different question, and the two compose. Durable execution engines make a workflow survivable; the recovery unit there is the activity the SDK happens to expose. Here the unit is the effect, and the same gate that decides whether a call may be repeated also decides whether it is authorized, what label its result carries, and what evidence it leaves. Governance below framework code is the part a durable backend cannot add later, because it has no declaration to check an argument against.

Do I need a database?

No. The default backend is redb, embedded and single-node — one file, no server. PostgreSQL implements the same store contract and is what several plane instances share when fencing and exactly-once have to be arbitrated by the database rather than by hoping the writers agree. Both pass the same conformance battery.

Is it production ready?

No, and the honest reason is a format rather than a feature: the durable record format is not frozen, so a shape change is a hard cut with no migration. The crate is published and pre-alpha. Status lists what will move, what is deliberately absent, and the checkable conditions the freeze is waiting on. The workable position today is to treat the export as the long-term artifact and the store as disposable.

What does “replay” actually do?

It re-executes your logic and satisfies every effect from the journal. No tool is called again, no clock is read again, no model is asked again. A run that took forty minutes and €40 of inference replays offline, for free, and reaches the same result — or reports precisely where this build disagrees with the recorded one.

Does it route or proxy model traffic?

No. It ships drivers for a few providers because a runtime with none cannot be tried, but the seam is the point: a `ModelProvider` you write is a first-class citizen, and metering, retry classification and replay work the same behind any of them. Routing and load-balancing belong to something else.

Does using it make a deployment EU AI Act compliant?

No. Compliance is a property of a deployment and its documentation, not of a dependency. What this provides is technical means — an append-only record of what an agent did, human oversight that binds, and erasure that survives the proof — mapped article by article, with the obligations it does not touch named as plainly as the ones it does. The mapping says which is which.

What it is not

agentplane does notUse instead
Ship a prompt library or IDEYour manifests; agentplane hashes and versions them
Route or proxy model trafficLiteLLM, Bifrost, your own ModelProvider
Implement a vector databaseLanceDB / pgvector behind a seam
Claim regulatory complianceIt provides technical means; compliance is the deployer's

Status lists what will move and what is deliberately deferred, and the security model is explicit about what it does not cover.