agentplane

Operations

Deploying, high availability, retention, observability, and a runbook for every state a run can get stuck in.

On this page
  1. Ownership and fencing
    1. A live run renews; only a dead owner’s lease expires
    2. The owner string identifies a process, not an agent
    3. Releasing frees the lease without forgetting the epoch
  2. Stopping an instance
    1. What a caller is told while an instance drains
    2. The endpoint race is the deployment’s to close
  3. Deploying a plane
  4. Two backends, one contract
    1. At-most-once admission is part of the contract
    2. The case layer
    3. Postgres
  5. What the gate costs
    1. One axis dominates, and it is not the one people ask about
    2. Read the gap, not the numbers
  6. The sweeper
    1. The recovery pass: who resumes a crashed run
    2. A finding has to be findable
    3. Answering a quarantine
    4. Retention for the admission index
    5. Taking the record away
    6. Disclosing one matter
    7. Re-deriving the policy verdicts
    8. Replaying an edited declaration
    9. One run’s timeline
    10. The command line’s exit statuses
    11. Which verb needs which feature
    12. What no audit can answer: the runs that never started
    13. Break-glass
    14. One matter, one scan
    15. The sweep writes its own history
    16. A capped tick says it was capped
    17. Outbound delivery: two ceilings, and what each one means
  7. Metrics
    1. The runtime does not measure durations
    2. Counters are emitted; gauges are observed
    3. Two rules, both guarded
  8. Observability
    1. What is deliberately never emitted
    2. The last mile: instrumented is not monitored
  9. The operator surface
    1. Serving it
    2. Identity comes from the request, never from its body
    3. Two gates, and the surface will not start without the second
    4. Checking a driver against the real thing
    5. Putting a tenant on your telemetry
    6. Per-tenant ceilings
    7. A tool’s rate ceiling
    8. The emergency stop
    9. One surface, many tenants
    10. What the endpoints are for
    11. No authenticator is shipped
    12. Claiming is what stops duplicated work
    13. Escalation widens the audience — and then leaves the expiry scan
    14. A task id mixes in the run it belongs to
  10. 🗄️ Retention and erasure
  11. 🧯 Disaster recovery
    1. The drill
    2. What an operator re-establishes
    3. Evidence
  12. 🚑 Runbook

Running this for real: topologies, the store contract, the background sweep, and what it reports about itself.


Ownership and fencing

Plane instances are stateless. Each run has at most one owner, held as a lease with an epoch.

Every append carries the writer’s epoch, and the store compares it inside the same transaction that writes. There is no window between “am I still the owner?” and the write for a paused instance to slip through, because there is no gap to slip into.

Two failure modes, deliberately distinct because they need opposite responses:

ErrorMeaningResponse
LeaseHeldSomeone else owns it and is aliveWait
FencedYour epoch is stale; you were taken overDrop the run. Never retry

Failover is not a special code path — it is the crash-recovery path. Lease expires, another instance claims at epoch + 1, resumes via replay. That is the payoff of building on replay: HA costs one lease table and an epoch column.

The claim is initiated by the sweep, and the subject matters. Fencing makes takeover safe and replay makes it correct, but neither makes it happen: every other resume has an event-shaped driver — an inbound message, a fired timer, an operator — so a run crashed mid-step with none of those pending would have no driver at all, appear in no backlog (it concluded nothing, and its wake was already consumed), and wait forever while looking exactly like work in progress. The sweep’s recovery pass is what closes that: an expired lease that still names an owner is precisely “an instance died holding this run”, because every clean exit — sealed, failed, suspended — releases. See the sweeper.

A live run renews; only a dead owner’s lease expires

Expiry answers is this owner dead? Without renewal it also answers a question nobody asked — a healthy run that outlives its TTL looks exactly like a crashed one, and agent runs routinely outlive a lease because one model call can. The run would be taken over and the original fenced mid-flight, having already done real work.

So the runtime renews while a run executes, at a third of the TTL, and stops the moment execution returns. Set the TTL with RuntimeBuilder::lease_ttl: it bounds how long a crashed owner strands its runs, not how long a run may take.

Anything under two seconds is refused at build. Both stores keep expiry in whole seconds and lapse on expires_at <= now, so a one-second lease is expired for part of every second it exists and no renewal frequency saves it — a plane configured that way would lose runs under load and nowhere else.

The owner string identifies a process, not an agent

This is the one piece of the mechanism you can defeat by accident. A lease is renewed without bumping the epoch when the claimant is the same owner — that is what lets a live instance keep its own run. So two instances sharing an owner string each read the other’s lease as their own, renew it, and both write under one epoch. Fencing is gone, and nothing reports an error.

Runtime::owner_id() therefore defaults to a value that is unique per process and per runtime instance. Override it with a real instance identity — a pod name — never with the agent’s name:

Runtime::builder(store).owner(std::env::var("POD_NAME")?).build()

Several instances of one agent are the normal way to run one, so the agent’s name is exactly the wrong choice. The agent is named by its manifest; the process is named here, and the two are different questions.

The owner lives in the lease table and never in the chain, so changing it has no bearing on replay.

Releasing frees the lease without forgetting the epoch

Every clean exit — sealed, failed, suspended — hands its lease back, so the next instance need not wait out the TTL. What a release must not do is delete the row: the epoch lives there, and without it append has nothing to fence against while the next acquire starts again at 1 — so a writer already fenced at 2 outranks the new owner and the mechanism inverts. Releasing marks the row expired instead, and the next takeover advances the epoch as any takeover does.

Releasing is something a run does, never something a shutdown does. The intuition runs the other way — hand the leases back on the way out and the next instance starts immediately — and it is wrong in a way the fence does not cover. A release says takeover is safe now. For a run still inside a tool call it is not: the next owner replays, finds the announced effect with no outcome, and performs the call a second time while the first is still in flight. The epoch bump stops the second append; nothing stops the second send. So an instance on its way out finishes what it can and leaves the rest to expire — which is the one signal the recovery sweep reads as an instance died holding this run, and is exactly what a run cut short by a shutdown is.

Stopping an instance

A process killed mid-step leaves an announced effect with no terminal record, and nothing in the journal can say whether that call reached the world. The effect’s declared recovery then decides — perform it again, or wait for a person — which is the right answer to an accident and an expensive one to schedule on every deploy. Draining is the difference between a stop somebody chose and a crash. Moving a plane to a new build starts with one → upgrading.

agentplane serve drains on SIGTERM and SIGINT. The signal does two things at once and then waits for four:

On the signalWaited for, under one grace period
Every listener stops acceptingRequests already being served, including the runs a blocking message/send is awaiting
Admission closes — Runtime::drainRuns this process put on a background task of its own
The sweep, the drill and push delivery, each finishing the tick it is in
Open subscriptions, which end at once rather than being waited out

Closing admission with the listeners rather than after them is the part worth stating. A request that arrives on an already-open connection after the signal is answered DRAINING and retries elsewhere, which is what that refusal is for — and it is what lets an open subscription end. A stream is a long poll over the journal, so a subscription to a run that sleeps for five working days would otherwise hold the graceful shutdown open for five working days.

--drain-secs (default 25) has to fit inside the supervisor’s own grace period, which is what sends SIGKILL afterwards — 30 seconds on Kubernetes and Docker unless raised. Runs still executing when it ends are named on stderr and at warn, and left to the recovery sweep on the path a crash takes.

In the published image the binary is the entrypoint, so it runs as PID 1. That is the right shape — the signal reaches the process that has to act on it, with no shell in between — and it is also why the handler is not optional: PID 1 has no default signal dispositions, so a SIGTERM with no handler registered is discarded and docker stop falls through to SIGKILL.

Two things a drain deliberately does not do. It does not gate resumes: a resume continues work this plane already owns, and refusing one would turn a shutting-down instance into a source of failed wakes for work the next instance is about to take. And it does not gate an agent commissioning another one — that is a step of a run the drain is itself waiting for.

An embedder driving its own server runs Runtime::drain(grace) beside their own graceful shutdown rather than after it — tokio::join! of the two — and reads DrainReport::unfinished. Serialising them either way deadlocks one on the other: a run awaited inside a request handler is finished by the server’s shutdown, and a subscription is ended by the drain.

What a caller is told while an instance drains

Admission answers RuntimeError::Draining — before the lease, the quota slot and the first append, so nothing is half-written. On the A2A surface that is -32031 with the ErrorInfo reason DRAINING, which is its own code beside QUOTA_EXHAUSTED and HALTED because it asks something neither of them does: a ceiling clears when a run finishes on this plane, a halt clears when a person lifts it, and this one clears the moment the caller reaches a different process. There is no window to wait out, so a compliant peer retries immediately.

The endpoint race is the deployment’s to close

Endpoint removal and SIGTERM are not sequenced against each other: a pod can receive the signal while requests are still being routed to it. No shutdown handler can serve a request that arrives after its listener is closed, so the fix is a preStop sleep covering the propagation delay, and a terminationGracePeriodSeconds large enough to hold the sleep and --drain-secs:

terminationGracePeriodSeconds: 60
lifecycle:
  preStop:
    exec: { command: ["sleep", "15"] }

There is no readiness route to flip, for the reason there is no /health at all — see the operator surface. A draining instance closes its listeners, so any probe that connects already fails.

Deploying a plane

Two artifacts run the same plane; both start from the files agentplane init --serve writes (getting started).

Compose (compose.yaml, written beside the others) runs the :full image of the version that wrote it against Postgres. The plane starts once Postgres answers its health check, as the user that owns the 0600 token file — never root, which init --serve refuses — with a read-only root filesystem; the manifest and policy are mounted read-only and the token file as a compose secret, never an environment variable. Every port is published on the host’s loopback only: A2A on 127.0.0.1:8080, MCP on 127.0.0.1:8081, the operator API on 127.0.0.1:9090. To serve other machines, publish 8080 and 8081 on an address they reach, set --url to the public A2A endpoint, and add a --mcp-allowed-host line for every name callers use to reach the MCP port. Postgres requires a password init --serve generated (postgres.password, 0600, read through POSTGRES_PASSWORD_FILE); the plane reads its connection string from store.env (0600) as AGENTPLANE_STORE, so no secret is on a command line. The network Postgres sits on is internal, but on Linux the host has an address on every bridge network, so the password — not the network — is what keeps a local process out.

The Helm chart (deploy/helm/agentplane/, not published to a chart repository) takes the manifest and policy as --set-file values, the token file only as an existing Secret, and a Postgres connection string as an existing Secret too (storeSecret.existingSecret), read into the pod’s environment rather than its args; a store value holding a password is refused. The pod runs non-root and read-only with every capability dropped; A2A and MCP share one Service and the operator listener has its own ClusterIP Service. Rendering fails for more than one replica without a Postgres store, since a redb file admits one writer, and the grace period is the drain plus ten seconds. With more than one replica the shared Service sets sessionAffinity: ClientIP, because a 2025-11-25 MCP session lives in the memory of the pod that opened it and any other pod answers it 404; behind an ingress that hides client addresses, route on the Mcp-Session-Id header instead. A changed manifest or policy rolls the Deployment: the chart reconciles nothing, because an upgrade is a rollout. serve reads the token file and the store once, at start, and never reloads them: a rotated Secret rolls the pod on the next helm upgrade (the chart hashes both Secrets as the cluster holds them), and otherwise needs kubectl rollout restart.

Two backends, one contract

JournalStore states three guarantees and requires them atomically: fencing, exactly-once, chaining. They are storage invariants deliberately — application logic can be bypassed by the next caller, a constraint cannot.

A second backend is where that stops being true, and the mechanism is worth being precise about. The new store is written from the same prose as the first. It encodes two guarantees exactly and something nearly like the third. Nothing catches it, because the suite that proves the runtime correct runs against the embedded store, and the new one gets whatever tests its author wrote — which are the tests for the parts they were already thinking about. The invariant they misread is by construction the one with no test.

So the contract is written once, in testkit::conformance, and every backend is run against the same battery. It ships rather than living in tests/ because an embedder bringing their own store needs it for the same reason.

Two design choices in the battery:

  • It reports every violation, not the first. Bringing up a backend is iterative, and stopping at the first failure hides whether the second is a separate bug or a consequence.
  • It fails if it checked nothing. A battery that silently runs zero checks reports success, which is the worst outcome available to it.

At-most-once admission is part of the contract

A RunAdmitted record carrying an idempotency key claims (tenant, key) in the same transaction as the append, and a second one under an issued key is refused with DuplicateAdmission naming the holder. The battery checks the refusal, that the named holder is the run that actually won, that the key is free before the append and taken after it, that a rejected batch spends no key, that retirement frees a key without touching the run it named, and that an unkeyed admission claims nothing — the direction that would otherwise break every plane never using the feature.

It also checks that a live lease is not stealable: LeaseHeld, distinct from Fenced, because a fenced writer must drop the run and a writer refused a live lease should wait.

The case layer

Five more stores, each settling one race:

StoreThe race
casestwo messages, one new matter; and two runs, one case state
eventsone message, two waiters
timersone wake-up, two sweeps
tasksone decision, two reviewers
batchesone item, two reservations

Each has a battery in testkit::conformance_case, run against both backends. The container tag is pinned in the test rather than inherited: the testcontainers-modules default is postgres:11-alpine, which has been end of life since November 2023, so the default would certify this backend against a release nobody should be running.

Postgres settles several more cleanly than the embedded store: UPDATE … RETURNING collapses a read-then-write into one statement, so there is no window to reason about because there is no second statement.

A paged read is part of the contract

JournalStore::recent_runs(after, limit) is the discovery index behind A2A task listing, and the battery checks three things a paged read gets silently wrong: the order is (updated_at, run) descending and total, the limit is honoured, and paging through it in twos reassembles the whole index exactly — no run served twice, none skipped.

The tie-break on run id is contract rather than detail. Both backends keep whole-second timestamps, so runs written back to back share one; without a tie-break the store may order them differently between two calls, and a cursor landing inside the tie drops or duplicates whichever moved. From a single page both look like a healthy listing.

The page boundary is also what makes the method checkable: a signature with no boundary has no boundary to get wrong, so a battery can pin nothing about it — and a listing that returns everything makes its one caller read every run’s complete journal on every request.

A sequential test cannot detect a race

Sequential checks prove the result is right, not that it is right for the right reason — a SELECT then INSERT returns the correct answer every time it is called one at a time.

So correlation also has a racing check, and it races hard: two concurrent callers serialise often enough that dropping the constraint which arbitrates correlation goes undetected. Eight racers across four keys catches it on every run, and mutation-testing the battery is what proves the check can fail.

  • A race test that does not reliably race reports green — an untested guarantee wearing a test’s clothing.
  • A race check corroborates, it does not prove. Passing means no interleaving found a violation; the constraint in the store is what makes absence real.

A store that serialises internally — redb admits one writer at a time — passes trivially and correctly, having no race to lose. That is not a reason to skip it: the check exists for the backend where the race is real.

A comment claiming a count-and-insert serialises “inside the row lock the write takes” reads as sound, and no such lock exists for inserts of different rows. So the guards suite races every store-side concurrency claim against a real PostgreSQL: quota admission, a tool’s rate ceiling across two planes, authority draws, task claims, timer sweeps, case correlation and case-state writes. A claim about concurrency that has only been read is a claim; raced, it is evidence.

Postgres

Three traps, each a plausible way to be nearly right:

  • Exactly-once is a partial unique index. A SELECT then INSERT has a window with two writers, and closing that window is the entire point.
  • Fencing reads the lease FOR UPDATE inside the appending transaction. Checking the epoch first and appending second re-opens the gap a paused instance wakes up into.
  • seq comes from the run’s own chain, never a sequence. Postgres sequences are non-transactional and leave gaps; a gap is indistinguishable from a deleted record during verification.

Each is held by a mutation that weakens the store and requires the battery to name the right invariant.

What the gate costs

The first question a runtime built on “the journal is the plan of record” invites. just perf answers it; these figures are from an Apple-silicon laptop and are quoted with the command precisely because they are hardware-specific.

Each control is measured on its own store and reported by its best of three runs, as the delta it adds to a bare effect. spread is what the method resolves — the gap between identical baseline runs — and a delta smaller than it is the machine rather than the control:

2000 effects × 3 runs, redb in memory, 3 policy rules
  journal     0.094 ms/effect      10586 effects/sec   canonicalize, chain, commit
  policy      0.120 ms/effect     +0.026 ms vs journal  one authorize per effect
  sink        0.098 ms/effect       under the spread  label gate, one protected field
  replay      0.003 ms/effect     291838 effects/sec   the read path, nothing performed
  spread      0.011 ms/effect                     across 3 baseline runs, what this resolves

400 effects × 3 runs, redb on disk, 3 policy rules
  journal     8.471 ms/effect        118 effects/sec   canonicalize, chain, commit
  policy      8.152 ms/effect       under the spread  one authorize per effect
  sink        8.137 ms/effect       under the spread  label gate, one protected field
  replay      0.004 ms/effect     250646 effects/sec   the read path, nothing performed
  spread      0.758 ms/effect                     across 3 baseline runs, what this resolves

One axis dominates, and it is not the one people ask about

The durability point is the whole figure. Authorization costs about 26 µs per effect at three rules, and the label gate over a protected field does not rise above what three identical runs disagree by. Against 8.5 ms of fsync neither is resolvable at all — the gate’s own checks are two orders of magnitude below the commit that records them. An adopter tuning this tunes the store, not the policy set.

Two consequences worth stating. Policy evaluation scales with the rule set, so a plane running forty rules pays proportionally more than the three measured here — still nowhere near the commit, but it is the axis that grows. And a control that reads as free here is not free everywhere: these are the costs of checking, and what a refusal costs the work it stops is a different measurement, taken against a deployment’s own bundle rather than against this loop.

Read the gap, not the numbers

Two fsyncs per effect. An effect crosses the protocol twice — EffectStarted before dispatch, its terminal record after — and both are durable commits at Durability::Immediate, because intent precedes action: the announcement must survive the process before anything reaches the world. The ~8.5 ms is that guarantee’s price, not an inefficiency waiting to be tuned away. Batching the two would break the only thing standing between a crash and an unrecorded payment.

Against what an effect normally is, this is noise. A model call is seconds; a tool call is tens of milliseconds. At 8.5 ms, journaling is well under 1 % of a real agent’s wall clock, and the whole design is priced for effects that reach the world.

It stops being noise when effects are cheap and many. cx.now(), a case read, a memory recall — a plan doing hundreds of those pays 8.5 ms each. If a step is looping over cheap effects, that loop is the cost.

On redb it is a plane-wide ceiling, not a per-run one. redb has a single writer, which is exactly what makes its exactly-once key and its fencing free of races — and it means ~118 effects/sec is the whole plane’s durable write budget on this hardware, not one run’s. A single-node plane running agents that call models will never notice. A plane running many concurrent runs of cheap effects will, and that is the signal to move to PostgreSQL, which is also the answer for more than one instance. It is a flag, not a rewrite: every verb takes --store postgres://… beside --tenant, including serve, in a build with postgres → which verb needs which feature.

Replay is not the same cost, by a wide margin. It performs nothing and reads history back — 2000× faster on disk. That matters more than it looks: a divergence check, a crash recovery and an offline audit all pay the read price, so the expensive half is the half that only happens once.

PostgreSQL is deliberately not measured here. It is a different machine, a different fsync, and usually a network hop; quoting a number produced on a laptop’s container would be worse than quoting none.

The sweeper

Until something runs on a clock, a deadline is a number in a table and an unclaimed event is a row nobody reads. That is the failure this runtime is built against — not a crash, but a silence.

One tick, several findings — not all of them alarms, and the first two routine healing that still means something died. Recovery runs first, because an abandoned run may be one step from meeting a deadline the passes below would otherwise breach; then deadlines, task windows, event redelivery and dead-lettering, and finally timers:

FindingWhat happens
An instance died holding a runThe run is taken over at epoch + 1 and resumed
A claimed event’s delivery diedThe delivery is finished; reported as events_redelivered
An obligation is approachingDeadlineTransition → Warned
An obligation passed unmetDeadlineTransition → Breached; the case is escalated — unless a run met the obligation since the sweep read it, which wins
A task’s window closedThe declared on_expiry is applied
An event nobody claimed aged outDead-lettered with a reason
A sleeping run’s instant arrivedThe run is woken; reported as timers_fired
A run with no lease — a lift, a release, a break-glass, an observed session, a sweep’s own — concluded but its seal diedThe run is sealed; reported as seals_finished. Nothing it did is retried — halt list and hold list say what stands
A wake was recorded but the resume diedReported as wake_failures; the lease lapses and the recovery pass picks the run up on a later tick
The log grew since the last anchored checkpointIt is submitted to the deployment’s witnesses; reported as cosignatures, witness_shortfall and witness_integrity

now is passed in rather than read, so the caller controls the clock. That keeps the sweeper testable at all, and lets a simulation drive a year of obligations through in milliseconds.

Not every field of SweepReport is an alarm, and not all of them are numbers. timers_fired is the system working — its own documentation says not to alert on it — and runs_recovered is routine healing whose real message is that an instance died. saturated is a set of flags saying which passes came back full and so may not have seen everything that was waiting; record names the sealed run holding the tick’s own evidence, when the tick did anything. Two fields sit apart from the counters. evidence_lost is the most serious thing the report can carry: the tick decided something — obligations breached, cases escalated — and the durable, tamper-evident account of who decided that and when could not be written. census_unavailable means the gauges could not be read this tick, so the census in the report is a default rather than a reading — a blind spot wearing a zero.

The witness pass is the one phase whose counterparty is somebody else’s server, which is why it runs last: a slow witness must not delay a breach or a recovery. cosignatures is evidence accumulating — the system working, like timers_fired. witness_shortfall is a plane running with fewer independent parties vouching for its history than the deployment declared it required, and nothing else will clear it. witness_integrity is a witness refusing on integrity grounds — the log shrank, forked, or claimed growth the witness could not verify — and it is reported even when the quorum was met, because two honest cosigners do not answer a third that remembers a different history. It is also the one sweep finding that goes to tracing as well as to the report: the audience for a witness says this history moved is not only the operator running the plane.

needs_attention() is the alerting predicate, and it enumerates exactly what a human should see: breached, tasks_expired or dead_lettered above zero, any saturated pass, recovery_failures or wake_failures above zero, evidence_lost, census_unavailable, and either witness count above zero. Recoveries that succeeded stay off that list — they are the plane healing, findable in the report and in the sweep’s own run. is_quiet() is the broader predicate: a healthy plane sweeps silently, so a non-silent sweep means something happened, even when it was only the system working.

The recovery pass: who resumes a crashed run

The candidate set is exact, not heuristic. Every clean exit hands its lease back — sealed, failed, suspended — so JournalStore::abandoned_runs answers one precise question: which leases expired while still naming an owner. That is the set of runs somebody was executing when their process stopped, and the sweep resumes each one: takeover bumps the epoch, the store fences the dead owner’s next append, and replay reads completed effects back rather than redoing them. A run whose last record is already its conclusion died between concluding and handing the lease back; the sweep finishes that — seals it or releases the lease — and never retries the work, whatever the outcome was. cargo run --example recovered_run walks the whole thing on two in-process instances: one dies mid-run, the other’s sweep finds it by its lapsed lease and finishes it, no stage repeated.

Three details are worth knowing at 3 a.m.:

  • runs_recovered above zero means an instance died, even though the runs themselves are fine. The healing is routine; the dying is not. A steady recovery rate with allegedly healthy instances is a contradiction — go look at why leases are lapsing (GC stalls and CPU starvation produce exactly this).
  • recovery_failures is one stuck run, not many. The run stays listed and is retried every tick, so a persistent count is the same run failing repeatedly — and nothing else will unstick it. The reason is in the log line of the failing tick. This is the one recovery number needs_attention() fires on. A recovery refused by the journal itself — a record this build cannot decode, a chain that does not verify, a canonicalization it does not implement, a recorded chain it cannot rehydrate — would be refused the same way on every tick, so it is not retried: the run is quarantined with the refusal named, and lands on the quarantine backlog for a person to reopen or abandon. Everything else, a plane-dependent refusal included, is retried.
  • The batch is small (32) on purpose. Recovering a run replays its journal and then executes live from the frontier — which may dispatch a model call — so a mass failure drains over several ticks with saturated.recovery up rather than holding one tick hostage. The flag means at least a batch was waiting, never that the batch was all there was.

Each takeover is written into the sweep’s own sealed run as run_recovered, because a takeover fences the previous owner and who fenced whom, and why must be answerable from the journal rather than inferred from an epoch gap. The note lands before the resume: a concluded resume releases the lease and leaves the recovery queue, so an account written afterwards would be the one write a crash could lose with no retry ever selecting the run again. A note that cannot be written skips the takeover for a tick that can write it, and the outcome stays out of the note — the run’s own journal answers it.

One state recovers to nothing, deliberately: a lease over an empty journal means admission acquired and died before its first append landed. No run exists — the atomic admission batch never committed, so nothing was declared, authorized or performed — and clearing the lease is the whole recovery.

The claimed-event window is covered by its own pass. A delivery can die between an event’s claim and the resume that consumes it — a crash in that window, or an owner that outlives the delivery’s bounded retry — and the message then belongs to nobody: claimed, so no longer waiting; undelivered, so no run holds it. The delivery parks the pair, and the sweep walks the parked pairs — never every registered wait, whose long legitimate members would fill its page first — finishes those deliveries and reports each under events_redelivered. A parked wait already holds its message, so no second message is matched to it while it waits for redelivery; a parked wait whose run has since been sealed is retired rather than tried, as is every wait a run held when a resume repairs the seal a crash left missing. The redelivery is the system healing; a persistent count means deliveries keep dying, which is worth asking why.

A finding has to be findable

Every conclusion this runtime reaches is queryable by whoever must clear it, without them already knowing which run, case or matter to open. A control that notices and does not deliver is closer to none than to half, because it also manufactures the belief that somebody was told.

Each backlog is a question, not an id:

curl -H "$AUTH" 'https://plane/runs?outcome=quarantined'   # history you cannot trust
curl -H "$AUTH" 'https://plane/cases?status=escalated'     # matters somebody must pick up
curl -H "$AUTH" 'https://plane/obligations'                # windows we missed
curl -H "$AUTH" 'https://plane/tasks'                      # decisions waiting on me
curl -H "$AUTH" 'https://plane/dead-letters'               # messages that reached nobody
curl -H "$AUTH" 'https://plane/push'                       # receivers that stopped accepting

All six page the same way: truncated says whether there is more, and the order puts the item you are most likely to want first — newest for runs and cases, longest-overdue for obligations. Ascending order is only safe on a listing that drains. A bounded query taken oldest-first is a page that stops changing: if nothing ever leaves it, a plane whose backlog exceeds one page returns the same rows forever and the thing that just happened is the one that never appears. Obligations are ordered that way because acknowledging a breach removes it — the head of that page is not permanent, and POST /obligations/acknowledge is what moves it.

/runs reads an index derived from the RunConcluded record inside append, in the same transaction, so it rebuilds from the chain and is never an authority. The last conclusion wins — a failed run moves to succeeded when a resume concludes it, so the failed backlog drains. Failure does not seal: a failed, exhausted or quarantined run stays open, and only succeeded, cancelled and abandoned freeze the journal. A quarantine is a pause rather than an ending — the runtime is saying it does not know — and freezing its chain would lock out the one record that answers it.

The single-run view’s sealed field is backed by Merkle inclusion, not by the presence of RunConcluded. That distinction is observable: a failed run has a conclusion and reports sealed: false, while a succeeded run reports true. Exhaustion remains structured in the journal and in operator push events, so automation can inspect the exact ceiling without parsing reason.

/cases defaults to escalated. An unrecognised status is a 400 rather than a quiet fallback — answering what is escalated with a list of healthy cases reads as an empty backlog, which is the most reassuring possible way to be wrong.

/obligations takes no status. A breach is the only obligation state anybody has to be told about; the rest are either still watched by the sweep or already answered. It reads the obligation’s own row rather than the case’s status, so it still answers after the matter is closed — closure is when people stop looking, which is when the record has to stand on its own.

It lists breaches nobody has accounted for, and that qualifier is what makes it a backlog rather than a ledger. POST /obligations/acknowledge names the case and the obligation and records who looked, taken from the authenticated caller rather than the body. The obligation stays Breached — what ends is the question, not the fact — and the account is readable afterwards on GET /cases/{case}. The call is idempotent and says which happened: recorded is false when somebody had already answered, because the first account stands and a retry must not rewrite who looked or when.

Alert on agentplane.obligations.breached, which is the same figure as a level rather than a page.

/dead-letters and /push are the two backlogs that are diagnoses rather than work items, and both carry less than the store holds. A dead letter means a correlation key does not match what a run subscribed to, so the listing gives the identity and the keys and not the message body — that is the counterparty’s content, and on a sealed plane it is not even readable without a key. A parked registration gives its config redacted, because the token a receiver registered is its own correlation secret and a listing is not where it is handed back. POST /push/rearm answers rearmed: false when nothing was parked under that name — already live, or never there — rather than reporting success to somebody who would then wait for a sweep with nothing to do.

Grant the read verbs explicitly. api:run.list, api:run.live, api:case.list, api:obligation.list, api:deadletter.list, api:push.list and api:halt.list are what an on-call person needs, plus api:obligation.acknowledge for whoever answers a breach — separate from reading the list, because the party allowed to see what a deployment missed is not automatically the party allowed to declare it answered. An allowlist built from route names alone will miss them; under a default-deny engine an ungranted verb means the backlog is refused to everybody. Enumerate action::ALL when writing rules rather than reading the route table. api:obligation.list is separate from api:case.list so a compliance function can be given what did we miss without the contents of every matter, and api:task.takeover is separate from api:task.claim so displacing an absent colleague can go to a queue lead without going to every reviewer. api:run.live is separate from api:run.list on the same principle: one enumerates what has finished, the other is a live map of what a tenant is doing right now.

Answering a quarantine

A quarantine is the runtime saying it does not know. An effect was announced, the process died or the provider went quiet, and whether the call reached the world is unanswerable from the journal — so nothing unwinds, because compensating around an unknown outcome is a refund for money nobody took.

A run is also quarantined when a resume is refused because the declaration or policy bundle this plane holds is not the one the run was admitted under; its reason names both digests, and it is answered by bringing back the recorded revision and reopening, or by abandoning. Either way the run stays listed by attention until somebody answers it.

That is the one conclusion a resume cannot clear, and the only backlog whose level moves because a person moved it. Three verbs, and they are deliberately three: supplying a fact is not deciding a run, and deciding a run is not declaring its outcome.

1 — Find out what is in doubt. The run view names the call, not just the situation:

curl -H "$AUTH" https://plane/runs/$RUN
{
  "run": "run_01K…", "status": "quarantined",
  "reason": "the provider could not establish whether the call landed",
  "undecided": [
    { "effect": "9f2c…", "step": 1, "phase": "forward",
      "kind": "tool://payments/charge", "doubt": "inconclusive" }
  ]
}

doubt says where to start looking. announced means the runtime never heard back — the crash shape, nothing known beyond the fact that the request left. inconclusive means something did report and could not tell, so there is a message to read and a provider that has already been asked once.

2 — Say what actually happened. Look the call up in the system that would know, and record the answer:

curl -H "$AUTH" -X POST https://plane/runs/$RUN/reconcile -d '{
  "effect": "9f2c…",
  "disposition": "landed",
  "output": { "charge_id": "ch_9RtQ", "captured": true },
  "note": "charge ch_9RtQ exists in the provider console, created 12:41Z"
}'

did_not_happen needs no output and leaves the effect safe to perform again. landed carries the result the run reads back. Three rules are worth knowing before you use it:

  • It supplies a missing fact and never replaces a recorded one. Only an effect listed under undecided may be answered. Anything else is a 409 — otherwise an operator could talk a run out of compensating work that is standing in the world, and the journal would show an orderly reconciliation while it happened.
  • Your name is on it. The record is the same EffectReconciled a probe writes, plus asserted_by. “The provider told us” and “somebody asserted it” are different evidence and the chain keeps them apart.
  • The value you supply is untrusted. Every other output in this runtime is labelled by the effect that produced it; there is no effect here, and you are not the provider that returned it. So it takes the conservative point of the lattice, and a run that needed that output to be trusted will be refused at its next gate and unwind. That is the honest ending — the alternative would make this the one place in the design where a person declassifies by typing.

3 — Hand it back, or write it off.

curl -H "$AUTH" -X POST https://plane/runs/$RUN/reopen  -d '{"reason":"charge confirmed in provider ledger"}'
curl -H "$AUTH" -X POST https://plane/runs/$RUN/abandon -d '{"reason":"two weeks of provider tickets; nobody can say"}'

reopen does not declare an outcome. It hands the run back to the runtime, which re-derives its verdict from a history that now holds what you established — and quarantines it again, on the record, if the answer was wrong or incomplete. The response is what the run actually reached, so you learn immediately rather than polling. One decision authorizes one pass: a run that quarantines again needs a fresh one.

abandon closes it where it stands. Nothing is unwound — the world keeps whatever the run left in it, because unwinding around an unknown outcome is exactly what quarantine exists to forbid and impatience is not evidence. The run seals as abandoned and leaves the quarantine backlog.

The doubt does not leave with it. Abandoning takes the run off the only listing that carried it, so what it left standing becomes an agentplane audit finding derived from the journal, where no later action can clear it:

run run_01K… is sealed but effect 9f2c… (step 1, inconclusive) never reached a
known outcome — nothing may resume a sealed run, so whether the call changed the
outside world is permanently undecided

Cancelling a quarantined run is refused, not recorded. cancel promises to unwind and put the world back, which is the one thing a run holding an unknown outcome may not do; the refusal names these two verbs instead.

Three verbs, three policy actions: api:effect.reconcile, api:run.reopen, api:run.abandon. Keep them apart in your rules. The person who can look a charge up in a provider console is often not the person who decides what the run does next, and neither of them is necessarily the person who may write off an unexplained mutation.

Retention for the admission index

Admission keys are kept until you retire them:

agentplane forget-admissions --store ./journal.redb --older-than-days 30

The window has no default, and that is the decision. Retiring a key reopens the door it closed, so a window shorter than your emitter’s retry horizon admits a second run on a timer. Runtimes that bound this automatically have to guess at that horizon; nothing here can know it.

The index is a row per admitted message and grows with inbound volume — a size your database monitoring already reports. The verb prints what it retired, so a pass that found nothing is distinguishable from one that said nothing.

Taking the record away

Two verbs read a journal and nothing else — no manifest, no source tree, no Rust toolchain — because that is what an auditor or a departing tenant holds:

agentplane export --store ./journal.redb > history.jsonl
agentplane audit  --store ./journal.redb > report.json
agentplane verify history.jsonl --checkpoint cp.note   # check a copy, offline
agentplane restore history.jsonl --store ./rebuilt.redb
agentplane policy check --bundle ./policy --from history.jsonl   # re-derive the verdicts

--checkpoint is the deletion check, and without it there is none. The Merkle root rebuilt from the file can otherwise only be compared with the file’s own header — which whoever dropped a run rewrites too — so the report lists deletion under not_checked rather than calling the file sound. Pass the checkpoint an earlier audit printed, or the tlog-checkpoint note a witness cosigned: the point is that it comes from somewhere other than the file being checked.

export writes JSON Lines: a header naming the log, its checkpoint and the canonicalization rule the digests were computed under; one line per record carrying prev_hash and hash, so the chain re-walks from the file alone; one line per case — the case layer is beside the journal, not derivable from it, so a file without it rebuilds a journal whose records name matters that no longer exist; and a trailer. The trailer’s absence is the signal — a file cut short by a full disk or a killed pipe ends without one, and every line in it is still valid JSON, so counting is no help to a reader who does not have the source. Runs that could not be read are named in the trailer rather than quietly missing, and the trailer’s case count is what catches the case layer stripped whole.

Case state travels as stored: on a sealed plane that is ciphertext, because an export of plaintext would quietly undo erasure — the key destroyed tomorrow would no longer reach the copy taken today. Two questions stay live rather than offline, and verify reports them as unchecked instead of passed: whether the blob bytes behind the exported digests are present, and whether sealed state’s keys still unwrap.

Those two are Runtime::drill’s job, run with the stores the plane actually runs with: every case’s blob digests read through that case’s own handle — the unit-scoped address the plane wrote to, opened through the sealing envelope when a ring is wired — then re-hashed (get, never has — presence without integrity passes over altered bytes), and sealed state proven to open with the plaintext dropped on the spot. Reading any other way would hold the references against a store the deployment does not use: on a sealed plane, every intact envelope would report as corrupt, the one verdict that pages. The report’s verdict has four answers, and the second is the one worth trusting the tooling for: intact; erased by design — a tombstone or a destroyed key is retention reporting itself, counted and never a finding; lost, which pages; and the bytes are gone and their tombstone does not read, which pages too — the erasure may well have run, and nothing left in the store can say so. A drill that alarmed on erasure would teach you its findings are noise, and that is how a real loss gets ignored six months later.

The same rehearsal has a CLI verb for deployments that never write Rust: agentplane drill opens the store the flags name, prints the report as JSON, and exits non-zero only on loss — erased-by-design counts stay informative. A store file holds no blob backend and no key ring, so those halves land in the report’s unchecked list rather than being silently passed; the library call on the running plane remains the complete form.

The report also carries releases and warrants — every point at which a label was raised, and what authorized each run. The gate on both is integrity, not innocence: a run whose chain did not verify shows neither, because nothing drawn from those records can be trusted, and a run that verified and then failed a check shows both, because that is the run an investigator opened the report for.

Read warrants with unadmitted, its complement: every run that verified is in exactly one of the two. A run with no admission record produces no warrant, and reported alone the warrant list is shorter than the run list with nothing saying why — so the runs that have no answer are listed with the outcome they concluded under. Some are unadmitted by design: the sweeper opens a run of its own for decisions it takes without a request. What the pair buys is that what authorized this run is a question the report answers for all of them, including with “nothing did”.

audit prints the report as JSON and exits non-zero on findings — but not on not_checked, which is a separate list and the one worth reading. An audit given no public key and no earlier checkpoint still walks every chain, and says in that list that it could establish neither authorship nor deletion. Supply them to narrow it:

agentplane audit --store ./journal.redb \
  --key plane-1=<64 hex chars> \        # the signer's Ed25519 public key
  --prior last-report.json              # the `current` field of an earlier report

--key is the authorship check, repeatable per trusted signer; add --require-signatures to make an unsigned record a failure rather than a note — off by default, because history written before signing was configured is legitimately unsigned. --prior is the deletion check: the report’s own current field, saved from an earlier pass, and a log that shrank or forked since then is a finding. The loop is deliberate — each audit prints the checkpoint the next one checks against.

Both verbs default to every sealed outcome. --outcome narrows and --limit bounds, and reaching the limit is never a pass. An audit records it in the report’s truncated field — {"limit": …, "reached": [outcomes]}, so it survives > report.json — and exits 5 when nothing it read was wrong. An export that reached it is refused before a byte is written, because a truncated file is framed exactly like a complete one: raise --limit, narrow --outcome, or pass --allow-partial to write it anyway, which still exits 5. See exit statuses.

Disclosing one matter

A request for one matter is answered with a package, not the whole export:

agentplane export --store ./journal.redb --case case_<ulid> \
  --to "Supervisory authority, ref 2026-114" --actor dpo@example > matter.jsonl
agentplane disclosures --store ./journal.redb --case case_<ulid>

The package carries the runs the case holds now (--run adds single runs), each sealed one with its inclusion path, and only the cases those runs belong to; the format page has its shape and how both readers verify it. The recipient checks it against a checkpoint obtained from you some other way: at the package’s size it is compared by root, at another size it is reported not compared. restore, replay --strict --from, policy check and grants refuse a package by name.

Before a byte reaches the destination the disclosure is recorded — recipient, runs, sealed or not, checkpoint, package digest, actor — so a later erasure names the copy. If the record cannot be written nothing is delivered. An unknown case or run, an --output that is a directory, or one whose directory does not exist exits 2 with nothing recorded. The package is staged as .<name>.dsc_<ulid>.partial beside --output, synced, and renamed into place once recorded; a process killed in between leaves that file, which nothing reads and you may delete. Once recorded the act stays recorded: a package written to a standard output that closes early is reported as recorded and not delivered.

Each case travels as its whole block — state, deadlines, blob digests, hold reason, and the ids of every run it holds, including runs the package does not carry. A package proves the inclusion of what it carries, never that it carries everything the matter holds. The register is a row in the store, like a hold: disclosures lists it and says so.

Re-deriving the policy verdicts

policy check reads an export and a policy bundle — the same --bundle loader serve --policy uses: one .cedar file, or a directory holding policy.cedar and optionally schema.json and entities.json, and nothing else — and evaluates every gated request the records support, offline. It opens no store and writes nothing. It needs the cedar feature (the :full image has it).

agentplane policy check --bundle ./policy --from history.jsonl --tenant acme
agentplane policy check --bundle ./policy --candidate ./policy-next --from history.jsonl --json

Each run is recorded (the bundle’s identity is the one its admission names, so it was evaluated), a mismatch (another bundle governed it; not evaluated), or ungoverned. A finding is a recorded permit the bundle refuses. --candidate adds, per run, the recorded permits the candidate would newly deny or cannot evaluate — measured against what happened, so a mismatched run is measured too. The report lists what it could not rebuild and why (sealed, erased, a refusal’s request, a gate the record cannot show was passed); security has the list. --tenant is needed on a multi-tenant plane: every request carries the tenant and no record does, so the report says whether it was supplied.

It exits 1 on a finding, a mismatch or a non-empty candidate diff, 5 when it evaluated nothing, and 0 otherwise. It checks agreement between a bundle and the record, not the record’s integrity: verify the file first.

Replaying an edited declaration

replay --strict re-executes a recorded run under the manifest it is handed and serves every model completion, tool result and peer reply from the journal. Handed an edited manifest, it answers the question an author has in review: would this change have made last week’s runs do anything differently, and where?

agentplane replay run_01M3… --manifest edited.yaml --strict --store runs.redb
agentplane replay --manifest edited.yaml --strict --from monday.jsonl --from tuesday.jsonl
run run_01M3… — diverged
  recorded: summariser 1.0.0 420d054e…
  candidate: summariser 1.0.0 9623cd59…
  (different digest)
  first divergence: step s0, forward phase, `model.complete` — history ek:024c…, this build ek:3dff…

Each run gets one verdict, and every verdict names both revisions — the declaration digest the run was admitted under and the one in hand:

VerdictMeans
verifiedEvery recorded effect was asked for again, in order, and the run reached the ending it recorded — a recorded failure included: the recorded ending is what is verified. Under a different digest it reads no divergence on this run, which says nothing about the next input
divergedThe first effect this build asked for differently, or beyond the record, or never asked for — with its step, phase and both keys. Or every effect matched and the run ended differently
cannot replayA reason that is not the edit, named: erased (sealed to a destroyed key), key absent (sealed, and this plane holds no key ring — nothing is known to be erased), canonicalization changed, entry point removed (the file no longer provides the capability the run executed), not concluded, or ended by an act (a cancellation, abandonment or sweep its steps did not decide)

It writes nothing — no lease, no conclusion, no index moves — and it reaches nothing: every provider, tool server and peer the manifest names is answered by a stand-in that refuses any call, so no provider credential is read and --mcp and --peer are refused rather than ignored. A model call’s identity includes its driver’s request profile (endpoint, schema mode, streaming), which is deployment wiring, so the stand-in takes it from the record: the verdict is about the declaration, not about this machine’s provider setup.

--from reads an export instead of a store: each file is rebuilt in memory, checked against its own checkpoint, and every run in it is replayed unless a run id names one. A corpus exits with its worst verdict — 4 if a run could not be read, then 1 if any diverged, then 5 if any could not be replayed, else 0 — so a CI job holding exports and a pull request’s manifest reads one status.

What it cannot see: a policy bundle edit (a strict replay replays recorded decisions and never asks the engine, which is what keeps a relaxed rule from re-judging old runs), and anything the declaration contributes outside the effects a run made — a raised ceiling diverges only on runs the old one stopped. A replay that serves some boundaries from the record and runs others live is refused, not offered: a live answer grafted onto old history is an observation the journal never recorded, and a model call made to explore an edit is a fresh run. A resume under a different declaration is still quarantined before anything replays.

One run’s timeline

agentplane history <run> --store <file|postgres://…> prints a run’s journal in sequence order, one line per record: sequence, kind, step, and the record’s payload. --from <seq> starts later. Every line passes through the escaping a reviewer’s rendering uses, so a bidirectional override or an escape sequence in a recorded string is printed as \u{…} rather than reaching the terminal. --json prints each record as GET /runs/{run}/history serves it; it is machine output, and JSON escapes only U+0000–U+001F, so DEL, the C1 controls and the bidirectional and invisible characters print as they are — read the text form on a terminal. A run the store does not hold exits 1, whatever --from names.

The command line’s exit statuses

One table for every verb, printed at the foot of agentplane --help, because a scheduler reads the status and nothing else:

StatusMeans
0Done, and the answer is yes.
1A finding or a negative answer: a failed run, an audit, verify, drill or policy check finding, a grant grants found unused, a strict replay that diverged, attention finding something, a content check a rule would refuse, a halt --lift / hold --lift / rearm that found nothing standing, a decide approval of a proposal the terminal cannot show or naming a --digest the task no longer holds, an invalid manifest under validate, a history of a run the store does not hold.
2Usage: the command as typed cannot be carried out — a flag, argument or input this binary refuses. clap’s own parse errors use it too.
3A run or a resuming replay stopped to wait, for a person, a timer or an event. A strict replay of a waiting run that reproduces the wait is verified, 0.
4Operational: a store, a witness, the network or a file could not be used.
5Partial: --limit truncated what was read — audit, export, tasks, waiting, halt list --lifted and hold list --released, each of which also prints truncated — policy check could evaluate nothing, grants met an incomplete export or calls it could not read, subject had its scan cut or met a run it could not read, or a strict replay met a run it could not replay (and none diverged).
6Unverifiable: verify or restore met an export under a canon this build does not implement — not damage, and not an outage; a build implementing that rule can read it.

A finding and an outage are different pages, which is the reason for the split.

Output. The verbs whose answer is a report — audit, verify, export, attention, tasks, waiting, drill and the other operator verbs — print JSON on stdout, always. The verbs whose answer is a sentence — halt, init, validate, digest — print text for a person, and one JSON document on stdout with --json; so does policy check, whose report is long enough that a person reads the summary and a pipeline reads the JSON. Everything else goes to stderr. serve --log-format json (AGENTPLANE_LOG_FORMAT) writes one JSON object per log line for a collector, and agentplane --version names the features the binary was built with.

Which verb needs which feature

--features cli builds every verb against a redb file, with every model provider and the fake one. What reaches further is one feature away, and a build without it refuses with exit 2 and names the flag to reinstall with. The :full image carries all of these; :slim carries only cli.

You runInstall with
any verb with --store postgres://…--features cli,postgres
run --mcp …--features cli,mcp-stdio
run --peer …--features cli,a2a
serve, and its --operator-addr--features cli,a2a-server,cedar
serve --mcp-addr, init --serve--features cli,a2a-server,cedar,mcp-server-http
policy check--features cli,cedar
rearma build with push, which a2a-server includes; the verb is absent otherwise
dev--features dev; no published image carries it, and the verb is absent otherwise

What no audit can answer: the runs that never started

A run refused before it exists — a policy denial on run:admit, a tenant ceiling, a standing halt — has no run id and so no chain to append to. Nothing about it is in the journal or in the report, so how often did policy stop a run from starting is a question a clean report answers with silence rather than with zero.

Manufacturing a run to record its own refusal would mean a chain, a seal and a Merkle leaf for something that never executed, and would let a caller grow the store by being denied. Every refusal after admission is journaled, which is why an effect-level denial has a PolicyDenied record and this one does not.

The number lives in the agentplane.policy.denials metric and agentplane.policy.denied telemetry, both carrying the action, with the durability of whatever collects them. It is deliberately not in not_checked: that list is what this audit could not check, and an entry in every report ever produced would train you to skip it.

Auditing open runs (--outcome failed, say) is not an alarm: an open run has no Merkle leaf, so it is checked on chain and signatures and the report says in not_checked that nothing pins its tail until it seals. The finding is the opposite case — a run whose own records carry a sealing conclusion, in a log that holds no leaf for it. That is history the log no longer commits to.

A sealed run is also held to its own transactional brackets: a GroupOpened with no GroupSettled under a sealing conclusion is a finding. In an open run that shape is the ordinary crash the resume repairs; under a seal nothing may resume, so whether the group’s members were taken or taken back is permanently undecided — a state no honest writer produces.

For a sealed run, the log’s leaf is held to the verified chain’s own head, before the tree math and independent of any checkpoint race. A truncated but internally consistent prefix of a sealed run verifies on its own, and the log’s genuine leaf verifies against the genuine tree — so an audit that checked each half without holding them to each other was verifying two halves of two different claims. A mismatch is a finding: records were removed or replaced after sealing, and the served history is not the one the checkpoint commits to.

verify is the drill, and it takes the file alone. It re-seals every record through the same function the store sealed with — so agreement is evidence about the bytes, not the file agreeing with itself — checks sequences are contiguous, holds every record to the run block it sits under (a relabelled block passes chain, leaf and Merkle checks, because those verify the bytes and only the label lied), then rebuilds the Merkle log from the positions each run block carries and compares the root against the checkpoint in the header. With --key it also verifies signatures, strictly: inside a signed history, an unsigned record is the one an attacker who cannot sign would add.

That last check is the one worth understanding. Delete a whole run from an export and every remaining chain still verifies perfectly: a chain links records within a run and knows nothing about its neighbours. Only the rebuilt tree notices the missing leaf. It is also the reason each run block carries its log position at all — without it an export is a transcript rather than evidence.

Exit codes: findings fail, not_checked does not. A pass with no public key has established less rather than failed, and the report says which.

Both verbs print one document, so a redirect cannot separate the verdict from the grounds. Beside the report is an anchor object:

FieldSays
obtained_froma witness’s prefix, or a file
cosigned_bythe witness keys whose signature over that checkpoint verified — for a file, those of a cosigned note checked under --witness-key; empty for any other file, which carries none to check
unreachedwitnesses asked that gave no anchor, and why
split_viewwitnesses that disagree, which fails the command

A clean report against a cosigned anchor and a clean report against a checkpoint somebody typed are different statements; the anchor is what distinguishes them. held_to on the report names the checkpoint itself — the anchor says how it was obtained, and never repeats the value.

Name a second --witness. Two witnesses holding one tree size with two different roots is the event witnessing exists to detect, and the one an operator auditing their own plane cannot find alone: every other check compares the store against something the store produced. Witnesses at different sizes are not a split view — they observed at different times.

How fresh the anchor is. A witness signs a timestamp with every cosignature. audit --max-checkpoint-age SECS judges each witness key’s latest signed time against the auditor’s own clock: older than the bound — or ahead of it by more — is a finding; without the flag, freshness is listed as not checked, and verify never judges it. A gap bounds when records could have been cut without a witness noticing, not whether they were. A plane keeps the times fresh by declaring serve --witness-submit URL --witness-key NAME=KEY --log-key NAME=PATH --witness-interval SECS: the sweeper re-submits an unchanged checkpoint once the interval passes, an interval shorter than --sweep-every is refused, and one with no witness is refused at build. The second reader checks the same cosignatures on a note anchor under its own --witness-key NAME=BASE64, and judges freshness under --max-checkpoint-age SECS (with --now RFC3339 for a fixed clock) by the same rule; a time whose cosignature did not verify is an unauthenticated number, and neither reader judges it.

A grader’s verdict, bound to what it judged. agentplane bind <export> --run R [--last-seq N] --content FILE --out SIDECAR writes an unsigned sidecar binding opaque verdict bytes to records 1..=N of a run; the grader signs it with its own key. verify --grader-verdict SIDECAR --grader-key ID=HEX reports each one bound, refused or not checked, and a refused one fails the command. It proves which records the verdict names and who signed it — not what the grader saw.

restore rebuilds a journal from an export and proves it by one comparison: equal Merkle roots at equal size. That is a far stronger statement than “the rows loaded” — it means every record, in every sealed run, in the order the log recorded them, rebuilt to the same commitment. A run restored this way strict-replays on a plane that never executed it.

Sealed is the load-bearing word, and it decides what to export. The log commits to runs that ended, so a file carrying no in-flight run rebuilds to exactly the same root at exactly the same size as one carrying all of them — a clean verify says nothing either way. And the runs still in flight are the ones no outcome index names: runs_by_outcome indexes conclusions, so a sleeping run, a run awaiting a message and a run waiting on a person are in none of them. agentplane export asks for them separately and says on stderr how many it carried; export::runs_in_flight is the same selection for an embedder. Narrowing with --outcome turns it off, because naming outcomes is a request for exactly those conclusions.

It rebuilds the case layer beside the journal: every matter is queryable again — case, correlation, the status worklist, the deadline sweep, the blob links erasure walks, and the legal holds that stop a retention pass from erasing a matter — and the conformance battery holds the import to every one of those read paths on both backends. A store that already holds a case refuses it: a restore rebuilds a case layer, it does not merge one.

Before anything is written, each record line must be replayable as written: its hash must cover its raw bytes and the chain before it, and its v must be the version this build writes. A file failing either is refused whole, with the store untouched — append re-derives every hash from what it is handed and stamps this build’s version, so replaying such a line would rebuild a history the export never committed to. A target that seals payloads as it writes (a journal wrapped by the keyring) is refused the same way: sealed payloads are restored as the ciphertext they are, so restore into the unwrapped store and open it with the ring. A failure after the first write — the store refusing a write, or a rebuilt hash missing the file’s — leaves a partial store that a retry refuses as already holding the run: discard it and restore into a fresh one.

It replays the ordinary append path rather than writing rows, and that is the safety argument: append maintains six derived indexes — case, exactly-once, outcome and its ordering counter, and both halves of the discovery index — and a restore that rebuilt five of them would produce a store that reads perfectly until somebody queries the sixth.

Two details make that reproduce the original bytes rather than similar ones. Runs are sealed in log-index order, because that order is the Merkle log — seal them in file order and the same leaves give a different root. And epoch is carried, not re-derived: it is inside the hashed body, so a run that ever changed hands would rehash under a single fresh lease, and those are exactly the runs a failover produced. Both backends fence only when a lease row exists, so restoring into a store with no leases writes each record under its own epoch.

The fencing token picks up where the history left off. The epoch is a fencing token, and the lease table that normally holds its high-water mark is store-local state an export does not carry — so a restored run has records reaching epoch 7 and no lease row at all. A first lease is therefore issued one past the highest epoch the run’s own journal records, not at 1. Restarting at 1 would give a run that had already changed hands a second ownership period wearing the number of its first, and the two would be indistinguishable in the only durable record of either — which matters beyond tidiness, because quota settlement decides which pass a spend belongs to by matching that number. The rule is checked in the store conformance battery on both backends.

What it cannot see is an epoch that was issued and never written under: a lease taken by an owner that died before appending leaves no record, so nothing durable knows the number was handed out. It bounds every epoch that ever wrote, which is every epoch that can conflict with the rebuilt history. The rest is the requirement every failover already has and the runbook owns — the old deployment must not reach the restored store.

What does not survive is named in the report rather than left to be discovered.

A restored wait is armed by nothing, and this is the entry that costs work. RestoreReport::awaiting lists the runs whose history ends in a wait, as ids rather than a count, because each one needs an act. The wait itself is journaled — RunSuspended carries the instant, the kind and the correlation — but what makes a wait happen is a row in a store the export does not carry: a timer, a subscription, and any worklist row the wait opened. So a restored suspended run has no timer to fire and no subscription to match, and because it released its lease cleanly when it suspended, the recovery pass does not see it either. Nothing in the system names it.

Resuming it repairs it, and not by new machinery: the runtime already treats an announced-but-unarmed wait as repairable, because a crash between the announcement and the registration produces exactly this state. A resume replays to the wait, finds no terminal record, and re-arms the timer, re-subscribes and re-opens the task row from the journal — until derives from journaled reads, so the instant is the one the original recorded. restore cannot do it: a resume runs the agent’s own code and a restore holds stores. So the runbook step is resume every run in awaiting, and a plane that skips it has restored a history and not the work.

Four layers are not in the export at all: the worklist and the decisions recorded against it, inbound events nobody claimed, webhook registrations with their delivery cursors, and governed memory beside the batch, quota and standing-authority ledgers. A decision a run already consumed survives, because that run journaled it; one nobody had consumed does not. Re-establish webhook registrations by hand — nothing journals a delivery cursor, so nothing can rebuild one.

Signatures: append attests as the restoring store’s signer, so a history signed by a key this store does not hold comes back unsigned — hashes and the root are unaffected, since a signature is taken over the chain hash and stored beside it, but authorship is gone. Configure the same signer if you are restoring your own log. Activity timestamps: the discovery index is rebuilt at restore time, so recent_runs orders by when history was restored. It is documented as ordering and cursor stability only, and nothing derives a decision from it.

What is exported is what the chain committed to. With a key ring configured that is ciphertext, deliberately: an export of plaintext would put a copy beyond the reach of key destruction and undo the erasure the key ring exists for.

Break-glass

Reaching another tenant’s data is the one exception to isolation, so it is recorded in that tenant’s journal — actor, roles, and a reason that cannot be blank — before anything is served:

let plane = planes.cross(&caller, &target, "INC-42: stuck settlement").await?;

Planes::cross hands back the plane only once the crossing is on that tenant’s record. That ordering is the control, and making it a door rather than a step is what makes it one: a break-glass that serves first and records best-effort works exactly as well when its own evidence is lost, which is the state an incident is most likely to produce.

The whole Caller goes in rather than its actor, roles and tenant separately, because those are one fact — who authenticated. Passed apart, a handler can record one operator’s name against another’s crossing, and the record is written, signed, and wrong.

The ordinary lookup, Planes::get, takes the same Caller and serves its tenant. That is what leaves cross as the way to reach another one: a signature taking a bare tenant id cannot tell mine from somebody else’s, so it serves both and the difference lives in whether the handler remembered which to pass. The narrow claim worth making is that no path reaches another tenant’s plane by accident, because none can name one — not that a cross-tenant read is impossible, since an embedder writing its own Authenticator decides what a credential means and can mint a Caller for any tenant. That seam is where a deployment defines identity; it is a deliberate act rather than a forgotten step.

cross is what an admin surface should use: it records the crossing and hands back the plane in one call, so no ordering is left to the caller. record_break_glass is for an embedder recording a crossing they made some other way.

Same-tenant crossings are refused, because recording routine access as break-glass buries the real ones. A tenant this process does not serve is refused rather than defaulted, exactly as the ordinary gate refuses one.

The crossing seals like any other run, so it enters the Merkle log, verifies offline, and lists with the rest:

curl -H "$AUTH" 'https://plane/runs?outcome=broke-glass'

GET /runs/{run} on one answers status: "broke-glass", decided_by with the operator, and reason with the words they gave. Both halves, because an incident review asks two questions and a surface that answers only why sends the reader to the export for the name.

Who may pull it is your policy engine’s decision, not this crate’s.

One matter, one scan

Show me everything about this matter is the question a regulated deployment asks, and there are two ways to answer it. Listing the case’s runs and reading each is a join whose cost grows with the case’s life — and it misses every record written by a run the case does not own.

A sweep is exactly that run. One tick may escalate several cases and belongs to none of them, so a per-run walk never reaches the record explaining why a case was escalated — which is the only reason for writing it down.

So the journal carries the case on the record and indexes it:

let history = store.case_history(case, 200).await?;

One range scan, tenant-first like every other key here, so a query that forgets the predicate returns nothing rather than another customer’s matter. Both backends are held to it by the same conformance battery, which checks the two halves separately — that the scan finds this matter’s records, and that it returns no other’s. Either alone passes for the wrong reason: a scan returning nothing satisfies the second, and one returning everything satisfies the first.

GET /cases/{id} includes it, with history_truncated when the bound bit, because a shortened list is shaped exactly like a complete one.

The sweep writes its own history

The sweeper makes the plane’s most consequential automated decisions: it breaches an obligation, escalates a case, expires a person’s task. Nothing asked it to — that is the point of it — so there is no run whose history explains why the state changed.

Without a record, why is this case escalated is answerable only from the resulting state, and state cannot distinguish the sweep breached this at 02:00 from somebody set it. No human was present to remember which.

So a tick that decides anything writes its decisions into a sealed run of its own: Swept { subject, action, detail }, one per action, with a typed action rather than a message. It inherits the chain, the per-record signature and the Merkle inclusion every other run has, so the external audit tool checks it without being taught what a sweep is.

let report = plane.sweep(now, grace).await?;
if let Some(run) = report.record {
    // Exactly what this tick decided, verifiable like any other run.
    let entries = store.read(run, 1).await?;
}

A quiet tick writes nothing. A healthy plane sweeps constantly, and a Merkle log filling with evidence of inactivity is where the somethings hide.

A breach is applied first and noted after; a warning or an expiry is noted first and corrected if it lost. A breach is conditional: a run that met the obligation after the sweep read it wins, and the store refuses the stale breach — so a refused breach writes nothing. An applied one is marked in the same write as owing its account, and the deadline_breached and case_escalated notes retire the mark. A crash between the breach and its notes leaves the mark, and the next tick notes every owed breach before it reads due — once. A warning, and a task expiry that meets an answer somebody already gave, are noted before they act; when the act does not apply the sweep appends not_applied for the same subject, so the last word about each subject is what happened to it. Read a subject’s deadline_warned or task_expired as withdrawn when a not_applied follows it. An expiry that meets a person’s answer also settles the task to that answer, so the next tick does not meet it again.

Two things are deliberately not in it, and both are worth knowing. Dead-lettered events are counted rather than named, because the event store reports how many aged out and not which — they stay in the report and the emitted event, and there is deliberately no SweptAction for them, because a variant nobody constructs reads as a capability. And a sweep run is not a plan: it seals as swept, never as succeeded, because a tick that breached forty obligations is not a plan that completed. GET /runs/{run} answers swept with no reason — what it decided is on its records, not in a one-line summary.

A capped tick says it was capped

Each sweep takes a bounded batch — 128 timers, 512 obligations, 512 expired tasks, 128 stalled deliveries, and 32 abandoned runs, the smallest because recovering one executes live from the frontier — so one tick is bounded. A sweeper still working through a backlog is a sweeper not noticing the next obligation, which is the failure the whole mechanism exists against.

The hazard is that a bounded query returns a list shaped exactly like a complete one. A tick that handled its cap and a tick that handled everything produce the same counters, and they are the two states most worth telling apart: the first means the backlog is growing while the report looks ordinary.

So SweepReport::saturated names which sweeps came back full, and needs_attention() is true when any did. A saturated tick means at least the cap was waiting — never that the cap was all there was.

let report = plane.sweep(now, grace).await?;
if report.saturated.deadlines {
    // More obligations were outstanding than one tick will take.
    // Sweep more often, or find out why they are accumulating.
}

Outbound delivery: two ceilings, and what each one means

A push sweep is bounded twice, and the two answer different questions:

BoundsSymptom when it is the binding one
run_once(at, limit)how many registrations one tick takesPushSweepReport::saturated — more were due than the tick would take
DeliveryWorker::max_in_flighthow many receivers are contacted at once (16 by default)the sweep is not saturated and still slow: it is waiting on receivers

Raising limit without raising max_in_flight buys a longer tick rather than a faster one; raising max_in_flight without limit leaves work behind that one tick was never going to reach.

Registrations are served concurrently; the records within one are not. A receiver still gets its records in journal order, and the cursor still advances only on a 2xx, because that is the only thing a cursor means.

Metrics

The runtime does not measure durations

Ambient clocks are lint-denied with three named escapes, each for a value that gets journaled or is store metadata. A fourth escape for instrumentation would end the rule, because timing is the most plausible-sounding reason anyone reaches for a clock. And a replayed run would re-measure durations belonging to calls it never made, so “effect latency by driver” would average network time with journal reads — the failure agentplane.effect.replayed exists to prevent, arriving through the metrics door.

Durations are therefore derived from spans, by the collector. The spans carry the mode and the replay flag, so a collector can compute latency and exclude replays, which an in-crate histogram could not.

Counters are emitted; gauges are observed

“Open cases” cannot be an increment-on-open, decrement-on-close counter. A crash between the state change and the emission loses a decrement permanently, and the dashboard slowly invents open cases that do not exist — plausibly, which is worse than obviously.

So gauges come from a census query against the store, and the sweeper emits them: it already runs periodically and already takes its now as a parameter, so no clock is read. The census is also the only consumer of a case’s opened_at, and the reason that column exists — a count cannot distinguish ten cases open for an hour from ten open for a month.

A gauge must never be read from a limit-bounded query. That is why census exists rather than by_status(..).len(): a paged count rises, flattens at the page size, and looks like a plateau exactly when it has become a backlog.

A counter cannot report a backlog, and the difference bites hardest on the one that matters. agentplane.quarantines counts the moment a run is set aside. It is monotonic and it lives in the process, so after a restart it reads zero over a backlog of forty, and a backlog that stopped growing reads exactly like one that was cleared. The number to alert on is agentplane.runs.quarantined — the level, counted from the store’s outcome index on every sweep, which is the same index GET /runs?outcome=quarantined pages.

It falls when somebody answers one — see Answering a quarantine. What it must not be read as is risk retired: an abandoned run leaves this count while whatever it left in the world stays exactly where it was, which is why the doubt is delivered as an agentplane audit finding derived from the journal rather than as a status somebody can clear.

Two rules, both guarded

A dimension is a variant, never a rendered message. Display on a budget error embeds the allowed and used figures; a label carrying those is one time series per distinct budget, which is how a metrics backend falls over. Every dimension comes from an as_str() accessor.

The catalogue is not a wish list. A declared-but-unemitted event leaves an empty panel, which at least looks wrong. A declared-but-unemitted counter reads as a hard zero — indistinguishable from “this never happens” — so an operator concludes the system is healthy from a number nobody wired up. tests/guards/layering.rs fails the build if a catalogue entry has no emitter.

Observability

tracing spans and events, so the runtime is usable by any subscriber — OTel, JSON logs, a test recorder — without the crate choosing an exporter.

agentplane.run                     gen_ai.operation.name = invoke_agent
                                   gen_ai.agent.name, gen_ai.conversation.id
                                   agentplane.run.id, .mode, .semconv
└── agentplane.step                agentplane.step.id, .capability; agentplane.phase
    └── agentplane.effect          .kind, .key, .attempt, .mutates, .replayed
                                   agentplane.mode; error.type when it failed
                                   gen_ai.operation.name = execute_tool | chat
                                   chat:          gen_ai.provider.name,
                                                  gen_ai.request.model,
                                                  gen_ai.response.model,
                                                  gen_ai.response.finish_reasons,
                                                  gen_ai.usage.{input,output}_tokens,
                                                  gen_ai.usage.cache_{read,write}.input_tokens
                                   execute_tool:  gen_ai.tool.name

Spans follow the OpenTelemetry GenAI semantic conventions where they apply: a tool call is execute_tool, a completion is chat. Each effect declares its own operation and its own attributes rather than having either inferred from its name, so a new effect type cannot pick up a label by accident — and the keys do not cross, because a completion naming a tool would make every model call look like a tool invocation.

Where the convention names a fact, the convention’s name is what is emitted. An agent’s name and a conversation’s identity are gen_ai.agent.name and gen_ai.conversation.id, not a second spelling in this crate’s namespace that no generic tooling reads. The conversation is this plane’s case, which is also what the A2A surface answers contextId with.

gen_ai.response.model is the one attribute a request cannot stand in for: it is the provider’s own word about which weights served the call, and it differs from the request whenever an alias resolves, a deployment is moved under a pinned name, or a gateway routes elsewhere. It is absent where the wire does not say — Bedrock’s Converse names no model on the way back — and it is never filled in from the request, which would report that a substitution had been ruled out when nothing looked.

error.type carries the class of fault on any attempt that failed — timeout, refused, rate_limited, metered. agentplane.outcome says whether an attempt succeeded and makes a failure countable; this says what went wrong, so “which driver fails how” is a group-by rather than a search through rendered prose. Without it, a dashboard built on the convention reports a plane with no failures at all.

agentplane.effect.key is the join between a trace and the evidence. Everything else on the span names a category; this is the journal’s own identity for that attempt, and GET /runs/{run}/history serves records under it. It names the attempt without carrying what it sent — which is the claim worth making, since a digest is still a commitment to the values it was taken over.

Effects that are not GenAI operations — reading the clock, arming a timer, writing case state — carry no such attribute at all. That is deliberate: labelling them would make the attribute useless to the tooling that keys on it, which is the whole reason to emit it.

What is deliberately never emitted

The conventions define Opt-In attributes that carry the content of a call: gen_ai.input.messages, gen_ai.output.messages, gen_ai.system_instructions, gen_ai.tool.definitions and gen_ai.tool.call.arguments. This plane emits none of them, and it is not a default to flip. A prompt is exactly where governed values arrive; the sink gates exist to keep those values inside a declared ceiling; and a trace exporter is an egress those gates do not cover. Sensitivity is a property of a value, and a span is not a sink that can carry one.

What a deployment gets instead is the shape of every call — who was asked, what it cost, which tool ran, how it ended — and the journal for the content, where the same values are sealed, labelled and erasable.

Three more the convention defines that this plane does not emit, each for its own reason. gen_ai.response.id, because no shared part of a completion carries one. gen_ai.provider.name on the run span, because the agent is executed here rather than served by anyone — the completion spans beneath it each name the provider that answered them. And server.address, because a driver’s endpoint is deployment configuration, identical for every call it makes: that belongs on the OTel resource your exporter sets, not repeated on every span.

The same rule reaches the loud events. A run’s failure reason, a quarantine’s, an undecidable effect’s detail and a compensation’s error are free text a provider, a peer or a skill wrote, and it quotes the request it refused — the caller’s data. The journal seals those fields under a key ring and an erasure destroys them; a log line carrying the same words would outlive both. So each of those events carries the run, the class of fault as error_type (undecidable, timeout, failed, …) and reason_digest — sha256: and the first sixteen hex digits of SHA-256 over the text — and never the text. Read the words through the operator API (GET /runs/{run}), which honours erasure; the digest is what joins the two.

Transport failures follow the same line one layer down: a webhook’s, a peer’s or a media source’s URL routinely carries a bearer secret, so an outbound error is rendered with the host it failed at and never the URL — in a parked push registration, a log line and an effect’s failure alike.

The conventions are still pre-1.0, so the revision targeted is pinned in telemetry::SEMCONV_VERSION rather than tracked. An upstream change becomes a deliberate migration instead of a silent shift in what your dashboards mean.

Further decisions worth knowing:

  • One span per effect attempt. A retried call shows as several spans rather than one long one, which is what makes “how often does this driver need a second try” answerable.
  • Replay is marked on every span. A replayed run re-executes its skills and emits spans again. An effect served from the journal is reported as an event with replayed = true, never as an effect span — otherwise “effect latency by driver” averages real calls with journal reads.
  • Spans attach to futures, never to threads. Span::enter returns a guard bound to the current thread; held across an .await it stays entered while the future is suspended, so whatever runs next is attributed to it. With concurrent siblings that silently reparents their work. Instrument is the only form that survives a suspension, and tests/guards/layering.rs bans the guard in async code.
  • The vocabulary lives in runtime::telemetry. A span name typed inline at twelve call sites is twelve chances to drift, and telemetry drift is invisible: the dashboard stops matching and nobody is told.

Every failure that must not pass silently has its own event target:

EventFires when
agentplane.run.nondeterminism_detectedReplay recomputed a different effect key
agentplane.run.quarantinedA run was set aside for a human
agentplane.run.abandonedA person closed a run whose outcome could never be established — nothing was unwound, and the doubt is now an audit finding
agentplane.run.unreproducibleA pinned read came back as different content — the durable record is not trustworthy
agentplane.run.recoveredThe sweep took over a run whose owner died holding it
agentplane.run.replannedA run changed its plan, and the successor names its parent
agentplane.run.failedA run concluded failed — an ordinary conclusion rather than an incident, findable under GET /runs?outcome=failed; the event carries the reason’s digest, and the API the reason
agentplane.effect.undecidableAn outcome could not be determined and guessing was forbidden
agentplane.effect.reconciledA probe was asked whether a call landed
agentplane.budget.refusedA limit refused an operation
agentplane.policy.deniedThe deployment’s rules refused an action
agentplane.saga.compensatedA completed step was undone
agentplane.saga.compensation_failedA compensation failed, leaving the run partly unwound
agentplane.event.dead_letteredAn event aged out with nobody waiting — a correlation bug
agentplane.deadline.breachedAn obligation passed unmet
agentplane.timer.firedA sleeping run’s instant arrived
agentplane.witness.integrityA witness refused this plane’s checkpoint: the log shrank, forked, or claimed growth the witness could not verify

That list is telemetry::LOUD_EVENTS, and this table is checked against it: tests/guards/docs.rs fails the build if the runtime promises an event this page does not name, because a table headed every is an alerting checklist and a short one is read as a complete one.

tests/guards/layering.rs fails the build if any of those has no emitter, and tests/process/telemetry.rs asserts on what a subscriber actually received rather than on what the source contains — an instrumentation test that greps is checking the author’s intent, not the runtime’s behaviour.

The last mile: instrumented is not monitored

Shipping no exporter is the right call for the same reason as shipping no policy engine, and it leaves a gap that a deployment has to close by hand. Three things sit in it, none of them obvious from the API docs, and cargo run --example observability is a runnable bridge that does all three and asserts on the result:

  1. Latency must exclude replay — and the runtime already makes this easy in a way worth stating, because it is not what a reader assumes. A replayed effect opens no span at all; it emits a debug event on the agentplane.effect target with replayed = true, plus the agentplane.effects.replayed counter. So a histogram built from spans is clean by construction. What is not safe is keying on the target and treating everything on it as a span: that view sees both, and it is the one in which replays quietly improve your p99 in proportion to how much recovery you are doing.
  2. Scheduling the sweep is scheduling your gauges. Gauges come from a census queried against the stores, never from increments — see below for why — so a plane with no sweep loop has counters and no gauges, and nothing says so. A dashboard with no data looks exactly like a plane holding nothing.
  3. SweepReport::needs_attention() is the alert predicate. It already folds in breaches, expiries, dead letters, saturation, failed recoveries, lost evidence and an unreadable census. Re-deriving it from individual counters is how the next failure mode ends up alerting on nothing.

The example’s module docs carry the OTLP wiring verbatim — pinned crate versions, the tracing-opentelemetry layer, and where each of the three rules goes — so the difference between it and a production subscriber is the exporter and nothing else.

Then schedule the two loops, or let the binary do it:

agentplane serve agent.yaml \
  --sweep-every 30      `# gauges, deadlines, task expiry, dead letters` \
  --drill-every 86400   `# the recovery rehearsal` \
  --drain-secs 25       `# how long a SIGTERM may keep working`

The operator surface

Feature http, off by default. A library embedded in someone else’s process should not open a port unless asked.

There are two surfaces, and the second one matters most when the first is gone. The HTTP API acts on a running plane; each CLI verb opens a store. On the embedded backend that division is load-bearing, because redb admits one writer process: a verb that opens the store is a verb that runs while the plane is down — which is exactly when somebody is reaching for it.

So every remedy agentplane attention names is a verb the CLI has, and every remedy GET /attention names is a route the API has — or, where the API has none (resuming a run, re-running a drill), the agentplane verb that does it. A remedy with nothing to run says no verb: and why. Each condition also lists its subjects — the run, task, <case>/<obligation>, message or <run>/<registration> the remedy’s verb takes, at most ten, newest first for run conclusions and most overdue first for waits and tasks, with unlisted counting the rest — so the next command can be typed from the answer alone:

agentplane attention --store ./journal.redb
# { "condition": "run.quarantined", "found": 1, "subjects": ["run_01JD…"],
#   "unlisted": 0, "remedy": "establish an undecided effect with `reconcile`, …" }

A failed run is listed only when it stands on landed work nothing undid — run.failed_with_landed_work, whose remedy is replay to finish the work or cancel to unwind it. Any other failed run is one a resume may clear, and is not a condition.

reconcile establishes an effect the runtime could not decide; quarantine reopens or abandons the run afterwards; tasks lists the worklist and decide <task> approve|reject answers one; acknowledge accounts for a breached obligation; cancel stops a run the ceiling or the halt will not release; rearm revives a parked push registration. Each takes --actor, recorded as asserted — nothing at a terminal verified the name, and the record says so, where the same act through the API records authenticated.

Two of them cannot finish the job from a terminal, and say so rather than pretending. quarantine records the decision and reports "applied": false: a terminal holds the journal and not the agent, so the next resume applies it. decide records the decision durably, completes the task and reports what the delivery did — Buffered, from a terminal that holds no agent — and prints the agentplane replay <run> --manifest … --store … that continues the run. cancel from a terminal records the stop and leaves it to the plane that holds the agent, which observes it on its next resume.

Only a plane that can drive a run judges it. A terminal holds no policy engine and no declaration, and a governed run was admitted under both; a terminal that compared its own empty policy with the run’s would quarantine every governed run it answered a task for. So a resume first asks whether this plane provides every capability the run’s plans name, and a plane that does not stops there — NoProvider, reported as Buffered by a delivery and as a recorded, undriven request by a cancel. A plane that does hold the agent under a different policy bundle or an edited declaration still quarantines the run: that is a real mismatch, and the one a rollout across a mixed fleet has to plan for.

Serving it

Embedders wire Api into their own process; the agentplane binary’s serve verb does the same wiring from flags, hosting one manifest or one room file per process.

A listener per audience. The peer surface binds to --addr; the operator surface this section documents is opt-in beside it on --operator-addr, and the MCP surface agent frameworks call on --mcp-addr. A network policy can then treat another agent calling in, a framework calling a tool and a person deciding a task differently rather than trusting one port with all three.

FlagDefaultWhat it decidesEnv
--storerequiredWhere the journal lives on disk.AGENTPLANE_STORE
--addr127.0.0.1:8080The peer surface. Loopback until you say otherwise.AGENTPLANE_ADDR
--operator-addroffThe worklist, task decisions and GET /runs?outcome=quarantined, on their own listener.AGENTPLANE_OPERATOR_ADDR
--mcp-addroffEvery agent in the file as MCP tools at /mcp, on its own listener — see MCP, being served.AGENTPLANE_MCP_ADDR
--mcp-allowed-hostloopbackA Host the MCP listener answers. Required for a non-loopback --mcp-addr. Repeatable.—
--mcp-allowed-originnoneA browser Origin the MCP listener accepts; any other present Origin is 403. Repeatable.—
--mcp-agentevery agentAn agent, by metadata.name, to serve on --mcp-addr; the others are left off. Without it every agent is served and each must declare spec.input. Repeatable.—
--policyno defaultThe Cedar policy set. No default, because a permissive engine and no engine are the same behaviour and only one of them looks governed.AGENTPLANE_POLICY
--tokensno defaultThe bearer tokens this plane accepts, each optionally carrying the scope and not_after of the chain that caller’s runs act under. A token under 32 bytes, or one published in this project’s examples, is refused at load.AGENTPLANE_TOKENS
--log-formattextjson writes one JSON object per log line.AGENTPLANE_LOG_FORMAT
--urlnoneThe A2A endpoint callers reach this plane at — /a2a under the public address. Goes on the Agent Card, so it is the public URL rather than what you bind; refused unless it ends in /a2a.AGENTPLANE_URL
--sweep-every30sHow often deadlines, task expiry, dead letters and due timers are swept. 0 runs the sweep from your own scheduler instead.AGENTPLANE_SWEEP_EVERY
--drill-everyoffHow often the recovery rehearsal runs — see disaster recovery.AGENTPLANE_DRILL_EVERY
--drain-secs25sBounds the stop — see stopping an instance.AGENTPLANE_DRAIN_SECS
--mcp NAME=COMMANDnoneAn MCP server to run and wire under NAME. Repeatable.—
--peer NAME=URLnoneAn A2A peer the manifest grants under tool://NAME/…, called with the token in AGENTPLANE_PEER_TOKEN_<NAME> (. and - as _; two names sharing a variable are refused). The token names nobody → peers. Needs a2a.—
--push-hostnonePermits A2A push notifications to that exact host. Repeatable. Without one, push is not wired and the Agent Card advertises it as absent rather than claiming a capability nothing serves.—

The three repeatable flags take no environment variable; everything else does, which is how a container image is configured without editing its command line.

Identity comes from the request, never from its body

This is the whole design, and everything else in the module follows from it.

Four-eyes is enforced in TaskStore::claim, which takes an actor and a set of roles. In-process both come from the embedder’s own code, which is trusted. Over HTTP they would come from whoever is on the socket — and a reviewer who can name themselves can name the person who proposed the action. That is not a bypass of the control; it is the control, inverted.

Discipline does not hold that. So the wire type has no field to hold it:

pub struct DecisionRequest {   // no `actor`. no `roles`.
    pub approved: bool,
    pub reason: String,
    pub amendment: Value,
    pub digest: Option<Digest>,   // the task version decided on
}

The handler builds the Decision from the authenticated Caller, because there is no other source available to it. A later maintainer cannot be talked into reading the body’s actor, since there is nothing to read.

What lands on the record is an Operator whose basis is authenticated, and this is the only surface entitled to claim it — an Authenticator named the caller before the route ran. The same approval given at a terminal records asserted, which says the name is whoever ran the command’s own account of themselves. An auditor can tell the two apart, which is half of what an approval is worth as evidence.

deny_unknown_fields is the other half. Without it, a body carrying "actor": "alice" is accepted and silently ignored — the integrator who wrote it believes they decided as Alice, the journal says Bob, and the disagreement surfaces at an audit months later. A 422 says so at the first call instead.

amendment is not a note. On an approved call task it is the call: the runtime dispatches the reviewer’s arguments in place of the model’s, labelled as the reviewer’s — trusted, provenance task:agent.approve_call, the original arguments’ sensitivity — and every sink gate judges them like any other value. An amendment that does not fit the tool’s declared argument schema dispatches nothing. On a rejection it stays what it reads as: recorded advice (“no, and here is what would pass”).

Two gates, and the surface will not start without the second

Authentication says who; it does not say what they may do. An operator surface that stops there hands every authenticated caller the whole plane. So every route runs gate(), which authenticates, resolves the caller’s tenant to a plane, and then authorizes through that plane’s PolicyEngine under an api: action — and Api::new returns an error if any registered plane has none, or if none was registered at all.

That refusal is deliberate. In-process an absent engine is a choice; on a socket it is a hole, and a permissive default is one nobody discovers until the port is reachable. DenyAll exists for wiring the surface up before the rules are written. One ungoverned tenant among governed ones is the one an attacker looks for, which is why the check is over every plane rather than the first.

Checking a driver against the real thing

just test-live exercises the OpenAI, Gemini and compatible-wire drivers plus the embedding wire against the actual APIs, loading keys from .env. Each battery skips on its own credential, so one key runs one battery and the rest say so. It is gated twice — AGENTPLANE_LIVE=1 and the key — and is never part of just ci: a developer with OPENAI_API_KEY exported would otherwise be billed for running the test suite, and would find out at the end of the month.

The compatible-wire battery runs on HF_TOKEN against Hugging Face’s router, or on CHAT_COMPLETIONS_BASE_URL pointed at a local engine — which is the more useful thing to do before trusting one.

Worth having because a stubbed provider cannot have the defects a real one finds. It accepts any request shape and returns whatever it is told to, so a driver that sends a malformed body, or mis-reads a response, passes every offline test. What these catch is exactly that: a tool declaration in the wrong shape for the API being called, a plan format no provider with constrained decoding accepts, a prompt field mistaken for the wire’s own.

Putting a tenant on your telemetry

Off by default, and the default is the interesting part. A tenant name is frequently a customer name, and an observability backend is usually the least protected system in a deployment: sampled into third-party services, on a dashboard nobody signs into, retained past every other record. A deployment that has not decided where its telemetry goes has not decided that customer names may travel there.

.tenant_label(TenantLabel::Name)

One policy, both signals. It is the tenant field of every metric event and agentplane.tenant on the run span — because which tenant is this answered on the metrics and unanswerable on the traces is the same decision honoured on one channel. On the run span only: every other span is inside that run’s trace, so repeating a per-plane constant on each of them is bytes without information.

Under the default the attribute is absent rather than blank, so a collector cannot read not disclosed as no tenant.

Cardinality is bounded by configuration, not by data. The label is this plane’s tenant, so the number of streams is the number of planes you wired. There is no request that can grow it — a tenant read from a request would be exactly the unbounded label that makes a metrics backend fall over.

There is deliberately no pseudonymous option. Hashing the name here would cover one of the many places a tenant already appears — store keys, blob paths, the policy request, the checkpoint origin published to witnesses, and the tenant field on an Agent Card served unauthenticated at a well-known path. A control that covers one exit and not the other nine is worse than none, because it invites the belief that the name is contained.

If customer names must not leak, do not put them in the tenant id. TenantId::new("t-9f3a") covers every one of those places at once, costs no code, and cannot fall out of step with a surface added later.

Per-tenant ceilings

Budgets bound one run. They do not bound a tenant: a caller that can start runs can start a thousand, each perfectly within its own ceiling, and the compute and the model bill are somebody else’s problem. RuntimeBuilder::quota sets a tenant’s limits and points at the store that accounts them. A plane built with Runtime::builder_on or builder_with already has the backend’s quota store with no ceilings, so halts and rate ceilings hold without this call; .quota states the limits.

.quota(store.clone() as Arc<dyn QuotaStore>, TenantQuota {
    max_concurrent_runs: Some(50),
    max_tokens_per_period: Some(20_000_000),
    period: Period::Monthly,
    ..Default::default()
})

The accounting is durable, and that is the whole point. An in-process counter is a ceiling that vanishes the moment a second instance starts — and it fails open, silently doubling when somebody scales out, which is exactly when it was needed. The reservation is one transaction that counts and inserts, so two instances racing for the last slot serialise and one loses.

What each ceiling bounds, stated precisely, because a limit believed to bound something it does not is worse than none:

Concurrency bounds runs executing. A slot is taken at admission and given back when the instance finishes with the run — including when it suspends, since a suspended run costs a row and not a thread, and holding its slot would mean a tenant waiting on a hundred approvals could start nothing. It follows that a resume is not gated: that work was admitted already, and refusing it would strand a run waiting on something that has now happened.

Spend bounds a period by reserving each run’s worst case at admission. A run can spend its own ceiling and end one call past it per step it has in flight — a call’s cost is known only when it returns — so admission holds ceiling + width × one call against the period, in the same store transaction that checks the ceiling and takes the slot, and refuses when settled spend plus every outstanding hold plus this one would pass the ceiling. The width is max_parallel_steps, or the admitted plan’s node count, and the run is dispatched no wider. What one call can cost comes from the declaration: each model role’s max_input_tokens plus its output ceiling, at its price (see the manifest reference); a Rust deployment states it on its Budget as max_call_tokens / max_call_minor_units. agentplane validate prints the figure per agent.

Under a spend ceiling, a run whose budget leaves that unit unbounded — no run ceiling, or no per-call bound — is refused with QuotaError::Unbounded naming the field. Admitted, it would hold nothing against a period it can spend without limit.

Every pass settlement moves that pass’s spend out of the run’s hold and into the settled total, and the pass that concludes the run releases the rest. A suspended run keeps what it holds, and a resume is never refused for spend. One live execution pass belongs to the period in which it starts: settlement charges it there even when midnight or month-end passes before it finishes. A resume in a later period moves what the run still holds into that period — released from the old one, held in the new one, unconditionally.

What this bounds. A period’s settled total stays within its ceiling through the work admitted in it, however many runs suspend, run concurrently, or admit from several instances at once. What still reaches past it: holds carried in by resumes from an earlier period (new admissions are refused until they settle); a model call that reports more input than its role declares, which fails but is billed as reported; an embedder’s own effects, bounded only as exactly as the per-call figure it declares for them; and compensating calls, which no ceiling gates. A commissioned run reserves and settles its own spend; the commissioning run’s own ceiling still counts it, and the period is charged once.

A QuotaError::SpentOut refusal names settled and reserved spend apart: settled spend resets with the period, reserved spend is held by open runs. Runtime::reservations(limit) lists the holders, and agentplane attention names the ones stopped waiting on a person — quarantined, exhausted, withheld or failed — since what they hold comes back only when somebody concludes them. A halt does not release a hold: the workload scopes stop admission, and only a subject: halt reaches work already running.

Wiring a quota store records every active run even when all ceilings are None. That keeps running() truthful for operators and means adding a limit starts from the work actually in flight, not from an empty ledger manufactured by the previous unlimited configuration.

Settlement is crash-safe rather than best-effort. Before a pass can dispatch an effect, QuotaPassStarted records its period and whether it owns the admission slot. Admission writes it with the run’s first records; a resume writes it with the first record the resume itself appends, so a resume that appends nothing — one reaching the conclusion the run already holds — leaves no marker, and a marker is never a run’s last record. At the end, QuotaStore::settle writes an exact (run, epoch) receipt, accrues that pass’s spend, reduces the run’s hold by it (releasing the rest when the pass concludes the run), and releases its slot in one store transaction. A lost acknowledgement repeats the same receipt and charges nothing twice; a changed retry is corruption. The run is physically sealed and its lease is released only after settlement succeeds. If settlement is unavailable, the lease expires still owned and the ordinary abandonment sweep derives the spend from the journal and retries it.

The journal and quota store may be separate systems, so this is deliberately not described as one distributed transaction. The journal is durable intent; the quota receipt is idempotent application. The protocol tolerates failure on either side of the call and converges after a transient outage. It cannot make progress through a permanent partition or survive independent destructive loss of one backend while claiming the other is complete.

A plane with no quota store resumes a run only when none of its passes accrued spend to a period, and refuses one that did, since that spend would go unbilled. It cannot give back the concurrency slot of an admission whose process died mid-pass; the sweep of a plane with the quota store releases it once the run is sealed, and counts it in SweepReport::slots_released.

RuntimeBuilder starts at Budget::unlimited(), so a Rust deployment that sets a tenant spend ceiling has to give its runs a ceiling and a per-call bound too — otherwise every run is refused as unbounded rather than admitted past the ceiling. A manifest without budgets already fails validation.

A refusal is RuntimeError::QuotaExceeded, deliberately not a policy denial: a denial means you may not and retrying is pointless; a ceiling means not right now. Over A2A it comes back as -32029 with the ErrorInfo reason QUOTA_EXHAUSTED rather than an internal error, so a peer backs off instead of retrying a “fault” immediately.

Two failure choices worth knowing. An unreachable quota store refuses rather than admits — a ceiling that yields when its accounting is down is one an attacker removes by taking the accounting down. And concurrency is tracked as a set of runs, not a counter, so a process that dies mid-run strands a slot an operator can name and release, rather than a number nobody can audit.

Naming them is QuotaStore::running_runs(limit). running() answers five of five, and one of those five may not be a run at all: a slot is taken at admission and given back at settlement, so an instance that dies in between holds one forever, indistinguishable from live work in a count.

let held = quotas.running_runs(100).await?;
let stranded = journal.abandoned_runs(100).await?;
// A slot whose run has no live lease is the one to look at. The recovery sweep
// resumes it and settlement gives the slot back.

The listing empties as runs settle, which is what makes it a queue rather than a record of everything that ever ran.

A tool’s rate ceiling

A grant’s rate_limit is counted in the same quota store, per tenant per tool, so it needs one wired: a plane built from the journal alone (Runtime::builder) whose declarations state a ceiling and that was given no quota store is refused at build. Each dispatch of a ceilinged tool costs one store transaction; an unceilinged tool pays nothing. An unreachable store refuses the call, as it refuses admission.

A run stopped by one concludes exhausted, and agentplane attention lists it with the other exhausted runs. The remedy is time, not a raise: wait for the window to pass, then agentplane replay the run. A replay inside a window that is still full stops it again on the refusal already recorded; after the window it records a re-admission beside the refusal and carries on. Changing a ceiling means publishing a declaration; a new revision does not reset the count.

The reservations are operational state, like delivery cursors: they are not in the export, and they age out with the widest window stated for their tool.

The emergency stop

Beside the ceilings sits a switch that is deliberately not one: Runtime::set_halt(&scope, &by, at, reason) stops new work from starting, across every instance, because the flag lives in the quota store rather than in the process — a stop that only stops the instance it was thrown on is the in-process-counter failure arriving during an incident. The refusal is its own error carrying the operator’s reason, never a ceiling: a ceiling says not right now and invites the retry somebody pulling this switch is trying to stop.

It names what it stops

A tenant-wide switch is the right answer when the plane is the incident and the wrong one when one agent of many on the plane is misbehaving — and hosting several agents is exactly what a multi-document manifest and A2aServer::hosting are for.

ScopeStopsReach for it when
HaltScope::Tenanteverything this tenant would startthe plane is the incident
HaltScope::agent("payments-clerk")every revision of one declared namethis agent is misbehaving and I do not yet know since when
HaltScope::revision(digest)one exact reviewed revisiona bad deploy — a fix published as a new version runs while the broken one stays stopped
HaltScope::subject("alice")everything acting for one principal, wherever it stands on a run’s delegation chain, including runs already in flightthe incident is a credential, not a workload: a laptop lost, a service account that turned out to be shared, somebody who has left

The first three ask what is running; the last asks who it runs for. That is read from the chain the run was admitted under — on a served plane, the caller’s, not the operator’s — and matches any link on it: the person at the root, each workload it was delegated through, and the one acting. Authority flows down the chain, so withdrawing alice stops the work alice asked for, including what a service is doing on her behalf, and leaves everyone else’s alone.

It is the only scope that reaches work already running, and it pauses. A run under a withdrawn credential stops at its next step boundary as withheld — or sooner, at its next call to a peer that is told who asked, which reads the halts before presenting a credential and drops any it holds for the withdrawn subject. Its completed work stands, and after the lift a replay continues it. It is not unwound — reversing correct work because a credential lapsed is a second incident. To reverse it, cancel it.

Both halves are journaled: AuthorityWithheld at the pause and AuthorityRestored at the lift, the second beside the first. A reader takes the last word.

agentplane halt  --store ./journal.redb --scope 'agent:payments-clerk' \
                 --reason "incident 42: looping" --actor ops-carol
agentplane halt list --store ./journal.redb      # what is stopped right now
agentplane halt  --store ./journal.redb --scope 'agent:payments-clerk' --lift \
                 --actor ops-dave
agentplane halt list --store ./journal.redb --lifted   # who lifted what, newest first

--scope takes tenant, agent:<metadata.name>, revision:<manifest digest> or subject:<principal on the chain> — which stops every run with that principal anywhere on its delegation chain: the person at the root, a workload they delegated through, or the one acting. Those are the forms HaltScope::parse accepts, and the refusal you get for a typo lists them.

--actor is required to throw one, and the row says it was asserted. The runtime cannot check an emergency stop: there is no verdict to re-derive and no policy that authorized the judgement, so the whole of its evidentiary weight is the name beside it. Through the operator API that name comes from the credential the authenticator verified and is recorded as authenticated; at a terminal nothing verified it, and what it proves is that whoever ran the command could open the store. Both are legitimate — the second is how an incident is handled when the plane itself is the problem — and the row keeps them apart so a reader two years on is not left guessing which they are looking at.

A lift is recorded, and --actor is required for it too. Before the row goes, the lift is written to the journal as the one HaltLifted record of a sealed run of its own, outcome halt-lifted: who lifted it and on what basis, when, and the whole stop it ended — scope, reason, who threw it and when. A lift that cannot be recorded does not happen. The record attests the instruction and the row as read, not that the removal succeeded; halt list is the answer to what is in force. Only the halt the record names is removed: one re-thrown between the read and the removal stays standing, and the lift fails naming the record’s run, as it does when the removal itself fails (a 500 over the API). The answer’s removed is false when another lift removed the row first. halt list --lifted (a page of --limit, default 100) and GET /halts?state=lifted read the lifts back, newest first, and the runs travel in an export like any other sealed run. Where the stop reached a run — the subject: scope, the one that does — the run’s own journal also holds both halves, with the operator from the halt on AuthorityWithheld.

Those commands open the store, and redb admits one writer process — so against the file an agentplane serve is holding they fail, saying so in those words. Two ways round it, and both are ordinary rather than workarounds:

# 1. The operator API reaches a running plane, whichever backend is under it.
curl -sX POST "$PLANE/halts" -H "authorization: Bearer $TOKEN" \
     -d '{"scope":"agent:payments-clerk","reason":"incident 42: looping"}'
curl -s "$PLANE/halts"  -H "authorization: Bearer $TOKEN"
curl -sX POST "$PLANE/halts/lift" -H "authorization: Bearer $TOKEN" \
     -d '{"scope":"agent:payments-clerk"}'

# 2. On the shared store, the CLI and a serving plane coexist.
agentplane halt --store "$DATABASE_URL" --tenant acme \
                --scope 'agent:payments-clerk' --reason "incident 42" --actor ops-carol

Throwing the stop and lifting it are separate capabilities — api:halt.place and api:halt.lift. Granting somebody the power to stop the plane says nothing about who may start it again, and one grant covering both would make that distinction unwritable. A lift over the API is recorded under the authenticated caller, and its answer names the lifter and the run holding the record; a lift at the terminal is recorded under --actor, as asserted.

A halt closes the door; it does not empty the room. Runs already executing carry on, deliberately: cutting them mid-saga leaves reversals unrun and turns one incident into two. To reach work in flight, cancel it — and GET /runs/live is the listing that tells you which, with the agent, the revision and the delegation subject beside each id. A slot marked stranded belongs to the recovery sweep, which resumes it; cancelling one unwinds work that was about to finish.

Withdrawing a credential:

# Nothing new starts for alice, and what is running for her pauses.
curl -sX POST "$PLANE/halts" -H "authorization: Bearer $TOKEN" \
     -d '{"scope":"subject:alice","reason":"credential withdrawn: laptop lost"}'

# What is paused or still finishing under it — cancel what you want unwound.
curl -s "$PLANE/runs/live?subject=alice" -H "authorization: Bearer $TOKEN"

Scopes are independent rows, not one flag the last writer wins. Halting the whole tenant while an agent is halted, and then lifting the agent’s, leaves the tenant’s standing — an incident that widens and then partly resolves is the ordinary shape, and a single overwritable flag gets it wrong in the direction that lets work through. Where several halts cover one run the narrowest match is the reason reported, because “the tenant is halted” told to the caller of agent 12 sends them to the wrong incident.

An ungoverned run — a skill registered directly on the plane, with no manifest — is stopped only by a tenant halt. There is nothing narrower to key it on, and inventing a match would stop work for a reason nobody could look up.

Where a halt is answered, it is answered as itself. A peer over A2A receives -32030 with the ErrorInfo reason HALTED and a fixed message that carries none of the operator’s words — not the ceiling’s QUOTA_EXHAUSTED, whose advice is come back, because a peer that backs off and retries is doing exactly what the switch exists to end. A batch pass returns the halt as its error and leaves the items it stopped pending: an admission that never happened is not an outcome of the item, and the next pass admits them again.

--reason is required to halt and refused to lift: the next person to look will be somebody else, possibly at three in the morning, and why is the whole question, while a lift needs no justification because it restores the default. Who is asked of both.

What a workload-scoped halt does not stop

Runs already executing, and suspended runs resuming — deliberately. Those are existing work, and refusing to let them continue would strand them mid-saga with reversals unrun, turning one incident into two. Work in flight is stopped by cancelling it (request_cancel), which unwinds what it did and records who asked. cargo run --example operator_stop runs both brakes side by side, which is the clearest statement of the difference.

The subject: scope is the exception, and it is the one an incident most often needs. It names an authority rather than a workload, so it does reach runs already executing — each pauses at its next step boundary as withheld, completed work standing, and after the lift a replay continues it. The response to POST /halts says which of the two you got, derived from the scope, because an operator who withdrew a credential and reads cancel to reach work in flight will unwind a week of correct work that the withdrawal had deliberately preserved.

One surface, many tenants

Api::new takes Planes, a registry keyed by tenant, so one process can serve several — a single-tenant deployment passes its one runtime. Which plane answers comes from the caller’s tenant, which the Authenticator derives from the credential exactly as it derives actor and roles.

The gate hands each route its resolved plane, and Api holds no runtime of its own, so a handler cannot read a store without having established whose it is. A caller whose tenant has no plane is refused rather than served by a default: a fallback would turn an unregistered tenant into somebody else’s data while looking like working software.

The gate runs before the path is parsed, so a denied caller cannot learn whether a run id exists by comparing a 400 against a 404.

What the endpoints are for

Each route’s shape — its parameters, body, every status it can answer and the body of each, and the api: action it asks the policy — is in the OpenAPI 3.1 document, which agentplane openapi also prints (a build with the http feature). It describes the operator API only: the A2A surface is described by its Agent Card and protocol, the MCP server by its tool listing. No operation takes a tenant; it comes from the credential. Every refusal is {"error": "<sentence>"}, including a body the API could not read. Every build serves every operation the document lists; a plane built without the store or feature an operation needs answers it 501 after the usual gate. The document is outside the compatibility promise: generate a client per version.

A standard-library Python client generated from it is kept in the repository at clients/python/agentplane_operator.py: one method per operation, the parsed answer returned, OperatorError (status and sentence) raised for every refusal and for every redirect, which it never follows: a followed redirect turns a POST into a GET and carries the token to whatever host it names. Copy the file; it calls operator routes and runs no agent code.

RouteThe question it answers
GET /runs?outcome=…What ended this way and has not been cleared? Newest first; defaults to quarantined. The matching gauge is agentplane.runs.quarantined — alert on that, open this
GET /runs/liveWhat is executing right now, and under whose authority? stranded marks a lapsed lease → the emergency stop
GET /drillWhen did this plane last rehearse recovery, and did it pass? {"drilled": false} for a plane that never has is a finding, not a clean answer → recovery drill
GET /attentionDoes anything here need a person, and which thing? One roll-up over every backlog below, each condition with its subjects and a remedy in this API’s routes. agentplane attention exits non-zero when something does
GET /runs/waitingWhat is this plane waiting on, and until when? Soonest due first, so what should have moved already sorts to the front → the recovery runbook
GET /runs/{run}What is this run doing — why is it not finishing, or why did it end, and on whose decision?
GET /runs/{run}/historyWhat did it actually do? The journal, record by record, from ?from=<seq>
GET /tasksWhat is waiting for me?
GET /tasks/{task}What is this proposal, and may I decide it?
POST /tasks/{task}/claimThis one is mine — don’t let a colleague duplicate it
POST /tasks/{task}/releaseIt isn’t mine after all; give it back
POST /tasks/{task}/takeoverThe colleague holding it is out — displace their claim, naming whose it was
POST /tasks/{task}/decideApprove or reject, as myself
GET /cases?status=…What is escalated and has not been cleared? Newest first; defaults to escalated
GET /cases/{case}What has happened on this matter, and by when must it end?
GET /obligationsWhat did we miss and has nobody accounted for? Longest-overdue first, and it outlives the case’s closure
POST /obligations/acknowledgeI have looked at that one — take it off the list, keep the record
POST /runs/{run}/cancelStop it — 202, because the run stops at its next boundary
POST /runs/{run}/reconcileI looked it up: here is what actually happened to that effect
POST /runs/{run}/reopenThe doubt is answered — judge the run again
POST /runs/{run}/abandonNobody will ever establish what happened; close it where it stands
POST /eventsThis message arrived; wake whoever wanted it
GET /dead-lettersWhich messages arrived and reached nobody — the keys they were filed under, so the mismatch is visible
GET /haltsWhat is stopped right now, why, and who threw it. ?state=lifted: the recorded lifts, newest first
POST /haltsStop work at this scope starting — the emergency stop, reachable while the plane it stops is running
POST /halts/liftLet it start again, recorded under the caller. A separate capability from throwing it
GET /holdsWhat are we still keeping against erasure, and on whose instruction? Oldest hold first. ?state=released: the recorded releases, newest first → legal holds
POST /holdsPreserve this matter against every erasure verb. Idempotent: the response names the hold in force, which a second placement does not replace
POST /holds/releaseLift it, recorded under the caller, so the next retention pass reaches the matter
GET /pushWhich webhook receivers a delivery worker gave up on, and what they said last
POST /push/rearmThat one is fixed — resume at the record it never acknowledged

POST /events speaks CloudEvents, and its own shape. A bus posts a CloudEvents 1.0 message in either content mode — structured (Content-Type: application/cloudevents+json) or binary (ce-specversion, ce-id, ce-source, ce-type headers with the data as the body) — and anything else is read as this plane’s own {id, kind, correlation, payload} — with non-empty id and kind, for the same reasons the CloudEvents shape requires its three. The type is the event kind the policy gate is asked about, either way. An envelope the plane has not understood is a 400, never a guess: an unknown specversion, a missing id/source/type, a control character in any of the three, data_base64, a repeated ce- header, or an extension name that is not lowercase alphanumeric or that shadows a core attribute. A store outage is a 503, so a bus retries what a 4xx would make it drop.

The source a run is woken under is the authenticated caller, never the one in the body — otherwise a caller would hold both halves of (source, id) and could deduplicate against another party’s messages by naming them. The producer’s own source is not discarded either: it rides inside the buffered event’s id, so a gateway relaying two counterparties that both number their messages from one still delivers two events rather than swallowing the second as a retry.

The agentplane. kind namespace is the plane’s own. A human task’s answer reaches its run as an event of such a kind, so any kind in it is a 403 here, and an A2A message addressed to a run waiting on a task is refused the same way. A task is decided with POST /tasks/{task}/decide, where the claim, eligibility and four-eyes run.

One CloudEvents attribute correlates: subject, the standard’s own “what this event is about”, becomes the key ("subject", value) — a run that expects to be woken by a bus correlates on the subject its counterparty will name. Extensions deliberately do not; a producer that must correlate on richer keys posts the native shape, where they are stated.

Two details carry more weight than the plumbing:

A suspended run says what it is waiting for. “Suspended” tells an operator a run is stuck; it does not tell them whether to approve something, chase a counterparty, or page somebody. The SuspendReason is on the record, so it costs nothing to answer properly.

That status is read from the run’s last record, not from whether a suspension appears anywhere in its history. Every run that has ever waited for a human has a RunSuspended in it, forever — scanning would report every completed approval flow as permanently stuck, which is worse than reporting nothing.

The worklist says when it was cut off. The response is an object, not a bare array, because an array cannot express it: a queue of 140 items paged at 100 returns 100 and reads exactly like a queue of 100. The flag comes from asking the store for one more than the page and dropping it — inferring it from len() == limit would cry wolf on every queue of exactly limit.

Each worklist item says whether this caller may decide it. A reviewer barred by four-eyes still sees the task — hiding it leaves them wondering where it went — and is told on the item rather than by a refusal after they have read the case and made up their mind. The flag calls Task::may_decide, the same predicate the store enforces, rather than re-implementing it: a second copy of an authorization rule drifts, and the copy that drifts is the one people read.

Each item carries the text a person should read, beside the value a program reads. justification is the structured value; rendering is the same task with every invisible or direction-changing code point escaped in place (\u{202E}), escaped saying whether any was, and every word that mixes scripts listed in mixed_script with where it occurs. agentplane tasks prints the same rendering. A proposal the plane cannot open carries withheld — sealed (no key ring here), erased or undecodable — and then neither justification nor rendering carries a proposal or evidence at all, never the envelope. Approving it is refused with 422 and the reason, before anything is claimed or recorded; rejecting it still records → what a reviewer is shown.

Deciding on a version

Every task view — GET /tasks, GET /tasks/{task}, the claim answers and agentplane tasks — carries digest, the version of the stored row (not of the served justification, which withholds what the plane cannot open). A decision may name it: "digest" in the decide body, --digest on agentplane decide. A row that changed since is refused before anything is claimed or recorded — 412, or exit 1 naming the tasks --show that re-reads it — and if it changes between the read and the claim, the claim taken is released. Naming no digest compares nothing.

No authenticator is shipped

Same reasoning as the policy engine and the tracing exporter. Authenticator is handed the whole header map, because a deployment may authenticate by bearer token, mutual TLS, or a signed header from a gateway, and a parser baked in here would be wrong for one and load-bearing for the other. What is asked of yours instead is a contract — testkit::conformance_auth, and the rule in it that is easiest to get wrong is that Missing and Rejected are different answers → testing.

Claiming is what stops duplicated work

decide alone makes the queue first-past-the-post at decision time: two reviewers read the same case in parallel and one of them discovers, at the moment they submit, that the work was wasted. claim reserves; release gives it back, and it is release that makes claim safe to use — without it, a reviewer who claims something they then cannot decide has parked it until somebody edits the database, so the queue learns not to claim and the reservation stops meaning anything.

Claiming is not advisory. TaskStore::claim runs four-eyes and role eligibility in the same transaction that reserves, so an ineligible reviewer is refused before they read the case rather than after they have made up their mind.

Eligibility outranks availability

A refused claim is a 403 or a 409, and they ask different things of the reader — this will never be yours versus try again, or ask Bob. Checked availability first, a barred reviewer asking for a held task would be told “held by Bob”, wait for Bob to release it, and then be refused for a reason nobody had mentioned — having learnt meanwhile who is reviewing what, from a queue they have no standing in.

The order is part of the TaskStore contract, and the conformance battery holds both backends to it:

NotFound → Excluded → WrongRole → NotPending → AlreadyClaimed

The permanent refusal wins over the transient one, because the transient one hides it.

The same classification rides every verb that can refuse, not only claim: an id that names nothing is a 404 wherever it appears — claim, decide, cancel — and a store outage is a 500, never a 409. A conflict tells its reader somebody else got there first; answered to a typo it sends an operator hunting for an interventionist who does not exist, and answered to an outage it teaches a client that a retryable failure is permanent.

A release by somebody who does not hold the task reports ClaimError::NotHeld rather than succeeding — otherwise the caller is told the task is free while the holder still has it — and deliberately not NotFound: “the id is wrong” and “it is not yours” call for different responses. The battery holds both backends to it.

Escalation widens the audience — and then leaves the expiry scan

When a task’s window closes unanswered, the sweep applies the policy the task was opened with. deny and proceed record a decision and resume the run. escalate does three things in one store transaction: the declared escalate_to roles join the audience (a union — the original reviewers remain eligible), the stale claim is cleared so the widened audience can actually take the row, and the state moves to escalated. Four-eyes survives it: whoever proposed the action stays barred, however wide the audience gets. The escalated row stays in the ordinary queue, ranked and claimable — that queue, filtered by the caller’s roles, is where the wider audience meets it.

An escalated task does not reappear in the overdue scan, although it is still pending and past due. Escalation is the one policy that leaves its task in that condition forever — it is answered by a person or answered never — so a scan that kept returning escalated rows would accumulate them at the head of its bounded, oldest-first batch until the batch held nothing else, at which point the deny and proceed policies of every task behind them would silently stop firing. Reviewer attention is a finite resource; a queue that can be flooded is an oversight control that can be switched off.

A task id mixes in the run it belongs to

An EffectKey is unique within a run — the journal enforces (run, effect_key) and needs nothing more — while the worklist is a table every run shares. Two runs of one plan reach the same step, at the same ordinal, with the same descriptor, so a task id derived from the key alone would collide, and TaskStore::open is idempotent by id: the second run’s task would silently not be created, and it would wait for an answer nobody is ever shown.

So TaskId::derive hashes the run in as well, and the field is private — the collision is unrepresentable rather than avoided by care. The ("task", …) correlation key inherits that, being derived from the id.

🗄️ Retention and erasure

A full-fidelity journal is simultaneously an asset and a GDPR liability. The window is yours to choose — a retention period is a legal decision and a crate that picked one would be choosing somebody else’s — but running it is not: Runtime::retain walks it, agentplane retention plan lists what a walk would erase, and erasure is the full account.

What exists:

  • The journal keeps everything, indefinitely. It is append-only and hash-chained: nothing in it can be edited or removed without invalidating every record after it. That is the property Article 12 wants and precisely the property an erasure request does not.

  • A record over 1 MiB is refused rather than written, so bulk content is pushed out of the chain by construction. Bytes that big belong in the content-addressed blob store, with only the digest journaled — see the cookbook. Deleting a blob is then a filesystem or bucket operation, and the chain still verifies afterwards because it only ever committed to the digest.

  • Blob bytes can be erased, and the erasure is recorded. BlobStore::expire drops the content and leaves a tombstone. A reader afterwards gets Expired with the date and the reason — not NotFound — because “retention did its job” and “data is missing and nobody knows why” are different answers and only one of them is an incident. Expiring twice keeps the first tombstone, so a retry cannot rewrite when the data went.

On scheduling, and why there is no TTL here. Every object store this runs on already expires objects far better than a sweeper could — S3 lifecycle rules, GCS object lifecycle, Azure blob lifecycle — and they run without your process being alive. Reimplementing that would be a worse copy of a solved problem.

The catch is worth knowing before you rely on it: a lifecycle rule deletes, it does not tombstone. A blob removed that way reads as NotFound, and the distinction between “retention did its job” and “data is missing and nobody knows why” is gone. So use lifecycle rules for bulk age-based expiry where that distinction does not matter, and call expire explicitly for erasure requests, where being able to say when and why is the entire point.

The erasure unit is the case, which is the only unit anybody actually names — nobody asks to forget a digest. Write bytes through cx.store_blob, which records the link at the one moment it is knowable (a digest cannot be reversed to find its case), and answer a request with one call:

let n = agentplane::blob::erase_case(
    Some(blobs.as_ref()), cases.as_ref(), Some(keys.as_ref()), &tenant,
    case, now, "art-17 request",
).await?;

The key-ring and tenant arguments exist with the keyring feature. On a sealed plane, passing the ring destroys the case’s wrapping key after every tombstone is written, so each sealed copy becomes unreadable at once; pass None on a plane that stores blobs unsealed, and the call is plain tombstoning.

Every blob that case produced is tombstoned with the same reason. Other cases are untouched — including ones that stored identical bytes, which land on the same digest by construction, so the link is what scopes the erasure rather than the content.

On a window, rather than one case at a time. The same act on a clock is Runtime::retain(older_than, at, reason), or:

agentplane retention plan --store ./journal.redb --older-than-days 2555

The CLI form lists — the binary wires no store that can erase, so it has no verb that claims to — and the pass itself is the library call. It erases closed cases only, and measures the window from opened_at: a case still open is a matter still running, and erasing underneath a live run turns a retention pass into an outage. Every pass returns not_erasable beside its count, and that list is the half that matters — a number with no coverage statement beside it is how a deployment comes to believe an obligation is discharged.

What still cannot be erased is anything written into a journal record. The chain is append-only by design, which is the point — keep personal data out of records rather than expecting erasure to reach it. Two ways to do that, and both are enforcement points rather than advice: write bytes through cx.store_blob so the chain commits to a digest, and declare a ceiling on what may be written down at all — spec.security.max_sensitivity_journaled in a manifest, or RuntimeBuilder::max_sensitivity_journaled on a plane of hand-written skills. The 1 MiB record refusal pushes bulk content out by construction, but a name, an address and an IBAN are a few hundred bytes and fit comfortably. Erasure has the full table of what lands where; regulation says the same in the obligations’ own terms.

🧯 Disaster recovery

Two numbers and a rehearsal. The numbers are properties of your backup schedule; what this page can state is where each one comes from and what the runtime contributes to it.

RPO — how much history a failure can cost. An effect’s intent is durable before it is dispatched, so the plane itself loses nothing it acknowledged. The exposure is entirely the gap between your last durable copy and the failure, which makes RPO a property of whichever of these you rely on:

CopyRPO
PostgreSQL streaming replication with a synchronous standby~0
PostgreSQL WAL archiving / point-in-time recoverythe archive interval
agentplane export on a schedule, against a shared storethe export interval
Embedded redb file snapshotsthe snapshot interval

An export is not a substitute for the first two. It carries the journal and the case layer and nothing else, by design — see what an operator re-establishes, below.

The export row names its backend because the condition is real: redb admits one writer process, so agentplane export cannot open a file agentplane serve is holding. On the embedded store the scheduled copy is a file snapshot and the export is what you take when the plane is stopped. On postgres the export runs beside the serving plane, so the interval is one you can actually schedule — agentplane export --store "$DATABASE_URL" --tenant acme.

RTO — how long until the plane serves again. Four terms, and only the last grows with how much work was in flight:

  1. Stand the store back up (restore, replica promotion, or agentplane restore).

  2. Verify it. agentplane verify checks an export offline against its own checkpoint; agentplane restore reports whether the rebuilt store commits to the same root at the same size, and exits non-zero when it does not.

  3. Start the plane. Nothing is replayed at startup; a run is replayed when it is resumed.

  4. Re-arm the suspended runs. Resuming each one is what repairs its waits: replay reaches the announced wait, finds no terminal record, and re-arms the timer, re-subscribes and re-opens the task row. Which runs those are is a query rather than something to have kept:

    agentplane waiting --store ./restored.redb      # soonest due first

    Or GET /runs/waiting on a serving plane, under api:run.waiting. Both answer from the journal, not from the timer and subscription tables — which is the whole reason they work here, since an export carries neither and those registrations are exactly what a restore is missing. A run leaves the listing by being resumed, so it is a worklist that drains rather than a page that stops changing.

    What it costs to read. An index maintained by the write path, holding one row per currently waiting run rather than one per suspension ever recorded — so its size is the backlog you are looking at, not the history. Postgres serves it from run_waiting_due in index order with no sort; redb ranges the tenant’s slice in key order and reads nothing past the page. Scanning history for RunSuspended is the answer that does not work, and not for cost: every run that ever waited carries one forever.

The drill

agentplane export  --store ./journal.redb > plane.jsonl
agentplane verify  plane.jsonl --checkpoint ./checkpoint.json
agentplane restore plane.jsonl --store ./restored.redb
agentplane drill   --store ./restored.redb      # every case's references, live

verify given no anchor reports deletion — not checked, and means it: a rebuild against the file’s own header proves self-consistency, which is exactly what an editor who dropped a run and rewrote the header achieves. Keep a checkpoint somewhere the plane cannot reach — an earlier audit’s output, or a witness.

restore writes into the store it is pointed at, which is normally a different tenant of a database that survived. That relabels the log, and the report says so in not_carried — the verdict is the commitment, not the name.

What an operator re-establishes

An export carries the journal and the case layer. Everything below lives in stores it does not travel with, and is listed so that none of it is discovered during an incident:

Not carriedWhat re-establishes it
Timers, event subscriptions, worklist rowsResuming the run. Replay reaches the announced wait, finds no terminal record, and re-arms from the journal. RestoreReport::awaiting names every run the file carried; agentplane waiting names what the plane is waiting on now, and keeps naming it until somebody resumes it
LeasesNothing: a first lease starts one past the highest epoch the run’s own journal records, so a fencing token cannot go backwards across a restore
Webhook delivery cursorsRe-registration, then POST /push/rearm. A cursor is how far a receiver got and nothing journals it
Standing authorityIssuing it again (AuthorityStore::issue). A recreated store holds none, and a draw on an authority it does not hold is refused as AuthorityError::Unknown
Record signaturesThe restoring store attests as its own signer. Hashes and the Merkle root are unaffected; authorship is not
Activity timestampsNothing — the index is rebuilt, the original instants are not
Blob bytesRestoring the object store. The file carries each case’s blob digests, which is what keeps erasure reachable, and never the objects
Key materialYour key management. A sealed plane restored without its ring holds ciphertext, and agentplane drill reports a sealed state that neither opens nor was destroyed as a finding

agentplane drill is what turns the last two rows into an answer rather than an assumption: it walks every case, holds each reference against the live stores, and reports unchecked for a store it was not given. A restore that reads as sound while every artifact is unreachable is the outcome it exists to prevent.

The verdict stays on the plane. An audit does not ask whether you can rehearse; it asks when you last did and whether it passed — and left in the job that ran the verb, that answer is a CI log that rotates. Every drill writes one row, which the next drill replaces:

agentplane drill --last --store "$DATABASE_URL" --tenant acme
# {"drilled":true,"at":"2026-09-20T04:11:02Z","sound":true,
#  "cases":812,"findings":0,"not_checked":1,
#  "origin":"acme-prod","log_size":91_204}

GET /drill serves the same answer under api:drill.read — its own capability, so a read-only auditor credential need not carry the operational roll-up as well. Three things to read off it:

  • drilled: false is a finding, not a clean answer. Nobody has rehearsed this plane is exactly what a missing log cannot tell you from a rotated one, so both verbs exit non-zero on it.
  • origin says which store it ran against, from the plane’s own checkpoint. A drill over a restored copy proves that copy recoverable and says nothing about production; without an origin the two read identically.
  • not_checked sits beside sound, not inside it. A pass over nothing is not a pass — the example above checked cases and was given no key ring, and a reader has to see that rather than infer it.

agentplane attention reports a failed rehearsal as drill.failed, and keeps reporting it until a later drill passes.

A message delivered to a restored plane before its run is resumed buffers, and a buffered message nobody claims dead-letters. So the order is: restore, resume everything awaiting names, then open the gates.

Evidence

The drill is a test, not a procedure somebody remembers. postgres_restores_a_plane_that_then_serves runs the whole sequence against a real PostgreSQL server, restoring into a tenant of a database another tenant is already using, and asserts each claim above: equal roots at equal size, records hash-for-hash, the matter and its obligation and its artifact, isolation in both directions, a lease past the journal’s highest epoch, a wait repaired by a resume — and then the part that separates a recovery from a backup, which is that the restored plane admits new work whose seal extends the log it restored. It ends by drilling the restored case layer without a blob store, because a report that said sound there would be saying it about bytes nobody had put back yet.

🚑 Runbook

SymptomWhere to look
A run is QuarantinedIt holds an effect whose outcome is unknown, or replay diverged. The record names the step. It will not be unwound automatically, and that is deliberate — reversing everything except the thing nobody can account for is worse than stopping
A run seems stuckIt is almost certainly suspended on an event, a timer, or a human. GET /runs/{id} reports why rather than only that
An event was dead-letteredNothing was waiting for it, and the grace window elapsed. GET /dead-letters names it and the keys it was filed under — the mismatch is usually visible by reading them next to what the run subscribed to
A webhook receiver stopped answeringIts registration is parked, not deleted: the cursor survives so nothing is lost. GET /push says which and what it answered last; POST /push/rearm resumes at the first record it never acknowledged
Budget exhaustedA ceiling doing its job, not a fault. The status carries the limit and where consumption actually reached, so it says what to raise it to. A tool’s full rate window reads the same way, and its remedy is time → a tool’s rate ceiling
LeaseHeld vs FencedOpposite responses. LeaseHeld means another instance is alive and you should wait; Fenced means this writer is stale and must drop the run, never retry