agentplane

The journal

An append-only, hash-chained journal with per-record signatures and a per-plane Merkle log — what it proves, and the claims it refuses to make.

The journal is the product. Recovery, audit, cost accounting and regression testing are all reads of one append-only structure, which is why they cannot disagree with each other.

The journal

Append-only, hash-chained, one row per record:

seq | kind              | effect_key | prev_hash | hash
────┼───────────────────┼────────────┼───────────┼──────
  1 | RunAdmitted       | –          | 0000…     | 9c2…
  2 | PlanFrozen        | –          | 9c2…      | 41b…
  3 | StepStarted       | –          | 41b…      | e07…
  4 | EffectStarted     | ek:3f9…    | e07…      | 55a…
  5 | EffectDone        | ek:3f9…    | 55a…      | d13…
  6 | Released          | –          | d13…      | 7f2…
  7 | StepFinished      | –          | 7f2…      | 8b6…

hash = H(prev_hash ‖ record_bytes).

A run that reaches a conclusion appends a RunConcluded record, so how a run ended is covered by tamper detection and a resumed run reads its own outcome from the history it just verified. A side table alone could not answer is this run finished? without inferring it from the last step that happened to finish.

The record keeps machine state machine-readable. An exhausted conclusion carries the typed BudgetExceeded verdict as well as its human reason, so idempotent redelivery returns RunStatus::Exhausted with the exact ceiling and counters; no projection parses prose back into control flow.

A reader knows what it is reading, or refuses

Every record carries a schema version, and it is compared on every read — not only when a parse fails, because the dangerous case is the one that parses. A journal written one shape ahead deserialises perfectly into an older build’s struct, with the fields that build has never heard of dropped on the floor, and every decision downstream is then made over a record nobody fully read.

Both directions are refused, and the policy is the strict one on purpose:

  • A bumped version is refused unless an upcaster can reach this build’s shape from the one on the record. The seam is consulted on every read rather than waiting for its first migration to also be its first exercise.
  • A field nobody bumped for is refused too. Most wire formats tolerate unknown fields, and they are carrying messages. A record is evidence: its fields are the inputs to authorization, retry and recovery decisions, so a reader that drops one reaches a verdict over evidence it did not see and reports it as an ordinary result.

Both are classified as a build skew, never as damage — UnknownRecordVersion for a version this build does not read, UnreadableRecordShape for a shape at the version it does. A rolling deploy that put a writer ahead of its readers reaching an operator as the history has been altered would spend the one alarm that has to stay believable. Each refusal carries its own remedy: which binary wrote the journal, and which one to run.

Before the freeze, the second is the only one you will meet. A shape change here is a hard cut and a hard cut does not bump the version — so the two builds either side of one agree about v and disagree about the shape. The classification is safe to make because of the order a read happens in: the stored hash is checked before the body is parsed, so a record that reaches the parse carries the bytes that were written and the chain commits to them. A failure after that point is a statement about the reader.

An upcast is a read-time view. The bytes the chain commits to are the ones that were written, so tamper evidence does not depend on the age of the reader — and a build that lifts an old record still verifies its link from the original bytes alone.

The bytes are frozen

tests/golden/records.jsonl holds one canonical record per kind and its chain digest; tests/golden/export.jsonl holds a sealed export every build must still verify offline. Reordering two fields or adding a skip_serializing_if rehashes every record this project will ever write and reads in review as a tidy-up — so it is a failing test, and re-blessing the corpus is a command somebody types rather than something a formatter does.

A conclusion is not always a closure. Only conclusions nothing may resume — succeeded, cancelled, abandoned — seal: the journal freezes (the store refuses further appends as a constraint, not a convention) and the run enters the Merkle log below. failed, exhausted, withheld and quarantined leave the run open, because each has a party who can honestly answer it — a resume reads completed effects back rather than performing them again, a raised ceiling continues an exhaustion, a lifted withdrawal continues a withholding, and a person answers a doubt — and a leaf published for a run its own resume may grow would be a checkpoint attesting a prefix of a moving history.

A quarantine is the case worth stating, because sealing one looks right and is not. It is the runtime saying it does not know: the story is not over, and a Merkle leaf claims a history is complete. A sealed chain refuses appends, so sealing a quarantine would refuse the one record that answers the doubt. One chain can therefore carry more than one RunConcluded record, and the last one is the run’s answer; the outcome index the operator queries derives from it in the same transaction, so a failed run that is resumed and succeeds moves between listings rather than being listed as failed forever.

What the chain proves, and what the signature adds

The chain is per run: prev_hash links to that run’s own head, and genesis is zero for every run. On its own it delivers exactly one thing:

No record was edited, reordered, or removed within a run, by anyone who cannot recompute every subsequent hash.

That last clause is the problem. Anyone who can run SHA-256 can rebuild a consistent chain, and the party holding the store can always run SHA-256 — which is the party an auditor is being asked to trust.

So every record also carries an optional KeySignature: a key id and a signature over the record’s chain hash. A hash says what the history is; a signature says who wrote it.

  • One signature per record is enough. The hash already chains, so signing record n’s hash transitively commits to every record before it. Rewriting any part of the prefix invalidates every later signature, not only its own.
  • It sits beside the hash, not in the body. Forced, not stylistic: inside the body, the hash would cover the signature that covers the hash.
  • Verification is lenient by default, strict on demand. A plane resuming its own history has no basis to reject an unsigned record — and a runtime that refused would make signing impossible to adopt incrementally. An auditor has every basis, and require_signature is the difference between “resume my history” and “prove this to me”.
  • A plane with no signer writes unsigned records, not self-signed ones. A self-minted key produces records that look attested and prove nothing, because the party being audited chose the key.
  • The signature carries no algorithm field. A self-described algorithm is how a verifier gets talked into checking a signature with something weaker than the one that made it. The verifier decides what it accepts.

The crate ships the seam (Signer, Verifier) and an Ed25519 implementation behind the signing feature. A deployment with workload identity — SPIFFE SVIDs, which is what the delegation model already assumes — plugs its own signer in, and then the key id on each record names the workload rather than merely a key.

Binding runs to each other

Signing binds authorship. It does not bind existence — and the per-run chain stops at the run boundary, so deleting an entire run leaves every remaining run verifying perfectly. The deleted run’s signatures leave with it, so those do not help either. What is left pointing at it is a case row: ordinary mutable data that goes in the same delete.

So sealed runs enter a per-plane Merkle log (RFC 6962 shape), and the store answers three questions:

  • checkpoint() — origin, size, and root over every sealed run, in the C2SP tlog-checkpoint shape so existing verifiers work.
  • inclusion_proof(run) — that one run’s position and the sibling hashes that prove it.
  • consistency_proof(old_size) — that the log has only grown since an earlier checkpoint.

Delete a run and the root moves; the deleted run can no longer prove inclusion.

“RFC 6962 shape” is two specific things, and both are checkable rather than asserted. Leaf and interior hashes are domain-separated by a prefix byte — without it an interior node’s preimage can be presented as a leaf, and a tree of n leaves reinterpreted as a different tree with the same root. A leaf hash is therefore its own type: a caller who skips the hashing step does not build an undifferentiated tree, they fail to compile. And the hashes themselves are pinned to values computed by another implementation, because a checkpoint is submitted to witnesses running somebody else’s code — a tree that agrees only with itself would pass every test here and be rejected by every witness in the network.

The third of those is what makes the other two mean anything. The root moves on every ordinary seal, so an auditor comparing two roots and seeing a difference has learnt nothing — legitimate growth and deletion-plus-growth look identical. A consistency proof shows every leaf committed to before is still committed to, in the same position. Without it the log detects a change and cannot say what kind.

Two details are decisions rather than implementation:

  • The log position is never a count of what survives. A count reuses a deleted run’s index, so a removed run can be silently replaced at the same position — and even the log size looks unchanged. The store never deletes a seal; a newest seal removed behind its back has its position reissued, which is what a witnessed checkpoint’s consistency proof refuses.
  • The proof does not authenticate its own parameters. An inclusion proof is checked against (leaf, index, size, root), all supplied by whoever offers it; the size and root come from a signed checkpoint. Expecting the fold to validate the size is asking the wrong component, and RFC 6962 has this shape for the same reason.

The audit an outsider runs

Every mechanism above is only checkable. Somebody has to check it, and if the only code that can is inside the runtime being audited, the party under examination is also the party running the examination. So audit::audit runs against a store it did not write, taking inputs the auditor holds:

GivenAnswers
nothingIs each run’s chain internally consistent?
a public keyWho wrote each record?
checkpoints from outsideHas anything been removed since they were issued?

Only the third detects deletion, and only because the checkpoints came from outside. The test that makes this concrete audits a store somebody deleted a run from twice: with no anchor it comes back clean — honestly, because there is nothing to compare against — and with one it fails. A second test does the same thing with the checkpoint fetched from a witness rather than saved earlier, which is the version an auditor can run without the operator’s cooperation.

Bring every checkpoint you can get, and the audit holds the store to each. They are independent observations, so reducing them to the strongest is the one thing that must not happen: an operator who forks a log and has a fresh witness cosign the fork ends up holding the longest history anybody has, and the honest observer’s shorter checkpoint is the only evidence of the divergence. A finding names which anchor the store failed to extend, because that observer holds the history it no longer has. Audit with one anchor and not_checked says what is still unknown.

That asymmetry is why AuditReport carries not_checked as prominently as findings, and why assert_complete fails on a skipped check as well as a failed one. Checkpoint also has a text form in the C2SP note encoding, because the one artifact that must leave the operator’s control cannot exist only as a Rust struct.

The quorum, enforced

  • Witness policy. WitnessQuorum::of(n) declares how many cosignatures suffice, and cosign_quorum holds each submission round to it. Three answers stay distinguishable: met; a shortfall, a finding to clear rather than a log line; and an integrity refusal — a witness that saw this log shrink or fork — reported even when the quorum was met. A run never waits on witnessing: it is retrospective evidence, gathered after sealing. The number itself is a deployment trust decision, and a checkpoint configured but never published is still only as trustworthy as the operator.
  • Split views — one history to one auditor, a different one to another — are refused by a witness that remembers, because the second history cannot prove it extends the first. Independence comes from choosing a witness run by somebody other than the operator; hosting your own proves nothing about you.
  • Where the submission happens. RuntimeBuilder::witnesses(witnesses, quorum) wires the set, and the periodic sweep submits — after sealing, off the run path, and skipped when the log has not grown since the last round that met the bar. SweepReport carries the three answers, so a shortfall reaches needs_attention and an integrity refusal also goes to agentplane.witness.integrity: the audience for a witness says this history moved is not only the operator who runs the plane. Build refuses a quorum larger than the witness list, in both directions — a bar no round can clear reports a shortfall every tick, which is how an operator learns to ignore the one that means something.
  • Where the anchor comes back. WitnessReader::latest(origin) reads tlog-witness’s monitor endpoint, and it is a separate type from Witness because reading is a different party’s action: submitting needs the log’s own signing key, reading needs nothing but the URL and the witness keys the reader trusts. That is the direction an auditor uses. agentplane audit --witness <prefix> --witness-key <name>=<base64> is the same call from the command line, and it is what makes the deletion check runnable by somebody the operator did not hand a checkpoint to. Naming several witnesses buys a check no single anchor makes: a split view is exactly two witnesses holding one tree size with two different roots.
  • The checkpoint is signed per submission, never once at configuration. A witness MUST verify the checkpoint signature against the public key(s) it trusts for the checkpoint origin, and that signature covers the note body — origin, size, root — which changes with every checkpoint. So LogKey holds a signer and a note key name, and a signature is made over each checkpoint as it goes out. A held signature would be correct for exactly one checkpoint and answered 403 Forbidden for every one after it.
  • A 422 is three different answers. The specification gives the status a size-zero checkpoint whose root is not the empty tree’s, a consistency proof that does not verify, and equal sizes with unequal roots. The first two are the client’s own inputs — refused here before a request goes out — and only the last is evidence about the log, because no proof-building mistake produces two roots for one size. A 422 on growth is WitnessError::Inconsistent, which names what is actually known: either the proof this log built is wrong, or its history moved.
  • Cosignatures are verified, not counted. HttpWitness::new takes the TrustedWitness keys a deployment accepts and refuses to build without at least one. Each signature line on a 200 is matched to a trusted key by name and four-byte note key id — signed-note’s conjunction, because a name is whatever the answering server typed — then checked as a C2SP cosignature/v1 statement: a big-endian timestamp leads the payload, and the signature covers the cosignature/v1 header, the time line, then the note body that was submitted. The header is what separates a witness’s observation from a log’s own note signature — same algorithm, same key length, different claim. A zero timestamp is refused: the specification says the cosignature MUST NOT omit the timestamp, and the observation instant is what separates a witness that is watching from one that answered once and stopped. The construction is pinned to the spec’s published example rather than to a round trip, since a signer and verifier written from one misreading round-trip cleanly. A quorum is otherwise a count of HTTP status codes, and every guarantee resting on an independent party observed this log would be a guarantee about string formatting.

Both backends maintain the log. redb advances a counter row inside the sealing transaction. Postgres allocates the next position under a per-tenant transaction lock, because several instances seal concurrently there, and a position must follow commit order: a position taken by a transaction that commits after a later one was checkpointed would land a leaf inside that checkpoint, and a witness would see a fork.

Positions have holes once a run is removed, deliberately. The tree is built by walking the log in key order, which yields dense positions with no holes — so the position a proof uses is the run’s rank in that walk, not its stored index. Handing back the stored index makes every run after a deleted one fail to prove an inclusion that is perfectly valid.

Hash the bytes you wrote

Records are hashed over their exact wire bytes, and those bytes are what the store keeps. Verification never re-serializes.

This matters when schemas evolve. If the chain were computed over the upcast form, then the first time a record shape changed, every historical hash would change with it — silently destroying tamper evidence for all past records, which is the one property the chain exists to provide. Upcasting is a read-time view; the chain is over history as written.

Schema evolution

Until the format freeze a shape change is a hard cut. The freeze commits the journal to evolving without rewriting history:

  1. Records carry (kind, v), and a shape change bumps v.
  2. Every shape ever written stays readable. Upcast on read; never rewrite.
  3. Upcasters are pure and total. Same input, same output, in this process and in one started a year from now.
  4. Hash the wire bytes (above).