agentplane

Erasure and keys

Cryptographic erasure that reaches backups, envelope encryption, key rotation and revocation, and the tenancy boundary.

On this page
  1. What lands where, and what can be erased
    1. A digest is not erasure
    2. Erasing on more than one instance
  2. Where a subject’s data went, before erasing it
  3. Retention on a window, and what it cannot reach
    1. A disclosed copy
    2. A legal hold is the one thing that stops a pass
    3. An expired address stays expired
    4. A tombstone is read or refused, never guessed
    5. not_erasable is the half that matters
    6. Admission keys are their own window
  4. Erasure that reaches the backups
    1. Rotation: sealed bytes never change
    2. An envelope says which construction it is
  5. Tenancy

An erasure obligation is discharged by making data unreadable everywhere, not by deleting the copy you can reach. This page is how that is done here, and what the boundary actually is.

For the trust model those keys sit inside, see security.


What lands where, and what can be erased

Decide this before personal data reaches a run. The journal is append-only and hash-chained: a record cannot be deleted without breaking the chain, so nothing that enters it is ever removed. What decides whether it can be erased is whether a key ring is configured — without one the bytes are permanent for the life of the store, and with one they are ciphertext whose key erase_case destroys.

The Erasable column below is therefore two answers, and the difference is one builder call:

DataWhere it landsErasable without .keyring(..)with it
Run inputjournal — RunAdmitted.inputno, verbatimsealed, per-case scope
Data subjects a run bindsjournal — DataSubjectBound.subject; labels carry only {run, index} references to itno, verbatimsealed, per-case scope
Model promptsjournal — inside EffectStarted.descriptor.args, because the prompt is part of effect identityno, verbatimsealed, per-case scope
Tool call argumentsjournal — same field, same reasonno, verbatimsealed, per-case scope
Effect outputs — completions, tool resultsjournal — EffectDone.output, and a reconciliation probe’s EffectReconciled.output, which is the same datano, verbatimsealed, per-case scope
Failure messages, notes, frozen plansjournal — EffectFailed.error (the message only), a compensation’s StepCompensated.outcome (its error text when it failed), Note.text, PlanFrozen.plan, which embeds the trusted input the plan was compiled fromno, verbatimsealed, per-case scope
Case state writes, status changes, deadline transitionsjournal (the effect) and the case storeno in the journal; the case store’s copy is overwritten, not erasedsealed in both, one scope
Inbound event payloads — a counterparty’s message bodyjournal (the awaited effect’s output) and the event storenosealed in both; the buffer’s copy is its own unit, keyed (source, id)
Human task proposals — Justification.proposed_action, the exact thing a reviewer is shown, and evidence, the trail behind itjournal (the task effect) and the task storenosealed in both; the task’s summary stays readable so a queue stays usable, and every entry keeps its trust label so a reviewer still knows who wrote a line whose words are gone
Memory item contentMemoryStoreyes — forget, forget_cascading, expiry sweep, in the live storeunreadable everywhere with an explicit EncryptedMemoryStore wrap: each item version has its own key, destroyed by the verb that erases it — not covered by .keyring(..), see below
Semantic-index vectors, derived from memory contentthe SemanticRetriever’s indexyes — every memory erasure verb tells the index what it removed (SemanticRetriever::forget), through IndexedMemoryStorea key ring does not reach them: an embedding is computed from plaintext, and a vector hidden from search but left in the index is reconstructible content
Blob bytes — cx.store_blob, fetched mediablob storeyes — expiry leaves a tombstone, in the live store onlyunreadable everywhere, backups included
Correlation keys, deadline names, task summaries, statusescase / task storeno, and deliberately — they are what the store is asked questions aboutunchanged: still readable
Admission keys — a message’s source/idjournal — RunAdmitted.idempotency_key — and the run_admission indexno, and deliberately: the index is looked up by a value the caller holds in the clear, and sealing it would leave a store that cannot refuse a redeliveryunchanged: still readable
Blob digest and classificationjournalno — a digest hides a high-entropy value, not a guessable one (below)unchanged

Read the first column as the floor and the second as what one call buys. Both are honest positions: a deployment that would rather refuse the data than seal it declares max_sensitivity_journaled and never reaches this table.

Without a key ring the rule is: keep erasable personal data out of the chain. Put the bytes in a blob and let the journal commit to the digest:

let digest = cx.store_blob(document_bytes).await?;  // erasable, linked to the case
// The chain records the digest and the classification, never the bytes.

cx.store_blob is the only way a value reaches a blob, and it is size-independent. There is no size-triggered spill: the 1 MiB Record::MAX_RECORD_BYTES ceiling refuses a record and names the limit, it does not blob for you. Anything under it is journaled inline whatever it holds.

That is deliberate. A size-triggered spill would make erasability depend on how long a value happened to be — the same field permanent for one customer and erasable for another. So references are the intended shape, not a workaround: journal a digest or an identifier, and fetch details through an authorised tool call. A digest stands in for the bytes only when the bytes cannot be guessed — see a digest is not erasure.

Prompts are the hard case, because the prompt is effect identity — replay reconstructs a run by re-deriving the same key, so a prompt cannot be redacted after the fact without making the run unreplayable. A prompt built from a customer record puts that record in the chain permanently. Pass a reference — a blob address, or an identifier the tool resolves — and materialize the bytes at dispatch, as governed media already does, or treat the run as retained data with a retention period on the whole store.

A digest is not erasure

Every digest here is unsalted SHA-256, and it stays in the clear after the key that sealed its plaintext is destroyed. Whoever holds the record can hash every candidate and compare. That hides a value only when the candidates cannot be enumerated:

Digest overAfter erasure
A document, a free-text prompt, any long unpredictable byteshidden: too many candidates to try
A yes/no, an amount, a date, a status, a name from a known listrecoverable by trying every candidate

This concerns sealed planes: on an unsealed one nothing is erased, and the arguments sit verbatim beside their key.

Where the runtime leaves such a digest, by record kind:

Record kindClear digestSurvives the erasure of
EffectStarted, EffectDone, EffectFailed, EffectReconciled, PolicyDenied, BudgetRefused, BudgetReadmittedthe record’s effect key, over the effect kind and its canonical arguments; step, phase and attempt are clear beside it, and the ordinal is smallthe arguments — and, for a memory.remember effect, whose arguments carry the item’s content digest, forget, erase_subject and sweep_expired too
Releasedvalue, over the canonical released value, and the effect key, whose arguments embed it; both are copied into the audit report and every exportthe released value

RunAdmitted (idempotency_key), CaseBound (correlation) and RunSuspended (the correlation keys it waits on, including a task’s) carry values clear by design rather than digests: they are the keys the stores are asked questions about, so a business key that is personal data is a key chosen badly. RunAdmitted (governed_by, policy_bundle), DeadlineRegistered (calendar_digest) and RunConcluded (chain_head) carry digests over nothing erasable — the deployment’s declarations, a calendar, and a chain over ciphertext, as every record’s hashes are. QuotaPassStarted, PlanFrozen, StepStarted, StepFinished, DeadlineTransition, Note, AuthorityWithheld, AuthorityRestored, IdentityBound, DataSubjectBound, GroupOpened, GroupSettled, QuarantineDecided, StepCompensated, RunCancelled, BreakGlass, HaltLifted, HoldReleased, Swept and Observed carry no digest.

Outside the journal:

DigestWhere it is clearSurvives the erasure of
dispatch identitya callee’s deduplication key, and the A2A message id derived from it — the callee’s copythe arguments
provenance claimthe signed claim a peer receives, carrying the arguments’ digest — the peer’s copythe arguments
reason digestevery loud event in the deployment’s logs, unkeyed and truncateda failure, quarantine or compensation reason the journal seals
blob content digestthe case’s blob rows, the blob’s tombstone, the backing address derived from it, and the key of any effect whose arguments carry itthe blob’s bytes
selection digesta Selected in a semantic retriever’s index, where the deployment keeps onethe memory item’s content

An effect whose arguments are one small value — {"approved": true}, an amount — leaves that value recoverable from its key. When such a value must be erasable, pass an opaque identifier the tool resolves instead of the value. Moving it into a blob or a memory does not help: both leave a digest of their content in the clear. Otherwise, accept the value as retained.

Declare the ceiling and the runtime enforces it. spec.security.max_sensitivity_journaled refuses, at dispatch, any argument more sensitive than a deployment is willing to make permanent — so the mistake is a refusal naming the blob pattern rather than a discovery at the first erasure request:

spec:
  security:
    max_sensitivity_egress:    secret     # what may leave
    max_sensitivity_journaled: internal   # what may be written down forever

Absent means unbounded.

A plane of hand-written skills declares the same ceiling in code, since it runs under no manifest:

Runtime::builder(store)
    .max_sensitivity_journaled(Sensitivity::Internal)
    .skill(Triage)
    .build()

Where both are present the stricter binds, the same rule a reviewed tool grant follows: a declaration may only tighten what the deployment allows. It is an enforcement point rather than a warning — a build-time lint that let the run proceed would be the advisory control this format refuses everywhere else.

Or seal the journal itself. SealedJournal::wrap(store, keys, tenant) seals the payload fields — run input, effect arguments (prompts, tool calls), effect and reconciliation outputs, failure messages, notes and frozen plans; the caller’s data, never the runtime’s routing — under the same per-case scope erase_case already destroys, so a single erasure reaches blobs and journal alike. The envelope’s associated data binds the tenant and the record’s identity, so ciphertext moved to another record fails to authenticate rather than opening as somebody else’s payload:

Runtime::builder_on(store)
    .keyring(keys)   // seals all of them, and blob payloads
    .build()

Run it: cargo run --example sealed_run --features redb,testkit,keyring writes a claimant’s details, shows them sealed in the store, erases the case, and verifies the chain afterwards.

One call, one guarantee. The wrapping happens at build(), so the order you write the builder in cannot lose it — registering a store after the key ring seals it just the same. An operator Outbox‘s store is swept up too: it keeps the destinations’ webhook bearer tokens — credentials like any caller’s — so SealedPush wraps it in the same pass. The decorators (SealedJournal, SealedCases, SealedTasks, SealedEvents, SealedPush) are public for embedders wiring stores by hand, but a plane should not need five correct decisions where one will do: a control that can be forgotten five times is one where forgetting looks exactly like remembering.

Governed memory is the one store this call leaves alone, and the exclusion is forced rather than an oversight. EncryptedMemoryStore serialises subject erasure against writes and legal-hold changes with a process-local mutex, so it holds its contract on a single-writer deployment and nowhere else. Wrapping it automatically would hand that adapter to an active-active PostgreSQL plane, where the mutex coordinates nothing and the hold race it exists to prevent is the result — a control that is worse than absent, because it reads as present. Its erasure unit differs too: each version of each item is its own, tenant/memory-item/<id>@<version>, and none of them is a case, so erase_case was never the act that reaches it. Erasing an id — by forget, the expiry sweep, or a cascade that reaches its current version — destroys the keys of every version it held; a cascade that reaches only a superseded version (a rolling summary whose v1 absorbed the erased source) destroys exactly that version’s key, so a backup taken before the cascade no longer opens it while the current version stays readable. Cascade.trimmed names each such id with the versions it lost. The sweep and a cascade erase rows before keys, so a key the ring refuses is owed and destroyed by the next erasure verb, which fails while it cannot; the debt is held in the process. A subject is not a key scope — erasing one destroys the keys of the items it held and leaves the subject writable under new ids. Wrap it where you can see the deployment:

let memories = EncryptedMemoryStore::new(inner, Arc::clone(&keys), tenant.clone());
Runtime::builder(store).memory(Arc::new(memories)).keyring(keys).build()

A semantic index is erased with the memory it was built from. Key destruction cannot reach it — an embedding is computed from plaintext — so every erasure verb of IndexedMemoryStore tells the SemanticRetriever what it removed, and an index that could not be told fails the erasure rather than letting it report success. What it missed is kept and delivered by the next erasure verb, the expiry sweep included, because retrying a cascade or a sweep finds the rows already gone; the debt is held in the process, so a restart before that loses it. semantic_memory(..) wraps the plane’s memory in it, which covers the expiry sweep; erasures made on your own handle reach the index only through the wrapper. erase_subject is the sealed store’s own verb and runs beneath anything the plane wraps around it, so the wrapper goes inside the seal, over the same retriever semantic_memory(..) is given — and build refuses a sealed store whose subject erasure would miss the index (BuildError::SealedMemoryMissesIndex), and a store already wrapped over a different retriever (BuildError::MemoryIndexedElsewhere):

let indexed = IndexedMemoryStore::new(inner, Arc::clone(&retriever));
let memories = EncryptedMemoryStore::new(Arc::new(indexed), Arc::clone(&keys), tenant.clone());

Erasing on more than one instance

Erasing a subject is not one operation. It reads the subject’s items, checks every legal hold, asks a KMS to destroy each item’s scope, and only then removes the rows — and between the hold check and the destroy, a write on another instance can add an item, or an operator can place a hold on one. Either makes the erasure wrong in a way nothing detects: the new item outlives an erasure that reported the subject gone, and the held item is destroyed anyway.

A held item refuses the whole erasure with StoreError::UnderLegalHold before any key is touched. A failure to remove the rows after the keys are gone is not reported as success: erase_subject returns an Erasure whose cleanup_failed names it — the erasure happened, no copy opens, and retrying the call removes the rows. SealedEvents::erase_event answers the same way.

EncryptedMemoryStore closes that window with a process-local mutex, which is correct on a single writer and silently nothing on an active-active plane. The lock is a seam, and the default is honest about being local:

// Single node: the default. `is_distributed()` answers false.
let memories = EncryptedMemoryStore::new(inner, keys.clone(), tenant.clone());

// Active-active: a session advisory lock in the database the plane shares.
let memories = EncryptedMemoryStore::new(inner, keys.clone(), tenant.clone())
    .coordinated_by(Arc::new(store.erasure_coordinator()));

Why a session advisory lock and not a row. A row taken with SELECT … FOR UPDATE needs its transaction held open for the whole erasure, and the erasure’s own writes go through the store’s other connections — so the row lock would be held by a transaction that cannot see the work it protects. A session lock is held by the connection, and PostgreSQL releases it when the session ends. That last property is the one that chose it: an instance that dies mid-erasure releases by dying, where a lease with a TTL must choose between stranding the subject and handing it over while the first instance’s KMS call may still be in flight.

Dropping the lease releases the scope. The usual argument against an RAII guard — releasing a distributed lock is async and fallible, Drop is neither — holds for a lease table and not for either primitive here: a process-local mutex releases by dropping its guard, and a PostgreSQL session advisory lock ends with the session, so dropping the connection releases it. Both are synchronous and cannot fail. That is what makes a cancelled erasure safe: put a timeout around a memory write and you abandon the work, not the subject.

Use under_lock anyway. It releases on the success and the failure path, and what it adds is the tidy half — an explicit unlock, so the scope frees immediately rather than when the connection finishes closing, and a release failure that reaches somebody. A coordinator of your own whose release needs a round trip carries no guard and gets neither.

The scope is per tenant, not per subject: forget, forget_cascading and set_legal_hold are addressed by item id and sweep_expired spans every subject, so choosing a subject-scoped lock would need a read that races the very thing the lock protects.

The pairing is refused at build. Both facts are in hand there: the journal store says whether two instances can write it (JournalStore::is_shared), and the memory store says whether its lifecycle lock spans them (MemoryStore::erasure_is_distributed). A shared store beside a process-local lock is BuildError::ErasureCoordinatorNotShared, naming the fix.

is_shared has no default, deliberately. A default of false would let an embedder’s shared backend answer single-writer by saying nothing, and a control that fails open when an implementer forgets is a property the runtime relies on rather than checks. erasure_is_distributed does default — to None, meaning “no lifecycle lock because there is no cryptographic erasure here”, which is the honest answer for an ordinary store and leaves the check inapplicable rather than satisfied.

Two properties make it worth having rather than merely present. Only the payload is sealed: seq, run, case, effect_key and the record’s own variant stay in the clear, so exactly-once, the case scan, the outcome index and the chain keep working with no key at all. And the chain commits to the ciphertext, so an auditor holding no keys still verifies the history of a run whose payloads have been erased — hashing the plaintext would have tied the tamper evidence to the key and destroyed both together.

A payload that will not open means one thing only: the key was destroyed. Every surface that reads sealed data leaves such a payload sealed rather than failing — an erased matter must not make a worklist, a case listing, a dead-letter list or an audit unreadable. Nothing else is forgiven. A key service you cannot reach, a wrapping-key version you retired, an envelope written by a newer build, and a payload that does not authenticate are each raised as themselves, because reported as erased they tell you data is gone that is sitting intact in your store. The sharpest case is a webhook credential: dropped, the notification goes out without the authentication it was registered with, so a KeyError::Unavailable there is a read that fails and a delivery that waits.

A run whose payloads were erased says so, and you can still close it. Resuming one answers RuntimeError::PayloadsErased rather than a parser’s complaint about a missing field — the journal is intact and still verifies; the data is gone because somebody asked. Such a run can never execute again, and it is still concludable: POST /runs/{run}/abandon (or Runtime::record_quarantine_decision with QuarantineDecision::Abandon) is recorded in the clear, needs no plan, and ends it. Abandonment rather than cancellation, because a cancellation promises to put the world back and the arguments it would need are what was erased.

Events are the one thing not erased by the case, and that is forced. An event is buffered before any subscription matches it, and one nobody claims becomes a dead letter — which by definition matched no case at all, and is kept indefinitely so an operator can find the wrong correlation key. There is no case to erase it with. So the event is its own unit, scoped by the (source, id) pair the buffer already deduplicates on, which is the finest granularity a request about one message could ask for:

keys.destroy(&event_scope(&tenant, "bank.example", "MSG-7"), at, reason).await?;

event_scope puts the length of source in front of the pair. A source is a URI and contains /, so a plain source/id join would give ("bus/x", "1") and ("bus", "x/1") one key, and erasing either would destroy both.

Its delivered copy is a separate matter and already covered: claiming an event journals the payload under the awaiting effect, sealed under that run’s case like every other journal payload.

An erased message is never delivered. Erasing a buffered event nobody has claimed moves it to the dead-letter list with the reason erased, in the same write that removes its payload, so no waiter claims the emptied row as the counterparty’s word. A message a run had claimed but not yet journaled — a delivery that crashed between its claim and its resume — is treated the same way: dead-lettered as erased, its claim released, and the wait it was parked with unparked. The wait stays open, bounded by its deadline, and the next matching message is delivered to it; the run is never handed an empty payload as a value, and no failure is recorded on its behalf. Every erased row is marked erased and no claim returns one. A sealed buffer checks the same on the way out: a claimed message whose key was destroyed is not handed to the run.

The composed claim is tested, not merely asserted. One erasure reaches every copy is a sentence about three mechanisms sharing one scope, and composition is where that kind of claim breaks — each decorator can be correct alone while sealing under a scope erase_case never destroys, which would report a successful erasure over readable data. A test wires all of it together, erases the case once, and asserts blobs, journal payloads and case state are all unreadable afterwards while the hash chain still verifies. A mutation that gives the case store a different scope is caught by that test and by nothing else.

Correlation keys, deadline names, task summaries and statuses stay readable everywhere by design — they are what the stores are asked questions about. A deployment that considers its business keys personal data must choose keys that are not. A deployment whose obligation covers them either erases the case — which destroys the scope key and takes the journal copies with it — or keeps the data out, which is what max_sensitivity_journaled refuses at the boundary.


Where a subject’s data went, before erasing it

agentplane subject <subject> --store … [--tenant …] [--limit N] [--json] answers the question that comes before an erasure. It selects the subject’s memory items exactly as an erasure does — every id the store holds for the subject, expired or not — and the runs whose DataSubjectBound names the subject, with their cases: the units erase_run and erase_case act on. It lists each outbound effect whose label names one of those items in its provenance, or one of those bindings in its data_subjects: run, case, effect key, kind, sink and bytes. A value is reported as influenced by an item or a bound run’s intake, never as containing it. It also lists the recall records that read an item, trusted or not, when their output can be read.

A run binds its subjects from its agent’s spec.data_subjects or RunTerms::subject, and the binding follows its input, its inbound events and its case-state reads to every effect they reach — into a commissioned run, too. The subject is sealed with the run, so an erased binding names nobody and its DataSubjectBound keeps no identifier in the clear. That is a claim about the binding, not about every place the identifier may also sit: a data subject is never read from a correlation key, because the key stays readable in every case record by design (above), and a deployment that also correlates on the identifier leaves it there whatever the binding says. A reference the runtime did not attach is a claim it cannot check: an embedder or a batch that admits a run with a Tainted input whose label already names references is believed, since no record says which run those references came from.

The report is a read. It appends no record, writes no memory, and takes no operator. It is evidence for whoever asks, not a statement that any access obligation is met.

What it cannot see is printed first, on every report, with each class marked when this report met one: a run that binds no data subject; a sealed binding with no key ring; an erased binding; data about the subject from a tool, model or peer, traced only once joined with what a bound run took in; a trusted item, whose recall leaves no source in any label; recall records that could not be opened; forgotten and swept ids; sinks whose arguments were not opened, named by kind only; the trace being per item rather than per version; a coded step that passes a value on without joining its inputs’ labels; and runs past the limit or unreadable, which also exit 5.

The terminal carries no key ring, so on a sealed plane it names a tool call’s sink by its kind and lists no recall and no bound run. The library’s subject::Trace::with_keys opens all three while the keys stand; a destroyed key is reported as erased, and any other ring failure fails the report rather than reading as an erasure.


Retention on a window, and what it cannot reach

erase_case and erase_run erase one unit. Retention is the same act on a clock: a window, applied to every closed case, by the plane.

erase_case refuses a matter that is still open, for the same reason the pass only selects closed ones: destroying the key freezes every run the case covers. Their plans and effect outputs stop opening, so none of them can be replayed, resumed or unwound again — the work is over whether or not anybody decided it should be. Conclude the runs, close the case, then erase. The refusal is typed (EraseError::CaseStillOpen) and nothing is destroyed before it.

let report = plane.retain(cutoff, now, "retention: 7 years from opening").await?;
agentplane retention plan --store ./journal.redb --older-than-days 2555

The verb plans; it does not erase. The shipped binary wires no blob store and no key ring — a redb file is a journal and a case layer — so nothing in it can make a byte unreadable, and a verb that walked the cases and printed erased: 0 beside a clean exit code would be a control that reads as having run. It answers the half it can, through the same selection rule the pass uses (retention::plan, so a listing and an erasure cannot disagree about which cases). The pass itself is Runtime::retain, from a plane built with the stores that can act.

A pass erases closed cases only, and the window is measured from opened_at. A case still open is a matter still running, and erasing the data underneath a live run turns a retention pass into an outage. opened_at is the anchor because retention rules are written as N years from the start of the business matter — and it is the only instant a case records, which is a fact this verb states rather than works around.

--older-than-days has no default, for the reason forget-admissions has none: a retention period is a legal and business decision, and a crate that picked one would be choosing somebody else’s. The pass’s reason lands on every tombstone and key destruction, so a later read says expired, on this date, for this reason rather than missing — which is the distinction the recovery drill’s verdict is built on.

A disclosed copy

A disclosure package is a copy outside the plane, and no erasure reaches it. What an erasure can do is name it: erase_case and erase_run return, and a retention pass reports in disclosed, one line per disclosure of what was erased — who received it, when, who sent it, and what the erasure does to that copy:

  • from an unsealed plane: a plaintext copy this erasure cannot reach;
  • a copy carrying sealed payloads: those sealed under the erased key become unopenable; anything that travelled unsealed — another case’s payloads in the same package among them — and the fields the journal keeps in the clear stay readable in it. Whether a copy is sealed is recorded per package, not per case.

A retention pass walks cases, so a case-less run’s disclosures are named by erase_run, not by the pass. The names come from the plane’s disclosure register, a row per disclosure in the store beside the holds — not a record on any chain. Whoever administers that store can edit or delete a row, so a line names what the operator recorded, and an empty list does not show that nothing left the plane. Without a register wired, the result says none was consulted.

A retention pass is automatic: it runs on a window nobody re-reads, and a matter under a preservation order looks exactly like every other closed case old enough to go. A hold is what makes that erasure fail instead.

# Preserve a matter. --reason and --actor are both required: one is what
# somebody reads in 2028, the other is who they can ask about it.
agentplane hold --store ./journal.redb \
  --case case_01JD... --reason "preservation order 2026-114" --actor compliance-dana

# Everything standing, for somebody who does not know which case to ask about.
# Each row carries who placed it and whether a credential or a terminal said so.
agentplane hold list --store ./journal.redb

# Release it, under a name. The next pass erases the matter normally.
agentplane hold --store ./journal.redb --case case_01JD... --lift --actor compliance-dana

# Every recorded release, newest first.
agentplane hold list --store ./journal.redb --released
let by = Operator::authenticated(caller.actor)?;   // or `asserted` at a terminal
cases
    .place_hold(case, &LegalHold { placed_at: now, reason: order.into(), by })
    .await?;

Run it: cargo run --example retention_hold --features redb

Nothing is half-done. The hold is read before the first tombstone and before any key is destroyed, so a refused erasure leaves no expired blobs and no missing key. Over HTTP it is GET/POST /holds and POST /holds/release, under api:hold.list, api:hold.place and api:hold.release — placing and releasing are separate, so whoever may authorise destruction is a grant you hand out separately from whoever may prevent it.

A release is recorded, and survives the erasure it permits. Before the row goes, the release is written to the journal as the one HoldReleased record of a sealed run of its own, outcome hold-released, stamped with the case: who released it and on what basis — the authenticated caller over the API, --actor as asserted at a terminal, where it is required — when, and who had placed the hold and when. It carries no hold reason, which is free text about the matter; names and instants only, so it is still readable in the matter’s history after the erasure destroys the case’s key. Only the hold the record names is removed: one re-placed between the read and the removal stays standing, and the release fails naming the record’s run. GET /holds?state=released and hold list --released read the releases back, newest first.

A held case reads as held, not as failed. RetentionReport::held is its own field beside failures, and a hold does not make is_complete() false — a control doing its job must not page anybody. retention::plan splits due from held too, so a dry run and a pass never disagree. The hold is read again inside erase_case, which is what honours one placed between the two.

What else could destroy it. Two things here destroy on a schedule, and a hold stops both: retention::retain (closed cases past a window) and MemoryStore::sweep_expired (memories past their effective expiry). Nothing else scheduled destroys anything — forget_admissions retires index rows and alters no history, sweep_unclaimed moves an event to the dead-letter listing with its payload, and journal records are append-only. A hold travels with its case through agentplane export and restore — instant, reason and operator — so the first retention pass on a recovered plane finds it. The deliberate verbs — erase_case, forget, forget_subject — all consult a hold; erase_run does not, because its unit is a run bound to no matter.

It says what is preserved now, not what ever was. Releasing leaves no history here. Kept rows would make their free-text reasons outlive the matter they preserved, and destroying them with the matter would destroy the record of the erasure’s own authorisation; the chain of custody belongs to the system that issued the order.

What it does not cover. A hold is placed on a case, the unit a preservation order names. Bind a run to a case if it must be preserved — that is also what gives it an obligation trail. Memory items carry their own finer hold (MemoryStore::set_legal_hold, listable with legal_holds), because memory is erased by item and by subject rather than by matter; an erasure it stops answers StoreError::UnderLegalHold naming the item.

An expired address stays expired

A blob’s address is its content, which makes one rule non-obvious and load-bearing: a write to an erased address is refused. Without it, a run producing the same bytes a second time lands on the erased object and puts the data back — silently, under a tombstone that still records when and why it went. That is not an exotic sequence. A resumed run re-stores what it stored before; a second run of the same matter does the same work.

The refusal is typed, StoreError::BlobErased, because retrying cannot help and the store is healthy. A skill reaching cx.store_blob gets it; a governed media fetch gets Rejected rather than the try again classification, so it stops re-fetching a URL whose artifact somebody removed.

A different matter storing the same bytes is unaffected: the erasure unit leads the storage address, so those are two objects. On a sealed plane the rule is belt and braces — erase_case destroys the scope’s wrapping key, so the write fails before it reaches the store at all — and on an unsealed one it is the whole guarantee.

A tombstone is read or refused, never guessed

A tombstone is the only evidence an erasure happened that outlives the bytes it describes, so it is a durable format and carries its own version like every other one here. A reader that cannot interpret one answers BlobError::UnreadableTombstone — a fourth state beside missing, expired and altered — and the recovery drill reports it as a finding rather than counting an erasure it never saw.

not_erasable is the half that matters

Every pass returns a coverage list beside its count, for the same reason DrillReport carries not_checked:

{
  "scanned": 1204, "erased": 318, "blobs_expired": 892, "failures": [],
  "not_erasable": [
    "no key ring is wired: blob tombstones cover the live store only, and journal payloads — run input, prompts, tool arguments, effect outputs — stay verbatim and permanent…",
    "journal records are append-only: the chain, the routing fields and the fact each run happened remain — by design…",
    "a run that belongs to no case is not reached by a case walk; erase one with `blob::erase_run`",
    "governed memory is erased by item and by subject, never by case — including memory a declaration keyed to `$case` or `$correlation/<namespace>`…",
    "an inbound event is its own erasure unit: the event buffer's copy, and every backup of it, is erased by `SealedEvents::erase_event`…",
    "media fetched under a named external retention policy belongs to that policy's unit, not the case…",
    "semantic-index vectors are derived from memory content and erased with it, by the memory erasure verbs through `IndexedMemoryStore`…"
  ]
}

A number with no coverage statement beside it is exactly how a deployment comes to believe an erasure obligation is discharged while the chain still holds the payload verbatim. Without a key ring, retention tombstones blobs in the live store and nothing else. Without a blob store the pass still runs — the unit is the key, and a plane that seals its journal and stores no blobs is an ordinary shape — and the tombstones it could not write are named. With one, the case’s key scope is destroyed and the erasure reaches every replica and every backup at once — because what was destroyed was never in them.

Admission keys are their own window

agentplane forget-admissions --older-than-days N retires the idempotency index, and it is a different decision from the one above: retiring a key reopens the door it closed. The window must exceed how long your emitter keeps retrying a delivery it has not seen a 2xx for, or a redelivery arriving after retirement admits a second run — the failure the key exists to prevent, delivered on a timer.

There is no default, deliberately. Other durable runtimes bound this for you — Restate expires an idempotency key a day after the invocation completes, Temporal’s dedup window is its namespace retention — and both are choosing a retry horizon on your behalf. Pick yours from the emitter, not from the index’s size:

# A webhook source that retries for 72 hours; a week is comfortably clear of it.
agentplane forget-admissions --store ./journal.redb --older-than-days 7

Absent a call, keys are kept forever. That is the safe default — the one that cannot silently admit a duplicate — and the size of the index is a fact your database monitoring already reports.

Erasure that reaches the backups

Dropping a payload’s bytes leaves the hash chain intact, because the chain only ever committed to a digest. That is the right shape and it is not sufficient: it erases the bytes in the live store. Every backup taken before the request, every replica, every snapshot nobody remembers still holds them.

Chasing those copies does not work. Backups are offline, offsite and frequently immutable by design — that is what makes them survive the incident they exist for — so a retention story requiring rewritten backups puts two guarantees in direct conflict, and whichever loses, loses silently.

So payload bytes are sealed under a data key, and the data key is wrapped by a key this crate never holds. Erasure destroys the data key, and every copy of the ciphertext becomes unreadable at the same instant — including the ones nobody can reach, because what was destroyed was never in them. A test restores a backup taken before the erasure into a fresh store and asserts it stays unreadable; that is the property deletion cannot provide.

Two operations fall out of one structure, which is the argument for it:

Erasuredestroy a scope’s wrapping key — every payload ever sealed under that scope is unreadable, everywhere
Revocationthe same act for a different reason; the blast radius a compromised key should have is exactly the scope it wrapped

Rotation: sealed bytes never change

Envelope encryption is usually sold on a third operation — rotate the wrapping key, re-wrap the data keys, leave bulk data alone. agentplane does not offer it, and the omission is a decision rather than missing work.

An envelope carries its wrapped data key inline, and the journal’s hash chain commits to the envelope bytes. That is what lets an auditor holding no keys verify a run whose payloads have been erased — and it is exactly what makes re-wrapping impossible: rewriting a journal payload’s envelope rewrites a record the chain covers, so it breaks the chain it sits inside. Re-wrapping the other stores would not help either, because a scope’s journal payloads and its case state share one wrapping key: the scope stays pinned to the oldest version any of its journal envelopes names.

So the rule is stated instead: sealed payload bytes never change, and the erasure scope is the rotation unit. A scope is already narrow — one case, one run, one memory subject — so a compromised wrapping key exposes that unit and nothing else, which is the blast radius rotation is bought for in the first place. Adding a key version is safe and needs nothing from agentplane; envelopes sealed before a rotation keep opening.

This is the model AWS KMS already assumes. KMS rotates a key’s backing material on a schedule, keeps every previous version in perpetuity, and picks the right one from the ciphertext on decrypt — you cannot select a version, and you cannot delete one. The only way to remove old key material is to delete the whole KMS key, which is precisely erasure. So on KMS, rotation is automatic and transparent, and nothing below can arise.

Vault’s transit engine is the one to be careful with, because it offers a lever KMS does not: min_decryption_version refuses to decrypt ciphertext below a floor. Since envelopes pin their key version for life, raising that floor past a live envelope makes un-erased history unreadable — an erasure nobody requested, that no retention record explains and no obligation discharges. (This is what rewrap exists for in Vault’s own model, and the reason it cannot serve that purpose here is the chain, above.)

agentplane cannot stop an operator moving that floor, so it makes moving it too far legible. KeyError::Retired is its own answer, distinct from a completed erasure and from a plain refusal, and it names the version the floor has to readmit:

the wrapping key version 'vault:v1' for scope 'acme/case-8f2…' has been retired
by policy — this is not an erasure and not a loss: the sealed bytes are intact
and become readable again if the key service's minimum decryption version is
lowered to admit 'vault:v1'

agentplane drill reports such a case as a finding that names the remedy, rather than as neither opens nor was its key destroyed — the sentence that would send somebody hunting for tampering while a reversible setting is the whole cause.

An envelope says which construction it is

Rotation-immutability has a consequence for the format, not just for the keys: an envelope is read by builds written long after it, for as long as it is retained. A mixed-version fleet mid-deploy, a rollback, and a restore from a backup taken by a newer plane are all ordinary operations that hand one build another build’s bytes.

So an envelope leads with the construction it was written to:

[u8 version][u32 len][wrapped data key][24-byte nonce][ciphertext ‖ tag]

The version is read before any offset is trusted, which is the whole point — a parser that reads a length first has already committed to a layout it may have no rule for. It is one number, exposed as keyring::ENVELOPE_FORMAT_VERSION, and it names the entire construction: layout, nonce width and AEAD together. Changing the cipher changes what the bytes mean, so a second AEAD is a second version rather than a second field — and a reader that picked between suites by trying them would be a decryption oracle, not a parser.

A version this build does not read is KeyError::UnknownFormat, beside Retired and for the same reason:

this sealed envelope is format version 2 and this build reads 1 — the bytes are
intact and not erased; they open under a build that reads version 2

Without the byte, that envelope would have reached the AEAD with its fields read from the wrong offsets and come back as the sealed payload did not authenticate — an incident sentence for a build skew whose remedy is which binary is running.

The rule binds readers too. A component that cannot identify what it is holding says so rather than staying silent: drill’s probe answers nothing to check only for state that was never sealed. State marked sealed whose envelope will not parse, whose version is unknown, or whose erasure scope names a different case is a finding — the last most of all, because erasing that case destroys a key which does not reach those bytes, so the data would survive the deletion request.

Three findings, and they send different people. A version this build does not read is a skew, and the remedy is which binary is running. Damage at a version it does read is an incident. Between them sits the header it reads the version of and cannot parse: nothing has authenticated the bytes at that point, so it is either damage or another build’s shape, and the report says so instead of choosing. A drill that guessed there would either page somebody for a rollback or shrug at a real loss, and it is the same alarm either way.

The version stays 1 until the durable-format freeze, and an envelope at any other version is refused rather than lifted. Pre-alpha shape changes are hard cuts, and here more sharply than anywhere else in the crate: sealed bytes cannot be rewritten into a new shape, so a deployment holding sealed data from an earlier release keeps it readable by staying on that release.

Four rules hold the semantics together, each removing a way an erasure could be quietly incomplete:

  • A scope yields one data key. Two keys in one erasure unit means destroying one leaves the other half readable.
  • An erased scope does not come back. Re-minting would let a late write land in a unit already reported as erased.
  • Erasure is idempotent, and the first tombstone stands — a retry must not rewrite when or why the data went, because that record is the evidence.
  • An erased read reports expired, never missing and never corrupt. Those three send an operator to three different places, and only one is an incident.

A blob stays identified by the digest of its plaintext, so every digest already written to a journal keeps meaning what it meant, and the digest is the envelope’s associated data — ciphertext moved to another address fails to authenticate rather than opening as somebody else’s payload. The storage address underneath is derived from the erasure scope and that digest (blob::unit_address), and the distinction carries an erasure guarantee of its own: identical bytes in two cases are one digest but two objects, so one case’s tombstones cannot destroy another case’s copy of the same document — and neither can its key destruction, because each case’s copy is sealed under its own scope. Deduplication ends at the erasure-unit boundary on purpose; a copy shared across two units would be one that one unit’s erasure either destroys wrongly or provably fails to reach.

The erasure unit is the case, on both sides, because the case is already the retention unit — bytes are linked to their case when they are written, and a second differently-shaped unit for keys would let the two disagree about what an erasure covered. cx.blobs() is where a skill gets a store already sealed to its case; a store held from the builder writes in the clear, and the two would disagree about what erasing the case erased. A blob write on a run under a key ring that belongs to no case is refused rather than quietly unsealed. The run’s journal payloads are a different matter: a record bound to no case seals under tenant/<run> — still an erasure unit somebody can name — and blob::erase_run is the verb that destroys it, the counterpart of erase_case for the unit that call can never reach.

Governed media is payload too, and takes the same route — scoped to its case, or to a named external retention policy when another lifecycle controller owns those bytes. A guard holds the raw store to exactly one reader, because a second write path reaching it directly is the shape of hole worth naming: everything works, the bytes are written, the run succeeds, and the erasure is quietly partial.

erase_case writes every tombstone first and destroys the key last. The order matters in one direction only: a crash between them leaves bytes that are still readable, which running the erasure again fixes. The reverse would leave unreadable bytes with no tombstone, so a later read reports corrupt and someone is paged for an integrity fault that is really a completed erasure.

Each payload gets its own data key, wrapped under the scope’s key and stored inside the envelope alongside the ciphertext:

[u8 version][u32 len][wrapped data key][24-byte nonce][ciphertext ‖ tag]

That shape is not a convenience — it is what a key-management service actually does. Vault’s transit/datakey and KMS’s GenerateDataKey both mint a fresh key per call and wrap it under a named key, so a design expecting a stable per-scope key cannot be implemented against either. The erasure unit is therefore the wrapping key: destroying a scope’s wrapping key makes every data key ever wrapped under it unopenable at once, however many payloads there were.

It is also what makes restore work. A backup holds the ciphertext and its wrapped key — everything needed to bring the bytes back and nothing needed to read them, because the wrapping key never left the service. Restoring into a fresh store, a new region, or a different operator’s hands yields ciphertext and a key nobody can open.

The KeyRing seam is where a deployment points at the thing that already holds its keys. VaultTransit speaks HashiCorp Vault’s transit engine over its HTTP API — four calls, no SDK — so the wrapping key is created inside Vault and never leaves it, and erasure becomes something this crate asks for and cannot undo. A single key-ring conformance battery is run against both the in-process ring and a real Vault, because the two fail in different places: one cannot get a status code wrong, the other cannot get a HashMap wrong. Vault reports a destroyed key as a 400 with a message, not a 404, and only the Vault run can hold the ring to reading that as a completed erasure rather than as a refusal indistinguishable from a permission problem.

One operational detail matters enough to state: a transit key cannot be deleted unless it was configured to allow it, so an erasure against a default key fails loudly here rather than reporting a success that did not happen.

A scope is never a Vault path. Scopes carry /, and an event scope carries a counterparty’s own source and id — ?, #, .. included — so written raw into a URL they would truncate two messages onto one key or walk the plane’s token to another path. Each scope maps to the transit key ap-<hex SHA-256 of the scope>; VaultTransit::key_name(scope) names it for whoever sets deletion_allowed on it. The battery runs against scopes shaped like the ones the plane writes, and checks that erasing one leaves its sibling readable.

MemoryKeyRing lives in testkit and is unreachable without it: it holds the wrapping keys beside the data they protect, so the feature gate is the guarantee rather than a warning in a doc comment.

Tenancy

RuntimeBuilder::tenant names the tenant a plane runs as, and the name is a validated type rather than a string: it refuses / and :, because a tenant called acme/prod would otherwise produce the same key scope as tenant acme with unit prod, and the two would be indistinguishable afterwards.

The tenant is a key component, never a filter. That distinction is the whole of it: a filter is a predicate somebody has to remember to write, and the query that forgets it returns another tenant’s rows. A key component means the same mistake returns nothing. Every test here hands the attacker a valid identifier belonging to the other tenant, because that is the realistic leak — not a guessed id, but a real one arriving through a path that never checked whose it was.

What is bound:

  • data keys are scoped tenant/unit, so one tenant’s erasure cannot reach another’s bytes — even when both use the same case name, which is exactly where a missing prefix collides;
  • the policy request carries the tenant, at admission and at every effect;
  • the redb journal, seal log, cases, events, timers, tasks and batches are tenant-keyed via RedbStore::for_tenant, which prefixes every key. The tenant leads every time-ordered index, so a sweep or a worklist ranges over its own rather than filtering another’s out, and counts are ranged rather than taken from the table — a whole-table count reports every tenant’s backlog as this one’s;
  • PostgreSQL carries tenant as the leading column of every primary key, unique index and foreign key, via PostgresStore::for_tenant. This is the backend that exists for several plane instances sharing one store, so it is the one where a forgotten predicate is both most likely and worst. It is checked against a real Postgres, and the check is adversarial in all three places that matter: a valid run id from another tenant, a correlation key two tenants both use, and an event whose kind and keys match another tenant’s waiting run.

Two correlation paths deserve naming, because both look like collisions and are not. A correlation key is a business value — document/DOC-1 means something different to every tenant, and two of them using it is ordinary. Left global, one tenant’s run would join another’s case and the two would share a history, a deadline set and an erasure unit. Worse, one tenant’s message would resume another tenant’s waiting run, handing it a payload nobody sent it. Both indexes lead with the tenant.

Blob paths lead with the tenant too, and the reason is erasure rather than reading. Blobs are content-addressed, so two tenants writing identical bytes — a standard form, an empty document, a common attachment — land on one object when the path has no tenant in it. Expiring it to discharge one tenant’s request then destroys the other tenant’s data and reports both requests satisfied: the request nobody made is marked done, and the data that should have survived is gone. Encryption does not fix that half; only the path does. The same argument holds at the unit erasure actually names: the erasure unit leads the storage address too (blob::unit_address), so one matter’s erasure cannot destroy another matter’s copy of the same bytes.

The plane and its stores must agree, and try_build checks. The tenant scopes the plane’s data keys; each store handle is scoped separately, so the two are set in different places and can differ. When a key ring is wired that difference is invisible: build() seals case, event, task and outbox state under the plane’s tenant, and an EncryptedMemoryStore seals memory under the tenant it was constructed with, while each store writes its rows under its own — both scopes are real, every run works — and an erasure destroys exactly the key it was asked for without reaching the rows, then reports success. That is the one failure a deletion guarantee may not have, so it is a startup refusal: JournalStore, BlobStore, CaseStore, EventStore, TaskStore, MemoryStore and PushStore each answer a tenant() question, and a disagreement fails the build naming the store and both tenants. A store that does not override the accessor answers default — right for a single-tenant deployment, and refused against a named plane, which is the safe direction.

Serving is tenant-aware: Planes maps an authenticated caller’s tenant to that tenant’s plane. Three details carry the weight.

The tenant comes from the credential, exactly as actor and roles do. It is the field that decides which store answers, so a body-supplied one would be a cross-tenant read with an authentication step in front of it.

The gate returns the plane with the caller, and the surface holds a registry rather than a runtime. A handler therefore cannot reach a store without having resolved whose it is, and every lookup on that registry names a caller rather than a tenant — so the accidental cross-tenant read is unspellable rather than guarded against. The deliberate one is Planes::cross, which records the crossing in the crossed tenant’s own journal before serving anything.

An unregistered tenant is refused, never defaulted. A fallback would turn an unknown tenant into somebody else’s data, and it would look like working software. Each plane also answers under its own policy engine, so one tenant’s rules cannot decide another’s requests.

On A2A the tenant is checked twice: against the card’s routing identifier and against the credential. Those are different questions — what the request asked for, and what the caller holds — and a peer with a valid credential for one tenant naming another is precisely the case where they disagree.

Quotas are per tenant too, and durable for the same reason the keys are: an in-process ceiling vanishes the moment a second instance starts, and it fails open. Concurrent runs and spend per billing period are both bounded — a run’s worst-case spend is reserved when it is admitted — refused at admission, and never consulted on replay — a ceiling crossed since a run happened must not turn its history into a refusal. See operations.