agentplane

Regulation

EU AI Act obligation by obligation, mapped to mechanisms that exist — and an explicit list of the ones that do not.

On this page
  1. 📅 Where the EU AI Act actually stands
  2. ✅ What the runtime gives you
    1. Art. 12 — automatic recording, enabling traceability
    2. Art. 14 — human oversight, and the ability to intervene and stop
    3. Art. 13 — a machine-readable description of the agent
    4. Art. 15 — accuracy, robustness, cybersecurity
    5. Recovery rehearsal, and retention, on a schedule
    6. Art. 26 — deployers keep logs
  3. ❌ What it does not give you
  4. 🧭 Other frameworks
  5. 🙋 If you are evaluating this for a regulated deployment
    1. Answers evaluators have had to ask for

agentplane is not compliant with anything, and cannot be. Compliance attaches to a system in a context, assessed by its provider or deployer. A library has no context.

What it provides is technical means: the record-keeping and human-oversight machinery that several obligations require, built because durable execution needs them anyway. That last part is the reason to trust them — these mechanisms are load-bearing for crash recovery and replay, so they are exercised on every run rather than only when an auditor asks.

This page maps obligations to mechanisms that exist, and is equally explicit about the ones that do not. A compliance page that lists only what a tool offers is a sales document.

This is the only place the mapping lives, and it is downstream of the mechanisms. Each row names something defined elsewhere — the journal, the worklist, the export, the witness — and those definitions are authoritative about what they do; a row is authoritative about which obligation they are being read against. So when a mechanism changes, the row moves. Nothing in the other direction: a statute does not become satisfied because a page says so, which is the sentence at the top of this one.


📅 Where the EU AI Act actually stands

The Digital Omnibus on AI — Regulation (EU) 2026/1744, published in the Official Journal on 24 July 2026 and in force from 27 July — moved the high-risk dates and left the transparency ones alone:

Applies from
Art. 50 transparency — disclosing that users are dealing with an AI, marking synthetic content2 August 2026 (in force now; watermarking of already-marketed systems graced to 2 December 2026)
Annex III high-risk obligations (Art. 9–15, 26) — standalone systems2 December 2027 (deferred 16 months from 2 August 2026)
Annex I high-risk — AI embedded in regulated products2 August 2028

The articles were deferred, not amended. The mapping below is unchanged by the Omnibus; only the calendar moved.

Art. 50 is not a runtime obligation. Telling a user they are talking to an AI, and marking generated content, happen in your interface — not in a journal. Nothing here discharges it, and the date being live now is exactly when a tool claiming “AI Act ready” is worth least.


✅ What the runtime gives you

Art. 12 — automatic recording, enabling traceability

The journal is the mechanism, and it is not an audit log bolted alongside the system: it is what the system executes from. A record that was not written is a step that did not happen.

RequirementMechanism
Automatic, over the lifecycleEvery effect is journaled before it is attempted; nothing reaches the world unrecorded
TraceablePer-record hash chain — any edit to any record invalidates every record after it
AttributablePer-record signatures naming the workload identity that wrote them (signing)
Detecting removalA per-plane Merkle log over sealed runs, with inclusion and consistency proofs — deleting a whole run is detectable, which a per-run chain alone cannot do
Checkable by someone elseAn offline auditor (audit) that runs against a store it did not write, and reports what it could not check as prominently as what it did
Which instructions produced a decisionThe system prompt lives inside the digested manifest (manifest), so a rewording is a version bump. A prompt composed in the deployer’s code has no version at all: it changes in a deploy, the journal faithfully records every run it affected, and nothing connects the two
How much left, not only whatEffectStarted.outbound_bytes — the canonical size of the value that crossed each sink, beside the label of what crossed it. Volume is the axis a sensitivity label does not have → volume is not sensitivity

A run refused before it exists is deliberately not in the journal, and an evaluator should know where to find it instead: the agentplane.policy.denials metric, carrying the action. So how often policy stopped a run from starting is a telemetry question by design rather than by omission — the reasoning, and what it would cost to journal it, are under what no audit can answer.

The limit, stated plainly: a checkpoint that never leaves the operator’s store is exactly as trustworthy as the operator. The Witness seam enforces the decision that matters — a checkpoint is cosigned only if it provably extends the last one seen, so a shrunken log and a split view (a second history of the same size) are both refused — and HttpWitness speaks C2SP tlog-witness, so the counterparty can be an existing public witness rather than a second process you also own. It verifies what comes back: a cosignature counts only if it is a valid Ed25519 signature with a non-zero observation timestamp, over the note that was submitted, by a key the deployment registered under both its name and its note key id.

Both directions are wired. RuntimeBuilder::witnesses(witnesses, quorum) submits on the periodic sweep, and agentplane audit --witness <prefix> --witness-key <name>=<base64> reads back what a witness holds — which is the direction that matters for an audit, because it puts the anchor in a reader’s hands rather than the operator’s. Naming two witnesses is a check the plane cannot run about itself: a split view is exactly two witnesses holding one tree size with two different roots.

What is missing is therefore not code but a counterparty. Until a second party runs a witness for your log, a witness you host yourself proves nothing about you. See status.

Art. 14 — human oversight, and the ability to intervene and stop

RequirementMechanism
A person can interveneDurable worklists: a run suspends and waits, costing a row rather than a thread
The right personCandidate roles, and excluding for four eyes — the proposer cannot approve their own action
Not answering is a decisionOnExpiry is declared up front — deny, escalate, or proceed — so an unanswered approval applies a stated policy instead of hanging. Escalation must name the roles it widens to, and the runtime widens them, so “a higher instance was involved” is a fact the worklist enforces rather than a label on a row
The ability to stopCooperative cancellation: the run unwinds what it did as a saga, and the record names who asked
Refusing to guessA run that cannot account for an outcome is Quarantined rather than unwound — reversing everything except the one thing nobody can account for is how a system refunds money nobody took
Declared, not rememberedspec.oversight puts approval in the reviewable file (manifest), so a declarative agent’s answer waits for a person by declaration rather than because a developer coded the call. Declaring it where nothing would apply it is refused, so the file cannot claim a human is in the loop when none is
A missed window reaches somebody who can answer itBreaches are listed in their own right, outlive the case’s closure, and leave the listing only when a person accounts for one — POST /obligations/acknowledge, recording who looked. The obligation stays Breached: what ends is the question, not the fact

Art. 13 — a machine-readable description of the agent

RequirementMechanism
What the agent is, in one reviewable artifactA manifest declares the prompt, grants, ceilings, models, result shape and oversight, and the runtime refuses effects the declaration never named. The A2A Agent Card is derived from the same file, so what a peer is told and what the runtime enforces cannot drift
The version that runs is the version that was reviewedThe registry addresses a manifest by digest, refuses to replace a published version with different content, and lets a caller pin the digest they reviewed — the form that survives the registry itself being compromised
Who published itA domain-separated publisher signature, verified on resolve, with publisher reassignment refused. An unsigned publish may adopt its first signature; an existing publisher may not be replaced
Facts a registry entry needs and a manifest could not holdmetadata.annotations — business owner, risk class, ticket — namespaced, never interpreted by the runtime, and covered by the digest
Which agents does this organisation runRegistry::names(), against a durable registry: both shipped stores implement the registry, so the inventory survives the process that published it

Because annotations are inside the digest, changing an owner is a version bump with a reviewer on it rather than an edit in a second system — and agentplane validate --require-annotation KEY makes a missing one a build failure, without the runtime reading any of it.

The limit, stated plainly: trust in publisher keys is a deployment decision. This crate never mints a key and never decides to trust one — resolve_verified returns the KeyId that signed, and what that identity is allowed to publish is somebody else’s policy.

Art. 15 — accuracy, robustness, cybersecurity

Not a feature but an evidence question, and the evidence is the assurance ladder: TLA+ models of the core protocols, an exhaustive crash-schedule sweep over every prefix of a real journal, store conformance batteries run against every backend, and a mutation sweep that breaks each guarantee on purpose to prove its test notices. Operations describes what each layer does and does not cover.

Recovery rehearsal, and retention, on a schedule

A control that exists and is never exercised is one an audit cannot count, and a retention policy nothing enforces is a document.

RequirementMechanism
Rehearse recoveryRuntime::drill holds every case’s blob digests and sealed-state keys against the live stores, sorting them into intact, erased by design, lost, and erased with no readable record of it. agentplane serve --drill-every 86400 runs it on a timer; agentplane drill runs one pass
Prove a copy without this crateagentplane verify history.jsonl --checkpoint cp.note recomputes an export from its own bytes; restore rebuilds a store and proves it by its own checkpoint
Enforce a retention windowRuntime::retain(older_than, at, reason) erases every closed case opened before the window: blob tombstones, and the case’s key scope destroyed, which reaches every replica and backup at once. agentplane retention plan --older-than-days N lists what that pass would erase, through the same selection rule, from a binary that wires no store able to erase
Know what retention did not reachEvery pass returns not_erasable, and it is the half that matters: without a key ring, journal payloads stay verbatim. A count with no coverage statement beside it is how a deployment comes to believe an obligation is discharged

Art. 26 — deployers keep logs

The journal is the log, it verifies offline against a store it did not write, and it comes out in a form nothing here has to be present to read:

agentplane export --store ./journal.redb > history.jsonl
agentplane audit  --store ./journal.redb > report.json
agentplane verify history.jsonl --checkpoint cp.note   # check a copy, offline
agentplane restore history.jsonl --store ./rebuilt.redb

--checkpoint is the deletion check, and without it there is none: a root rebuilt from the file proves only that the file agrees with itself, which is what an editor who dropped a run also achieves. Supply one an earlier audit printed, or fetch one with --witness and --origin — how the checkpoint is obtained is what the verdict rests on.

Both verbs take a store and nothing else — no manifest, no source tree, no Rust toolchain — because that is what an auditor holds. The export is JSON Lines: a header naming the log, its checkpoint and the canonicalization rule the digests were computed under; one line per record carrying prev_hash and hash, so the chain can be re-walked from the file alone; one line per case, carrying the matter’s status, version, correlation, obligations and blob digests — the case layer is beside the journal, not derivable from it, and a regulator’s question is usually about a matter; and a trailer. The trailer’s absence is how a file cut short by a full disk or a killed pipe is told from a complete one, any run that could not be read is named in it rather than quietly missing, and its case count is what catches the case layer stripped whole. Case state travels as stored — sealed stays sealed, because an export of plaintext would quietly undo erasure.

verify takes the file and nothing else — no store, no manifest, no toolchain — and re-seals every record through the same function the store sealed with, so agreement is a statement about the bytes rather than about the file agreeing with itself. It then rebuilds the Merkle log from the positions the export carries and compares the root against the checkpoint in its own header. That last step is what catches a whole run deleted from the middle: every surviving chain is internally consistent, because a chain links records within a run and knows nothing about its neighbours.

restore is the other direction, and it proves itself the same way: the rebuilt store must commit to the same root at the same size as the export claimed. The log’s name is deliberately not part of that verdict — a recovery is normally into another tenant of whichever database survived, and a relabelled log with identical contents is still the history. The report names the relabelling beside everything else the file could not carry. Signatures do not survive unless the restoring store holds the original key, which the report also says rather than leaves to be found.

It exports what the chain committed to, which with a key ring configured is ciphertext. That is deliberate: an export of plaintext would put a copy beyond the reach of key destruction, and undo the erasure below.

The retention floor is a control, not a hope. Art. 26 makes a deployer keep logs for a minimum period, and an erasure request can arrive inside it — the one place a retention ceiling and a retention floor point in opposite directions. The automatic pass cannot tell them apart: a matter under a preservation order looks like every other closed case old enough to sweep. A legal hold is what refuses the erasure.

agentplane hold --store ./journal.redb --case case_01JD... \
  --reason "Art. 26 retention floor; supervisory request 2026-114" \
  --actor compliance-dana
agentplane hold list --store ./journal.redb   # everything standing, with reasons and who placed them

The standing holds are listable by somebody who does not already know which case to ask about, which is what makes the floor auditable: the question at review time is what are we still keeping, and on whose instruction. The mechanics are in erasure.

What it does not do is decide the floor for you. There is no default period here for the reason --older-than-days has none: the window is a legal determination, and a crate that picked one would be choosing somebody else’s.

Erasure is possible without breaking the record — for data you kept out of the chain. BlobStore::expire drops a blob’s bytes and leaves a tombstone; the hash chain still verifies afterwards because it only ever committed to the digest. So you can prove what happened and that the record is unaltered without retaining those bytes, which is what makes Art. 26 and Art. 17 compatible rather than opposed.

A record is never deleted — the chain is append-only. But erasability is a wiring decision, not a size one: the 1 MiB record ceiling is a refusal, not a router, and sensitivity does not track volume. A name, an address and an IBAN are a few hundred bytes. Three mechanisms decide it, and they compose:

What it does
.keyring(..)Seals journal payloads under a per-case key. The chain commits to ciphertext, so destroying the key erases every copy — live store, replica, backup — and the history still verifies without it → erasure
cx.store_blobPuts bytes in a blob at any size and journals only the digest
security.max_sensitivity_journaledRefuses at dispatch, before the announcement, when a value above the ceiling would be journaled

So customer data in a prompt is permanent on an unsealed journal and erasable on a sealed one. Erasure and keys partitions it row by row.

Erasure is answered by case, which is the unit a request actually names. erase_case tombstones every blob that case produced and leaves other cases alone. Bytes are linked to their case when written, through cx.store_blob, because a digest cannot be reversed afterwards to discover what matter it belonged to.

What is still missing: a scheduled TTL — object-store lifecycle rules do age-based expiry better than a sweeper could, at the cost of deleting rather than tombstoning. And on an unsealed journal, personal data that reached a record cannot be removed: wire a key ring, or keep it out through cx.store_blob and max_sensitivity_journaled. Governed memory is the one store .keyring(..) deliberately does not wrap — its erasure unit outlives the case and its adapter is single-node by contract — so wrap it explicitly.


❌ What it does not give you

ObligationWhy not
Art. 9 risk managementThere is a policy seam and a Cedar adapter, but no risk-tier model. agentplane policy check measures a candidate bundle against the requests an export records; nothing proves properties of a policy set
Art. 50 transparency to usersAn interface obligation, not a runtime one (above)
Anything about your modelBias, accuracy, training data, and evaluation are properties of the model and its use. This is a runtime
A conformity assessmentA person does that, about a system, in a context

🧭 Other frameworks

ISO/IEC 42001 and the NIST AI RMF consume the same artifacts — the journal answers “what happened and can you prove it” regardless of which framework asks, so no framework-specific integration is needed or planned. The export is the integration: JSON Lines goes into whatever collects evidence.

The AI-logging standards. ISO/IEC 24970 (AI system logging; FDIS ballot, stage 50.20, 28 August 2026) and CEN/CENELEC prEN 18229-1 (AI trustworthiness framework — Part 1: Logging; draft issued for public enquiry 5 June 2026) both specify the logging of an AI system’s events. Against either, the journal is the log — every effect recorded before it is attempted, and hash-chained — and the export, under its published format, is the copy a second party reads without this crate. Neither is final and neither has been mapped clause by clause, so no conformance to either is claimed.


🙋 If you are evaluating this for a regulated deployment

Three questions worth asking of any tool in this space, including this one:

  1. Is the audit record the thing the system executes from, or a copy? A copy can disagree with reality; this one cannot, because a step that was not recorded did not run.
  2. Can someone outside the operator check it? Here: partly — the chain and signatures verify offline, but without external anchoring you are trusting the operator not to have removed a run wholesale.
  3. What does it refuse to tell you? The auditor reports skipped checks beside findings, and status lists what is not built. Anything that reports only findings is telling you about its coverage by omission.

Answers evaluators have had to ask for

Each lives on a page you may not have opened, so they are collected here. If you arrived with a security control catalogue rather than a statutory one, this table is still the right one.

QuestionAnswer
Which format-freeze conditions are not met?None; the freeze is an act not yet performed, and it lands with the checks that hold a frozen vector and the operator vocabulary’s spellings, which are not yet in place → the conditions
Who threw an emergency stop, or placed a legal hold — and who lifted it?On the row while it stands, with what established the name — authenticated when a credential named it, asserted when somebody with the store typed it → the stop. A lift or release is journaled as a sealed run of its own (halt-lifted, hold-released) naming the lifter the same way, listed by GET /halts?state=lifted and GET /holds?state=released
How much data did a run send?EffectStarted.outbound_bytes, per effect → volume. Budget::max_egress_bytes bounds it, exactly — the size is known before dispatch, so the call that would cross the ceiling is the call refused
How often did policy stop a run from starting?The agentplane.policy.denials metric, by action. Not the journal, and deliberately
What does max_denials actually bound?The refusals a model can probe — sink refusals, which come back as REFUSED and let the loop continue. An engine denial ends the run in that same loop → what it counts