Manifest reference
Every field an agent declaration may carry, what enforces it, and what an absent value means.
On this page
An agent declaration is a YAML file. It is content-addressed, so editing it changes the digest a consumer pins and the journal records; and it is parsed with deny_unknown_fields, so a field this page does not list is a hard failure, not a warning. max_tokns: 100 in a permissive parser silently means no token ceiling at all, which is the single most dangerous thing a configuration format can do.
Two rules apply to everything below, and they are the reason the file exists:
- Every field is enforced or refused. There is no advisory tier. A control the runtime cannot bind is removed from the format rather than left as reviewable intent, because a security document that appears to enforce something it does not is worse than one that says nothing.
- Absence means something, and it is stated. Where silence would be expensive — an unbounded budget — it is refused. Where silence is an ordinary wiring decision — no declared model — it is allowed.
In the tables below, required in the Default column means there is no default: omitting the field is a parse error.
Every manifest on this site is parsed by the crate’s own validator in CI, so nothing here is a snippet that has never been run.
Editor validation
The format ships as a JSON Schema, generated from the same types the parser deserializes into, so it cannot structurally drift from what parse accepts. One modeline gives any editor running the YAML language server autocomplete, hover documentation and inline errors:
# yaml-language-server: $schema=https://hupe1980.github.io/agentplane/agent.schema.jsonagentplane schema prints the same document for vendoring or CI linting. The schema is the format’s shape — unknown fields, missing fields, wrong types and wrong enum spellings fail it exactly as they fail the parser. The semantic refusals documented on this page (an unstated budget, a control nothing performs, an incoherent topology) run only in the parser, so a document the schema accepts still owes agentplane validate a pass before it is a thing to deploy.
The whole document
apiVersion: agentplane.hupe1980.github.io/v1alpha1 # the only value
kind: Agent # the only value
metadata:
name: support-triage
version: "2.0.0"
annotations: # opaque, never read
example.com/business-owner: "Support Ops, F. Meier"
spec:
execution: { kind: tool-calling, max_turns: 6 }
identity: { role: "...", constraints: "..." }
topology: { mode: single, role: specialist }
security: { max_sensitivity_egress: internal, max_delegation_depth: 0 }
capabilities: { provides: [support.triage] }
models:
privileged: { provider: anthropic, model: claude-sonnet-5 }
budgets: { max_tokens: 120000, max_steps: 5 }
tools:
- ref: "tool://tickets/read"
mutates: false
max_sensitivity: internal
description: "Read a support ticket."
output:
schema:
type: object
additionalProperties: false
required: [severity]
properties: { severity: { type: string } }
oversight:
approval: required
deadline: { name: refund-review, kind: working-days, params: { n: 1 } }
memory:
formation:
subject: "$correlation/customer" # per-party, resolved at run time
purpose: "support"
instruction: "Record durable facts about this customer."metadata
| Field | Required | Notes |
|---|---|---|
name | yes | Non-empty. Identifies the agent to an operator; never a grant — a rule keyed on a name grants authority to anyone who types it. |
version | yes | Free-form and compared only for equality. The crate does not parse semver, because it has no version-ordering decision to make and pretending to understand a scheme it never checks would invite one. |
annotations | no | Facts about this agent that the runtime never reads. Namespaced keys, covered by the digest. See below. |
metadata.annotations
A governance catalogue asks a registry entry for name, id, business owner, technical owner, business goal, platform, data sources, tool access, autonomy level and risk classification. The manifest answers most of those better than a registry would; the ownership ones are facts no spec field should enforce, and a document that cannot hold them at all gets a second registry kept beside it, keyed on name + version + digest and drifting from the file. Two sources of truth about one agent is the defect this format exists to remove.
metadata:
name: pattern-compliance-auditor
version: "2.0.0"
annotations:
example.com/business-owner: "Compliance, F. Meier"
example.com/technical-owner: "Platform, T. Nguyen"
example.com/risk-class: "C"
example.com/ticket: "GRC-2291"This does not weaken the rule that a field either has an enforcement point or is refused. Three properties make the whole map intent by construction:
- The runtime never reads it. No key here reaches a gate, a grant, a prompt or a decision, and there is no accessor that turns one into behaviour. So no value here can become a security decision — which is what an advisory control would be, and this format has no advisory controls: a field named as a security, governance, budget or oversight control is enforced by the runtime or refused by the parser.
- The digest covers it. Changing an owner changes the manifest digest, so it is a version bump with a reviewer on it. That is exactly what a governance process wants from an ownership record and what a wiki page cannot give it.
- Keys are namespaced, in Kubernetes’ grammar.
prefix/name: a DNS-subdomain prefix (lowercase labels joined by dots, at most 253 characters) and a name of at most 63 beginning and ending alphanumeric with-,_and.between, with 256 KiB of keys and values in all — so an entry carries into a Kubernetes object unchanged. The prefix is required here where Kubernetes makes it optional: an unqualifiedowneris exactly the name a future first-class field wants.agentplane.hupe1980.github.io/is reserved askubernetes.io/is there, so an annotation can never shadow a field the format grows.
Who may read them is the same line Kubernetes draws. In a cluster, annotations are configuration for controllers all the time — an ingress controller reads its rewrite rules from them, cert-manager its issuer, a cloud provider its load-balancer type — while the API server itself never acts on one. Here the runtime is the API server: it enforces spec and never reads an annotation, because a field that changes behaviour belongs where it is validated, versioned and refused when wrong. Your wiring is the controller: the map is public on Manifest, a registry resolve returns it, and a deploy pipeline, a dashboard or a controller that turns a manifest into a cluster object may read example.com/replicas and act on it — that is what the namespace is for. What the runtime refuses is only to be that controller, and Kubernetes’ own history says why: kubernetes.io/ingress.class was read by controllers as configuration with no schema and no version until it had to be promoted to a real field (ingressClassName) to get them. Anything an agent’s behaviour should depend on goes in spec, and the format grows a field for it.
A blank value is refused: a key that answers nothing reads, to a reviewer, like a question that was answered. Everything else in metadata is still closed — businessOwner at the top level is still a parse failure, and this is one deliberate door rather than a general loosening.
spec.execution
Declaring this makes the agent fully declarative: the runtime supplies the behaviour and you write no Rust. Omitting it means the behaviour is a registered Skill, and the manifest governs its boundary rather than its conduct.
The difference is what the digest covers. A declarative agent is content-addressed in its entirety; a coded one only as far as its declaration reaches.
| Field | Default | Notes |
|---|---|---|
kind | required | completion, tool-calling, planned or call. |
max_turns | 8 | The loop’s turn ceiling, and a planned agent’s step ceiling. A ceiling, not a suggestion: a budget also stops a runaway loop, but only after paying for every turn. |
kind is a closed enum on purpose. A configuration format whose behaviours are open-ended is one nobody can review, because the reviewer would have to know what the string does.
completion — one model call, answered in the output.schema shape, using the privileged model and the identity prompt. No tools, no second turn.
tool-calling — call tools until the model stops asking, then answer. The model is offered exactly the tools spec.tools grants, with the descriptions and argument schemas declared there; the name it returns is matched byte for byte, and one matching nothing comes back as a failed call so the model can correct itself. Arguments carry the completion’s own untrusted label, so protected fields and the egress ceiling decide. output.schema rides every turn, not only the last — which turn answers is the model’s choice, so there is no moment before dispatch at which “this is the final turn” is known, and a schema attached only where the runtime guessed the answer would land is a contract the model can step around by answering a turn early. A turn that asks for tools is untouched; the turn that answers is provider-constrained during generation, and the settled answer is validated against the schema before it is returned. An agent still asking when max_turns runs out fails rather than returning half-formed reasoning as its answer.
planned — plan first, then execute without the model. One privileged call reads the run’s input — which must be trusted, refused otherwise — and answers with a plan in a bounded schema: which granted tools to call, in what order, with which arguments — as JSON text, since constrained decoding has no free-form object. The runtime validates every step against the grants and executes the plan itself. Step outputs travel as references ($step0/customer/email, $input/payee), resolved with labels intact and never read by a model — so a hostile tool output cannot steer later steps, and a protected field is satisfiable by binding to a trusted source. A parse step hands a prior output to the quarantined model under a declared schema and a runtime-injected have_enough_information bit whose false fails the step. The trade: a plan cannot react to what it discovers. Choose planned when the task’s shape is known up front and the data is hostile; tool-calling when the shape is the discovery.
call — dispatch the one granted tool with the run’s input as its arguments, and answer with its result. No model is declared or called: this is the kind an agent framework with a model of its own routes a tool call through, so the plane governs the effect rather than running a second agent. The dispatch is a planned step’s — the arguments held to the tool’s declaration (the catalogue’s, else the grant’s arguments), the grant, its protected fields, the egress ceiling, the budget, and the approval gate when the grant asks for one. spec.input binds, and is the caller’s half of the same shape: input that does not validate against it, or against the tool’s declaration, fails the run before any effect. Refused at parse: not exactly one grant, no spec.input, a spec.input with an object lacking additionalProperties: false (the input is the arguments, so an open object lets a caller send ones nobody reviewed), a spec.models, spec.memory or spec.output block (nothing in a call would read them), oversight.approval: required (the answer exists only once the call has happened — gate the grant instead), and a mutating grant with no protected_fields carrying a trust, source or value rule.
A served caller’s input is untrusted and labelled Internal, from peer:<actor>. So a field marked require_trusted refuses every served call and passes only for the operator’s own --input; a field a caller may fill names that caller in allowed_sources (["peer:app-1"]) or carries a one_of menu; and the grant declares max_sensitivity: internal to receive it at all.
A tool-calling or planned agent granting a tool with no description is refused at parse: a bare name makes the model guess, and the guess is refused at the field check after the tokens are paid for.
spec.identity
| Field | Required | Notes |
|---|---|---|
role | yes if the block is present | Non-empty. What the agent is for, in one line. |
constraints | no | How it must behave. Separate from role because the two are reviewed by different people and change on different schedules. |
There is no workload_id field: nothing would read it, so it would be an identity claim the runtime never checks. Workload identity is configured on the plane and recorded in the journal (IdentityBound). Same reasoning as capabilities.requires.
The prompt lives here so that rewording it is a version bump rather than a deploy nothing records. A prompt composed in Rust has no version at all: it changes, the journal records the run, and nothing connects the two.
Where a long procedure goes: constraints. There is no separate instructions field, and adding one would be surface without semantics — the prompt is exactly role, a blank line, then constraints, so a third field would concatenate the same way while giving a reviewer one more place to look.
role is one line because it answers what is this agent; constraints is unbounded and is where a hundred-line numbered procedure belongs. That it lives in the digest is the point rather than a cost: in a regulated domain, editing step 7 of a procedure should change the identity consumers pin, and should show up as a diff with a reviewer on it. A procedure held in code has no version at all.
The block is optional because an embedder may compose its prompt in code — in which case the digest does not cover it.
A coded skill can still take its prompt from here. StepCtx::manifest() hands a skill the declaration it is running under, and Identity::system_prompt() renders role, a blank line, then constraints — the same string a declarative agent gets:
use agentplane::manifest::Identity;
let system = cx
.manifest()
.and_then(|m| m.spec.identity.as_ref())
.map(Identity::system_prompt)
.unwrap_or_default();So behaviour can stay in Rust while the prompt stays in the reviewed, digested file. What a coded agent gives up is coverage of its conduct, not of its prompt: the manifest still pins what it says, and cannot pin what the code then does with the answer.
There is no templating — no {variables}, no dynamic-instructions callback, no state injection. A templated instruction has no reviewable identity, and the instruction slot is the trusted slot: run-time values spliced into it would wear its trust. Per-run data goes in the input, journaled and labelled, beside the instruction.
A prompt naming a tool the agent was never granted is refused. An ungranted name comes back to the model as a failed call — deliberately, so it can correct itself and never gets the tool it nearly named. The cost is that a procedure naming an ungranted tool fails quietly: the model asks, is refused, improvises, and the step silently does not happen, with nothing in the journal saying the instruction was unfollowable. So a tool named in role, constraints or memory.formation.instruction must be one spec.tools grants — in either spelling: tool://server/name, which is the reviewer’s, and server__name, which is the model’s, and therefore the one an author writing a procedure reaches for.
It only sees those two. “call list_overdue_processes” names a tool in prose, and prose is not something this crate can tell from an ordinary noun — a check that guessed would refuse manifests over the word “search”.
spec.topology
| Field | Default | Values |
|---|---|---|
mode | single | single, collaborative |
role | specialist | specialist, orchestrator |
reason | none — required for collaborative, refused otherwise | parallel-disjoint, distinct-authority |
Three combinations are refused, and each refusal is the point:
specialistwithmax_delegation_depthabove zero — a specialist that may hand off is an orchestrator nobody reviewed as one. A specialist’s effective ceiling is zero even when the numeric field is omitted.singlewith a role other thanspecialist— there is nobody to orchestrate.collaborativewith noreason— collaboration costs roughly an order of magnitude more tokens and opens the whole inter-agent failure surface, so why it is warranted belongs in the file where a reviewer can disagree with it. Areasonon a non-collaborative mode is refused too: a justification for something the agent does not do reads in review as one that was required.
distinct-authority is the reason worth emphasising, because neither side of the public multi-agent debate raises it: the best reason to split agents is often security, not capability. If a sub-task needs credentials the parent should not hold, delegating to a narrower agent is least privilege.
There is no routed/router. Choosing one agent before a run starts is deployment dispatch, and accepting YAML the runtime never executes manufactures confidence.
spec.security
| Field | Default | Notes |
|---|---|---|
max_sensitivity_egress | none — each sink’s own ceiling binds, which for model.complete is public | public, internal, confidential, secret. Combined with each sink’s own ceiling at dispatch; the stricter wins. |
max_sensitivity_journaled | unbounded | The highest sensitivity an argument may reach an effect whose arguments the journal records — may this be written down forever, where egress asks may this leave. Refused at dispatch, before anything is recorded; .keyring(..) is the seal it answer → erasure and keys. |
max_delegation_depth | role-dependent | Checked against the configured identity and against every delegating sink, including in-plane commission. |
content | none | Rules over a value’s content, and uses of a registered checker → below, and what they prove. |
spec.security.content
rules are deterministic; checks hand a value to a registered checker and map the categories it reports. Ids are unique across both, and are what a refusal names.
| Rule field | Notes |
|---|---|
id | Required. |
match | Exactly one of pattern (a regular expression on the linear-time engine; luhn: true keeps only matches whose digits pass the Luhn check), contains (literal substrings; case: fold ignores case) or invisible: true (the code points that render as nothing). |
at | Any of admission: true; sources — model.complete, tool.call, event.await, memory.recall; sinks — model.complete, tool.call, media.fetch. A kind at a position it does not occupy is refused. |
fields | JSON pointers narrowing which subtrees are read. Absent reads the whole value. |
then | refuse, {classify: <sensitivity>} or, at sinks only, {redact: <token>}. |
| Check field | Notes |
|---|---|
id | Required. |
checker | The name a checker was registered under with RuntimeBuilder::content_checker. An unregistered one refuses the build. |
at | sinks, and sources: [model.complete, tool.call]: a check runs as an effect of the step. Never admission. |
on | From the checker’s declared categories to refuse or {classify: <sensitivity>}. A category the checker does not declare refuses the build. |
Every string leaf and object key is read after Unicode NFC. Anything outside these shapes — warn, flag, allow, a score, a second matcher — is refused at parse, naming the rule.
spec.capabilities
| Field | Notes |
|---|---|
provides | The capability names this agent answers to. Runtime::run(capability, input) dispatches on these, and a plane refuses to build if two agents claim one capability. A declarative agent that omits it provides [metadata.name], resolved before the digest, so spelling the default out does not change the digest. |
A coded agent may provide several capabilities — each has its own skill behind it, and the build refuses a declared capability no registered skill serves. A declarative agent provides exactly one, refused at parse otherwise: the capability never reaches the prompt, so a second name would be a distinction nothing executes. Two capabilities are two documents in one room file.
There is no requires twin. Parsed and digest-covered but never enforced, it would be a control the runtime does not check — exactly what a reviewable file exists to eliminate. A field that only documents intent belongs in prose, not beside enforced ceilings.
There is no SKILL.md, no kind: Skill, and no free-form spec.config: instructions live in identity.constraints, on-demand references are spec.context grants, executable helpers are tools, and behaviour shared between agents is an agent of its own, granted as tool://agent/<capability>.
spec.models
| Role | Notes |
|---|---|
privileged | The model trusted with tool calls and decisions. |
quarantined | The model that reads untrusted material and holds no authority. |
quarantined is refused where nothing would select it. Two things point a model at untrusted-derived content on their own: a plan’s parse steps (execution.kind: planned) and memory.formation. A completion or tool-calling agent with neither sends every call to the privileged model, so declaring the second role there would read as dual-model isolation while one model did all the work — a control the file claims and the runtime never applies. A coded agent may declare both: its skill chooses, so the roles are a reviewed allowlist rather than something a tier selects from.
Both are { provider, model } plus three optional per-role ceilings — max_tokens, a per-call output ceiling; max_input_tokens, a per-call input ceiling; and reasoning_effort, an explicit reasoning depth — and an optional pricing, what the role’s tokens cost. The ceilings are enforced on every call the role serves, the quarantined role’s included, so a memory-formation extraction or a parse step runs under the reviewed bounds rather than the driver’s defaults. provider is the name a driver was registered under. The agentplane binary ships openai, anthropic, gemini, bedrock, chat-completions and fake; an embedder registers its own with RuntimeBuilder::provider, and a name nothing was registered under is refused at build rather than at the first call. The pair is refused when both roles name the same provider and model: two roles behind one model keeps the label and removes the control it stands for.
What quarantined does: it is part of the reviewed model allowlist, memory formation runs on it when declared, and a planned agent’s parse steps run on it — no tools, a bounded schema, and nothing handed back to the privileged path but success or failure. The agent’s answer stays on the privileged model. The runtime does not route ordinary completions between the two by content.
Absent means wired in code. models: {} means no inference at all, declared on purpose — a rules-only agent is a legitimate design, and saying so is what distinguishes it from one whose model wiring somebody forgot.
What one call can cost
max_tokens is sent to the provider, so output is bounded before a call. Input is not — it is whatever the conversation has grown to — so max_input_tokens is held against the input the provider reports: a call that sent more fails, billed as reported, and its answer is not handed on. 0 is refused for either ceiling, since no prompt is empty and no answer is either.
Together they state what one call through the role can cost: max_input_tokens plus the output ceiling (max_tokens, or the default every call is sent with) in tokens, and those tokens at the role’s pricing in money — input at the dearest of the input, cache-read and cache-write rates, because which of them a call’s input lands in is the provider’s decision. The run’s per-call bound is the dearest declared role’s, and it is unbounded while any role omits max_input_tokens.
That figure is what a tenant’s spend quota reserves beside the run’s ceiling — a run can end one call past max_tokens or max_minor_units per step in flight — so under a spend quota an agent whose per-call bound is unbounded is refused at admission (per-tenant ceilings). agentplane validate prints the worst case per agent — the ceiling plus the width times one call, and for a tool-calling agent max_turns times one call — derived from the file alone, naming every term the file leaves unbounded instead of printing a total for it.
There is no fallback role. Fallback changes behaviour and must be explicit orchestration, not decorative configuration the runtime never executes.
spec.budgets
Absent is refused. An unstated ceiling is unbounded spend, and that is a decision that has to be visible — declare budgets: {} to mean it.
| Field | Unit |
|---|---|
max_steps | steps |
max_effects | effects |
max_tokens | tokens, across every model call in the run |
max_minor_units | money in minor units — cents, not euros. A float would make a budget that fails to bind by a rounding error, and it is unsigned, so a negative ceiling is a parse failure rather than a ceiling that un-spends itself. Requires pricing on every declared model role → money |
max_replans | replans |
max_wallclock_secs | seconds, named for its unit so a manifest cannot mean minutes → what it costs |
max_denials | policy refusals, before the run is stopped |
max_parallel_steps | how many of a plan’s ready steps run at once |
max_egress_bytes | bytes sent into sinks — a tool call’s arguments, a peer call’s payload, and a model call’s whole request: prompt, tool declarations, every earlier tool result, continuation state and each granted media artifact at its encoded size. Reads cost nothing → volume |
What a model costs
No driver knows what your contract with a provider costs, and this crate ships no price table — rates change, differ per model, and a guessed number is a ceiling that binds in the wrong place without saying so. So a model role states its price, in minor units per million tokens, and every call it serves is priced from its reported usage at the effect boundary, rounded up:
models:
privileged:
provider: anthropic
model: claude-sonnet-5
# Your rates, not these: minor units per million tokens.
pricing: { input: 300, output: 1500, cache_read: 30, cache_write: 375 }All four rates are required; a provider with no cache-write charge states 0. A metered failure — a stream that died after generating — is priced exactly as an answer is. max_minor_units beside a role with no pricing is refused at load: that role’s calls would report no money, and the ceiling would never bind on the agent’s largest cost. The price is covered by the manifest digest, so a rate change is a version bump; it is not part of any effect key, because it changes what a run is billed rather than what the provider is asked.
The one ceiling on how much left
Every other ceiling here bounds work: steps, calls, tokens, money, time. An extraction sized just under any of them passes, because a label answers what may this value touch and never how much of it went. max_egress_bytes bounds the quantity.
Unlike the metered ceilings, it is exact. A token cost is unknown until the call returns, so those refuse once consumption has reached the limit and the run overshoots by one operation; an outbound size is in hand before dispatch, so the effect that would cross the ceiling is the effect refused and nothing over the limit is ever sent. The refusal names what the call would have sent, so you can tell a ceiling that is too low from one call that is too big.
0 is meaningful, like max_replans and max_denials: it says this agent may read and may not send.
It is a ceiling, not a detector. Forty times the median for this capability is a threshold you set against your own traffic, and EffectStarted.outbound_bytes — the same figure, per effect — is what makes that an ordinary query over the journal. A ceiling bounds the worst case; the figure catches the case that stayed under it.
What max_wallclock_secs costs, and what it stops
It is the one ceiling that reads a clock. A run declaring it pays one journaled clock read per step boundary; a run that does not declare it pays nothing. Journaled is what makes it replayable, and it has a consequence worth knowing before you set it: such a history replays under a raised ceiling and not under a build that removed the field, because the reading is part of the step.
It stops the next step. Nothing here interrupts a call already in flight — a ceiling that could would have to abort mid-effect, which is the thing the whole runtime is arranged to avoid.
It counts second boundaries rather than measuring a duration. Elapsed time is the difference between two journaled unix timestamps, so the same 1.4 s of work reads as 1 or 2 depending on where in a second it began. That is what makes the ceiling replayable, and it means you set this as an outer bound on a run, never as a stopwatch on a step.
A run permitted 1s and refused at 1s elapsed has not hit an off-by-one: every metered ceiling here stops the next operation once the figure is at the limit. The refusal says so rather than leaving you to infer it.
Budgets bind the whole run including delegation: commission is an effect, so a sub-run’s reported spend is billed to the run that ordered it.
Every ceiling above except max_replans and max_denials is checked before the work and against every effect — a clock read and a tool call each cost one — so 0 is refused at parse: it does not mean “no spend”, it means the run is refused its first operation of any kind, forever.
max_parallel_steps is the one that bounds width rather than total work. A plan’s ready set runs concurrently, and each step in flight holds a connection, a share of the provider’s rate limit, and one operation’s worth of spend the metered ceilings have admitted and not yet billed — so this is also what bounds how far max_tokens and max_minor_units can be overshot. Omit it and the plan’s own width is the bound, which is the right default for a graph you wrote and the wrong one for a graph anything else may widen.
max_replans and max_denials are the exceptions, and 0 is accepted for both, because each counts an event that may never happen. max_denials is the declarative half of the refusal side channel: a refusal carries a uniform message so a model cannot tell one denial from another, but the refused/allowed bit itself still leaks once per attempt, and what bounds that channel is bounding the attempts. It is counted after the refusal, so max_denials: 0 means the first refusal ends the run. Read operationally it is the same control: a run stuck in a denial loop has stopped making progress, exactly like one that replans without bound.
spec.tools
| Field | Default | Notes |
|---|---|---|
ref | required | tool://server/name, transport-neutral — see what server may be below. |
mutates | true | Whether calling it changes the world. The cautious default. |
max_sensitivity | public | The highest sensitivity this tool may be sent. |
description | none | What the model is told. Required by a tool-calling or planned agent. In the digest, because text that steers tool selection belongs where the system prompt does. |
arguments | derived | JSON Schema. Omit it for a typed Tool: the schema comes from the Rust argument type, and stating it twice is refused because a second copy can only drift. |
requires_approval | false | A person approves this call, seeing the exact tool and arguments, before it is dispatched. Needs spec.oversight and a kind that calls tools (tool-calling, planned or call); refused without either. See approve with an amendment. |
protected_fields | none | See below. |
rate_limit | none | { count, window_seconds }: at most count calls in any window_seconds, across every run of the tenant. See rate ceilings. |
Approving with an amendment
A decision’s amendment is the call. An approving reviewer’s substitute arguments dispatch in the model’s place, schema-checked and judged by every sink gate as the reviewer’s own trusted value, with provenance task:agent.approve_call. spec.oversight supplies the approvers, the obligation bounding the wait, and what happens when it closes.
What server may be
A ref names a tool, not a transport. server may be an MCP connection, tools compiled into the binary, agent for an agent on this plane, or the id of a registered A2A peer — a deployment decision made at build, so one manifest runs against an in-process double in a test and a real server in production.
A peer grant dispatches through StepCtx::call_peer: a delegating hop that extends the run’s chain, counts against max_delegation_depth, and is held to this grant’s fields and ceiling.
Rate ceilings
tools:
- ref: "tool://payments/refund"
rate_limit: { count: 20, window_seconds: 3600 } # twenty an hourA run’s budget sees one run; this sees every run. The count is kept per tenant per tool reference in the quota store, so instances sharing a store share it, and every ceiling any declaration on the plane states for one tool binds every agent calling it. The window slides: twenty an hour admits twenty in any hour, not forty across an hour boundary.
A call past the ceiling is refused before it is announced, recorded on the run as a budget refusal naming the tool, the ceiling and the count reached, and the run stops exhausted. It costs the run’s own budget nothing and is not a policy denial. Resume the run once the window has room; a resume inside a still-full window stands the refusal without recording a second. Replay reads the refusal back and never asks the count.
A retry of one call, and a call re-dispatched after a crash, spend once. Two runs making the same call each spend. An undo is counted and never refused. Nothing is refunded early: a call that was reserved and never announced ages out with the window, because nothing proves it did not reach the world. The instant is each instance’s clock, so a window across instances inherits their skew. A ceiling on another agent’s plane counts only where that plane’s declarations are held.
Refused at parse: a zero count, a zero window_seconds or one longer than 31 days (2678400 seconds, how long a store keeps each call’s row, whatever the window of the declaration that recorded it), and a ceiling on an agent grant. Refused at build: a plane with no quota store.
protected_fields
The field-level rule. Each entry is an RFC 6901 JSON Pointer plus the constraints that path must satisfy:
tools:
- ref: "tool://ledger/transfer"
mutates: true
description: "Move funds between accounts."
protected_fields:
- path: /recipient
require_trusted: true # untrusted data may never select this
- path: /amount
allowed_sources: ["operator:treasury"] # only these provenances
- path: /category
one_of: [refund, adjustment] # only these exact values
- path: /memo
max_sensitivity: internal # this field's own ceilingOrdinary content fields may stay untrusted beside them. That is the whole design: a mutating tool with no protected fields refuses untrusted arguments outright, and declaring which fields a model may influence is how you permit the useful part without permitting the dangerous part.
one_of is the content discipline — the select-from-a-menu pattern made declarative. Every entry was reviewed, so an untrusted influence choosing among them discloses only which approved option was chosen; matching is exact structural equality, never a near miss corrected. It layers over the other rules rather than substituting for them: a field with both a source rule and a menu refuses an allowed source answering something nobody enumerated. What the manifest deliberately cannot express is anything richer — a pattern or format admits values nobody enumerated, and a judgment about one specific value belongs in a coded skill’s release, which carries basis and evidence per instance.
On a mutating grant, at least one protected field must carry a trust, source, or value rule — require_trusted: true, allowed_sources naming where the value must come from, or one_of enumerating what may stand in it. A grant whose every entry carries only a sensitivity ceiling is refused at parse: declaring protected fields is what lifts the whole-object taint gate, and a ceiling bounds how secret an argument may be, not who authored it — so the model’s own untrusted completion would fill every authority-bearing field unconstrained. A source rule names the concrete source an effect’s output carries: tool://server/name, model:{provider}/{model}, agent/{capability}, or an operator identity of your own.
Which grants went unused
`agentplane grants –from A grant is named only where a call’s arguments were readable. On a sealed export the terminal opens nothing, so every grant with no readable use is not established, menus are not derivable, and the sealed and erased counts stand beside the grant table. The library’s Exact MCP context reads, separate from action-granting tools. They remain untrusted data, but which external prompt/resource may enter an agent and at what sensitivity is still a reviewed, digest-covered decision. Prompt arguments are outbound data and Use The mirror of Declared here, that shape is covered by the manifest digest: the schema a model was offered on a given run is the schema somebody approved, and narrowing it is a version bump a consumer can pin against rather than a deploy nothing records. An agent that genuinely takes anything omits the block, which is the honest way to say so. Every object needs The rest of the constrained-decoding subset is not checked here, because providers spell it differently, but the OpenAI driver refuses the call before sending if it is missed: every property in Only meaningful beside The agent registers the obligation, opens a task carrying its actual answer, and returns only on approval. It applies to both execution kinds — a Nothing is written until the answer is approved. In particular An oversight block that performs nothing is refused: See human oversight. The mode an advisory agent needs. A A condition is one JSON Pointer and one of five total operators — Why this may hold a predicate when Two refusals keep it honest: Rows are opened after the approval gate and after formation, so every row in a worklist corresponds to an answer that was actually returned. The task carries the answer itself; it is untrusted model content, deliberately — a worklist whose rows had to be trusted could only carry findings nobody needs to look at. What a declarative agent reads from, and writes to, durable memory. Both halves are optional; the block is refused if it declares neither, and refused beside a coded skill (which calls There is deliberately no semantic search here. Similarity is computed over item content, so anything able to write a memory is a ranking signal — an attacker who cannot taint a value can still decide which clean values a model is shown, and no label shows it. A deterministic recall’s order is a fixed rule no stored item can move, which is what makes it safe to spell as one reviewed line. Ranked retrieval is The memories are folded into the prompt under Each item arrives carrying its own label, so a recall widens nothing: the same egress ceiling governs the model call and the same protected-field rules govern every tool the answer reaches for. Two consequences: Forms bounded durable facts from each declarative answer. Refused without a declared The model proposes bounded key/content pairs, both strings; the runtime derives ids, taint, provenance and retention. Trust is never taken from what the content says. Both retention windows are bounded at each end. Zero is refused, because it expires what it just wrote. So is anything longer than the span between the first and last instant a timestamp can name (631,107,417,599 seconds), which is a window there is no A subject is the unit A binding resolves the subject from something the run already established: Four rules, each a refusal: The keys a binding resolves against are the ones recorded on the run’s A plane declaring either half with no memory store is refused at The extraction runs on the quarantined model when Whose data a run takes in. Each entry is The resolved subject is sealed with the run, so a key ring is what lets a report match it and an erasure of the run’s case takes it with it. A hand-wired run names its subjects with grants::Grants::with_keys opens sealed arguments while the keys stand. A refusal names no grant, so a digest with refused tool calls marks an uncalled grant unused, refusals not attributable.--propose writes <dir>/<digest>.yaml: the manifest with whole unused grants removed, one at a time, each kept if validation refuses the result. It never adds a grant, changes a kept one, or touches a budget, and it is never applied, signed or published — a person reviews it and publishes a new revision. The verb exits 1 when a grant is unused, 5 when the export was incomplete or some calls could not be read.spec.contextcontext:
prompts:
- server: templates
name: summarize
max_input_sensitivity: internal
output_sensitivity: internal
resources:
- server: knowledge
uri: kb://support/rules
output_sensitivity: internal
task_input:
- server: templates
max_input_sensitivity: internalmax_input_sensitivity bounds them. output_sensitivity raises the returned label when the server may disclose classified content; neither field can make server output trusted. Duplicate or blank grants are refused. URIs are exact — no wildcard whose interpretation can disagree with the MCP server’s URI parser.task_input permits answering a server’s outstanding input requests (tasks/update) — an elicitation is a server asking this plane for data, which is the direction an operator most needs to have said yes to. Per server rather than per task, because a task id is minted at runtime and nobody can review a name that does not exist yet; only an input ceiling, since the update returns nothing to label.McpAccess::from_manifest(server, manifest) to avoid restating these grants in code. The gate compares the grant’s ceilings against what the wiring declares, so a coded agent whose McpDataSafety disagrees with the reviewed manifest is refused at dispatch — in both directions, looser and tighter, because two artifacts stating one decision must agree.spec.inputField Notes schemaJSON Schema, digest-covered. The shape a caller must send. Optional, except for execution.kind: call, where it is the call’s arguments, must be closed (additionalProperties: false on every object), and every input is validated against it; schema: {} is refused for the same reason it is on spec.output — it permits anything while looking answered.spec.output, for a caller the result contract never had to consider: a model composing the arguments. Calling run_under from your own Rust, you hold the shape on both sides; an A2A peer’s message is the sender’s problem. Neither stays true once the agent is offered as a tool to somebody else’s model — an MCP catalogue, an Agent Card’s skill declaration — because then the model is handed a shape and composes against it.spec.outputField Notes schemaJSON Schema, digest-covered. Handed to ModelCall::expecting, so it enters the effect key — editing it makes a replay report divergence rather than reinterpreting a stored answer. schema: {} is refused: it permits anything while looking answered.additionalProperties: false — here and in a declarative agent’s spec.tools[].arguments, refused at parse rather than closed for you, because this file is digest-covered and a rewrite would make the document a reviewer signed and the shape that runs two different things. Open is schema: {} one level down, and it is the rule with teeth at dispatch: a driver handed an open object either refuses the call or generates without constraint. A coded agent is exempt — nothing generates against its schemas.required, every array with items, optionality as type: [string, null], no oneOf, allOf or default, and a type on every subschema. agentplane::model::strict_schema_problem answers for one schema, naming the rule it breaks.spec.oversightexecution. Declared next to a coded agent it is refused, because nothing there would apply it, and a file must not claim a human is in the loop when none is.Field Default Notes approvalrequired required gates every answer; tools-only gates only the grants that set requires_approval; none gates nothing and leaves the deciding to triage. See belowapproversanyone Roles that may decide. Empty means anyone — worth choosing on purpose rather than by omission. deadlinerequired The obligation that bounds the wait: { name, kind, params }. The agent registers it, which is why the declaration carries more than a name. kind and params reach the deployment’s Calendar unchanged, so “one working day” means whatever that domain says; a count n that is not a positive integer is refused at parse.on_expirydeny What happens when the window closes. deny refuses the answer. escalate widens the audience and keeps waiting: the escalate_to roles join the reviewers, the stale claim is cleared, and the task leaves the expiry scan — it is answered by a person or answered never. proceed acts unattended.escalate_to— Roles added to the audience when a task escalates. Required by on_expiry: escalate, because widening is escalation’s one enforceable meaning; refused beside any other policy. escalate also needs bounded audiences: an empty approvers already means anyone, which no list can widen.allow_unattendedfalseExplicit consent required for on_expiry: proceed, so acting with no human is a greppable decision somebody made rather than an enum variant they picked off a list.triagenone Tasks opened beside a completed answer. See below. tools-only is the shape most deployments want, because gating a tool-calling agent’s answer is a review that arrives after the tool already ran. Under it, every mutating grant must declare requires_approval — refused at parse otherwise, since a mode that gates tool calls while a call that changes the world needs nobody is a declared control nothing enforces. Read-only grants stay ungated. None of the three modes is a predicate: “require approval when severity is high” changes what the agent does, and that is one step from an if.tool-calling agent has already touched the world by the time it answers, which is the case that most needs a person.memory.formation runs after the decision, because a memory formed from a refused answer would be read by the next run as established fact — a control that governed the reply and not the write would govern the less important half.approval: none with an empty triage and no grant asking for approval is a declaration that reads in review as a human control and is not one.spec.oversight.triagetool-calling agent that grants no mutating tool cannot act at all — its arguments come from a model completion, so a mutating call with no protected_fields is refused by the taint gate on every run, which is why that grant is refused outright. For a whole class of agents the other two modes are therefore both wrong: tools-only gates nothing because there is no mutating call to gate, and required suspends every run until somebody approves a report.triage says: return the answer, and open a task when it says something a person must see.oversight:
approval: none
deadline: { name: unused, kind: hours, params: { n: 4 } }
triage:
- name: breach
summary: "a regulatory deadline was missed"
audience: [grid-operations]
priority: high # low | normal | high | urgent
when:
- path: /deadline_status
equals: BREACH
deadline: { name: triage-breach, kind: working-days, params: { n: 2 } }Field Default Notes namerequired Unique within the block. The task’s kind is agent.triage/<name>, so a worklist can filter on it and a rule cannot collide with a runtime kind.whenrequired Conditions, all of which must hold. Empty is refused: a rule matching every answer is a task per run written as a filter. summaryrequired What the worklist row says, in the words a reviewer reads — in the file, and digest-covered, for the same reason the system prompt is. audienceanyone Roles the row is offered to. deadlinerequired The obligation bounding the row. Its own, not the block’s: how long a run waits for approval and how long a row may sit are different questions. prioritynormal How the row is ranked. equals, in, at_least, at_most, exists. There is no nesting, no or, and no negation. Rules are independent: two matching rules open two rows, because two findings are two desks’ work.approval may not. A triage rule changes nothing about the run. The answer is produced, validated against output.schema, returned, and the memories are formed identically whether a rule matched or not; the only effect is a row in a worklist. That is reporting, and reporting is the one place a declaration can carry a condition without becoming control flow.triage requires spec.output. A predicate over an answer with no declared shape is prose about a document nobody wrote.type: object with additionalProperties: false whose properties lack the field. Deliberately narrower than a validator: anything the walk cannot decide ($ref, anyOf, an open object) passes, because a check that guessed would refuse valid manifests. A rule that can never fire reads in review exactly like one that does.spec.memoryStepCtx::recall and StepCtx::form_memories at the moments it chooses).memory:
recall: # read, before the model is called
subject: "$correlation/malo"
purpose: clearing
limit: 5
formation: # write, after the answer
subject: "$correlation/malo"
purpose: clearing
instruction: "Record stable facts stated in the source."StepCtx::semantic_recall.spec.memory.recallField Default Notes subjectrequired Which pile to read — a literal or a binding, exactly as formation writes it. See below.purposenone Restrict to one retrieval partition. Absent reads every purpose under the subject. limit5Between 1 and 50. Selection is most trusted first, then newest. refresh_accessfalseSlide each selected memory’s sliding-retention window forward, as a second journaled effect. /memory, beside the trusted /system instruction and the caller’s /input, as a list of {id, purpose, content, written_at}. The key is present even when nothing was recalled — a prompt whose shape depends on what the store happened to hold is one no reviewer can read against the manifest.security.max_sensitivity_egress fails the run at the model call rather than being filtered out — a silent drop would make the answer depend on a ceiling nothing in the transcript mentions.execution.kind: planned may not declare a recall. That kind refuses untrusted input because its plan is compiled from what the planner reads, and a recalled memory is untrusted whenever whatever wrote it was.spec.memory.formationprivileged model.Field Default Notes subjectrequired Where the memories are filed — a literal, or a binding resolved per run. See below. purposerequired Mandatory retrieval partition. instructionrequired What the extraction model is asked to record. max_items3Between 1 and 10. retention_secondsnone Fixed expiry. Seconds, and the window is bounded at both ends — see below. access_retention_secondsnone Sliding expiry, refreshed by an explicit journaled touch. Same bounds. max_sensitivitypublicCeiling on what the forming model may be shown. now to add it to. A window inside the bound can still land past the last instant when the run’s own clock is near it, and the step that forms the memory refuses that too.Binding the subject to the party a run is about
forget_subject erases, so a literal one pools every customer, meter and matter the agent ever reasoned about under one key. One party’s facts are then recalled into another party’s run, and an erasure request naming one person cannot be satisfied without destroying everybody’s. Under a data-protection regime that is a defect, not a caveat.Binding Resolves to $correlation/<namespace>The value of one of the run’s correlation keys. $caseThe case id. $input/<pointer>An RFC 6901 pointer into the run’s input — only if that field is trusted. anything else A literal, exactly as written. $$x is the literal $x.memory:
formation:
subject: "$correlation/malo" # one pile per metering point
purpose: clearing
instruction: "Record stable facts stated in the source."$ value is refused, never taken as a constant. Reading $correlaton/malo as the string "$correlaton/malo" would file every customer’s memories under a typo, and nothing would look wrong until an erasure request.$input requires a trusted field. A subject taken from untrusted input is whoever supplied it choosing whose memories this run writes into — strictly worse than the pooling bindings exist to fix. Correlation keys need no such check: correlation is a deterministic lookup performed at admission from keys the deployment’s edge supplied, and no model touches it.build, because nothing could ever resolve it.CaseBound journal record — not the case’s keys as they stand now. A case accumulates keys over months, so re-reading them would let a resumed run resolve a subject the live run never saw. A hand-written skill reaches the same values with cx.correlation() and cx.correlation_value(namespace), so reading these memories back does not mean guessing at a naming convention.build. For formation the cost of finding out late is the point: it runs after the answer, so left to run time it fails once the run has already paid for its model calls.spec.models declares one, and on the privileged model otherwise — see spec.models above for why. Write the instruction extraction-only, with fabrication refused: record stable facts stated in the source; do not infer addresses, dates or identifiers that are not literally present.spec.data_subjects$input/<pointer> or $case from the memory-subject grammar, resolved once per run into its DataSubjectBound journal record; every value derived from the run’s input, its inbound events and its case-state reads is then attributed to it, and agentplane subject follows that data to the effects it reached (erasure).data_subjects: ["$input/customer/id"]$correlation/<namespace> is refused at parse. A correlation key is readable by design in CaseBound, RunSuspended and the case store, so a subject read from one would be sealed in one record and left in the clear in the others, and an erasure would leave it behind. Read the identifier from the input with $input/<pointer>, and correlate on a key that is not personal data.$case resolves to the case id, which identifies nobody.$input may name an untrusted field, unlike a memory subject. A data subject attributes and decides nothing — no gate and no policy context reads it — so a value naming the wrong party produces a false listing, not a write into another party’s pile. The report marks such a binding asserted by untrusted input.RunTerms::subject.What is deliberately not in the format
routed/router topology, and a fallback model role — both would be YAML the runtime never executes.