Agentic and LLM systems

Model attack trees for LLM applications and AI agents: the five attack-surface zones, prompt injection and tool-misuse paths, probabilistic guardrails, and bypass-rate measurement.

LLM applications and agents are where path-based reasoning pays off most. A prompt injection is rarely the goal. It is the first step of a path that continues through planning, tool use, and egress, and the defences along that path are often probabilistic. Enable the agentic profile next to default:

# .specify/extensions/attacktree/attacktree-config.yml
profiles: [default, agentic]

# Five zones

The profile tags nodes with the zone of the agent they target, following Christian Schneider’s scenario-driven approach:

ZoneWhere the step happensTypical attacks
inputevery channel into the context: prompts, retrieval, email, tool output, tool descriptionsdirect and indirect prompt injection, tool description poisoning
reasoningplanning and goal interpretationgoal hijack, sensitive context leakage
toolstool execution with the agent’s privilegestool misuse, delegated credential abuse, egress
memorycontext, working memory, long-term storagememory poisoning that persists across sessions
inter-agentmessages between agentsspoofed or tampered messages, cascading compromise

Walking a goal zone by zone is a reliable way to find the paths. A typical exfiltration path needs a step in input (instructions reach the agent), reasoning (the plan shifts), tools (a data-reading tool runs), and tools again (data leaves). That is an AND of four steps, and every one of them is a place for a control.

# Vectors and references

The profile proposes 15 attack vectors with references to the OWASP Top 10 for LLM Applications 2026 (owasp-llm-2026), the OWASP Top 10 for Agentic Applications 2026 (owasp-asi-2026), MITRE ATLAS (mitre-atlas), and MAESTRO layers. Proposals are a checklist, not a template: the model records a vector only when the spec makes it possible.

# Probabilistic controls

Guardrails, injection classifiers, goal-lock checks, output filters, and human approval reduce attacks without blocking them. Mark them probabilistic: true:

- id: control.injection-classifier
  name: Injection classifier on retrieved content
  kind: preventive
  nodes: [node.inject-via-document]
  effect: medium            # assumed until measured
  probabilistic: true
  validation:
    - Send 300 injection probes through the retrieval path; count how many reach the planner

The effect is an assumption until a micro attack simulation measures it. Converge then records the measured bypass_rate, and the tree uses it. In Monte Carlo runs, probabilistic controls are sampled as well, so goals that depend on them show a wider spread.

Two rules from practice:

  • Layer different failure modes. Two classifiers that both fail on the same payload family are one control, not two. Pair a probabilistic control with a deterministic one, such as a tool allowlist, an egress allowlist, or scoped credentials, on the same path.
  • Give humans the context. Human approval that shows “Send email? Yes/No” without the reasoning chain approves the attack. Write a validation step that checks what the approver sees.

# The same example, zone by zone

The shipped example is an agentic assistant with a retrieval path and a ticket tool. Its goal goal.abuse-ticket-credential combines an injection (input) with a call beyond scope (tools). Its most likely path, however, is an operator reading the credential from configuration. Path-based ranking surfaces exactly this kind of finding: the AI-specific path is real, but not the cheapest.