Starving the Fire
Table of Contents
Every firefighter learns the fire triangle on the first day of training. Fire is not a mysterious force that requires courage to defeat. It is a system with three dependencies, fuel, heat, and oxygen. Remove any one and the fire cannot sustain. The entire discipline of fire prevention is built on this insight. A grease fire in a kitchen? Do not throw water on it. That spreads the burning oil. Cover the pan with a lid. You have removed the oxygen. An electrical fire? Cut the power at the breaker. You have removed the heat source. A wildfire advancing on a town? Bulldoze a strip of land ahead of the fire line. You have created a firebreak, removing the fuel from the fire’s path. The fire reaches the break and dies. Firefighters do not ask the fire to burn less. They identify which element is most practical to remove given the context, and they remove it.
In June 2025, Simon Willison named the lethal trifecta for AI agents. The three elements are access to sensitive data, exposure to untrusted content, and the ability to communicate externally. When all three are present, an attacker can trick the agent into stealing your data. In October 2025, Meta formalized this as the Agents Rule of Two, which states that agents must satisfy no more than two of the three properties. Both frameworks identify the threat. Neither provides a systematic answer to the question that matters. Which leg do you break, and how?
The AI security industry is trying to make the fire burn less. It calls the effort “guardrails.” This post argues that the defense is not a better filter. It is a firebreak. Each leg of the trifecta maps to specific architectural controls from this blog’s framework, including scoped identity, deterministic policy, context isolation, and egress restriction. The choice of which leg to break depends on the deployment pattern, not a universal prescription. The fire triangle teaches the same lesson. The technique depends on the type of fire.
The three elements #
Willison’s trifecta identifies three capabilities that, individually, are why agents are useful. Combined, they create a direct path to data exfiltration.
Sensitive data is the fuel. An agent that cannot access private information is an agent that cannot help with most real work. Customer records, source code, financial data, credentials, internal documents. The value proposition of an AI agent is that it operates on your data so you do not have to. The more useful the agent, the more fuel it carries.
Untrusted input is the heat source. Any mechanism by which attacker-controlled content enters the agent’s context window is an ignition point. Direct input from users is the obvious vector. Indirect input is the dangerous one. A document the agent summarizes, an email it reads, a web page it browses, a Model Context Protocol (MCP) tool description it loads. As I showed in The Architecture of Inevitability, the self-attention mechanism processes all tokens identically regardless of origin. A token from a poisoned email competes for attention weight against the system prompt through the same softmax computation. The attacker does not need access to the prompt. They need to place content where the agent will read it.
External communication is the oxygen. The agent can send data out. An HTTP request, an email, a tool call, a rendered markdown image, a pull request, even a clickable link. Without this pathway, a compromised agent is a fire in a vacuum. The injection succeeds inside the model, but the stolen data has nowhere to go. Meta’s Rule of Two extends this leg to include state changes. Writing to a database, modifying a configuration file, executing code. This is a meaningful extension. An agent tricked into deleting production data does not need an exfiltration channel. The damage is the state change itself.
Willison’s framing covers data exfiltration. Meta’s covers state modification. Both describe the threat. The gap is the prescription. “Avoid the trifecta” is correct but incomplete, like telling a firefighter “avoid fire.” The question is which element to remove, using which technique, for which type of deployment. The remainder of this post maps each leg to specific architectural controls from this blog’s framework and provides a decision framework for common deployment patterns.
Every major system has burned #
The trifecta is not a theoretical risk. Willison maintains a catalog of production exploits spanning every major AI platform. The pattern repeats with mechanical regularity.
In June 2025, researchers disclosed EchoLeak (CVE-2025-32711, CVSS 9.3), a zero-click prompt injection in Microsoft 365 Copilot. The attack required no user interaction. An attacker sent a single crafted email to the target. Copilot, processing the inbox, ingested the email’s hidden instructions. Those instructions directed Copilot to retrieve sensitive data from the user’s context, encode it into a URL using Markdown’s reference link syntax, and auto-fetch the link, exfiltrating the data without the user clicking anything. All three legs were present. Copilot had access to the user’s emails and documents (fuel), it processed an attacker-controlled email (heat), and it could render URLs that triggered outbound requests (oxygen). The prompt injection classifiers Microsoft deployed to detect prompt injection were bypassed by phrasing the malicious instructions as if they were addressed to the email recipient, not to an AI assistant.
In February 2026, the SANDWORM_MODE attack demonstrated the trifecta across layers. Nineteen malicious npm packages impersonating popular developer utilities injected a rogue MCP server into the configurations of Claude Code, Cursor, and VS Code. The MCP server’s tool descriptions contained embedded prompt injection that instructed the AI assistant to read ~/.ssh/id_rsa, AWS credentials, and npm tokens, then exfiltrate them without the developer seeing a prompt or confirmation dialog. As I described in
Trusting the Label, the tool description entered the context window as tokens and competed for attention weight against the system prompt. The description was the execution.
In January 2026, Microsoft 365 Copilot read and summarized emails it was explicitly configured not to touch. For twenty-eight days, sensitivity labels said no, data loss prevention (DLP) policies said no, and Copilot read them anyway. The fuel was present despite the policy. The fire burned through the label.
I contributed to this pattern myself. In Weighting the Switch, I described granting my AI agent blanket write access to my home directory because the per-action approvals were slowing me down. What I did not describe was the full picture. The agent also had MCP servers connected, tools that could browse the web and fetch URLs. I had all three legs of the trifecta active in a single session. Access to my filesystem (fuel), exposure to web content through MCP tools (heat), and the ability to write files and make network requests (oxygen). The architecture was a lit match in a room full of fuel with the windows open. The agent never exploited it. But the path was there, and I did not see it until I mapped the trifecta against my own setup.
The National Institute of Standards and Technology (NIST)-partnered red teaming competition in 2025 quantified the scale. Nearly 2,000 participants attacked 22 frontier AI agents. Policy violations occurred 20 to 60 percent of the time on the first query. After ten queries, nearly every attack succeeded. These are not edge cases. They are the expected behavior of systems where all three elements of the trifecta are present.
Smoke detectors are not sprinklers #
The industry’s dominant response to the trifecta is the guardrail, a probabilistic classifier, frequently a large language model (LLM) itself, that processes the agent’s input or output and returns a confidence score. As I argued in The Architecture of Inevitability, this defense is recursive. An LLM-based guardrail processes tokens through the same self-attention mechanism as the model it guards. It suffers from the same softmax-forced commitment. The same adversarial techniques fool it.
Mindgard and Lancaster University tested six production guardrail systems from Microsoft, Nvidia, Meta, Protect AI, and Vijil. They bypassed every one using character obfuscation, adversarial perturbations, and emoji smuggling. The attack success rate reached 100% against Protect AI v2 and Azure Prompt Shield. Gray Swan’s benchmark of Claude Opus 4.5, the highest-scoring model in their evaluation, found a 4.7% success rate per attempt. At one hundred attempts, 63%. An agent that processes two hundred requests per day faces those odds daily.
The strongest counterargument is that guardrails catch 95% of attacks, and 95% is better than nothing. Willison’s response is correct: in application security, 95% is a failing grade. Parameterized queries do not stop 95% of SQL injection. They stop 100%. The NX bit does not prevent 95% of buffer overflows. It prevents all of them. The standard for a security control is not “usually works.” It is “works when it matters.”
A smoke detector is useful. It alerts you to danger. It belongs in every building. But no fire code in the world treats a smoke detector as a substitute for a sprinkler system, a fire-rated wall, or an emergency exit. The smoke detector is a probabilistic signal. The sprinkler is a deterministic response. Guardrails are smoke detectors. They belong in the stack. They are not the defense. The defense is the firebreak, an architectural control that removes one element of the triangle regardless of the fire’s sophistication.
Breaking the triangle #
The fire triangle teaches that removing any one element is sufficient. The same holds for the trifecta. The architectural question is which element is most practical to remove for a given deployment, and which controls accomplish the removal.
Starving the fuel #
The most direct way to limit what an attacker can steal is to limit what the agent can see. If the agent holds a token scoped to a single customer record, a successful injection yields one record, not the entire database. The fuel is present but measured in drops, not barrels.
In
Stamps Without Passports, I defined the six requirements for agent identity. They are task-specific, delegation-aware, scope-attenuated, short-lived, cryptographically verifiable, and attestable. Scope attenuation is the mechanism that starves the fuel. When a customer-support agent delegates to an order-processor agent, the authorization server issues a new token with narrower scope: inventory:read:sku-8842 instead of inventory:read. Each delegation hop narrows permissions. The token expires in fifteen minutes. Even if the agent is compromised for the full duration, the blast radius is one SKU for fifteen minutes.
Cedar enforces this at the policy layer. A Cedar policy that restricts an agent to two database tables evaluates structured data against formal logic, 42–60 times faster than Open Policy Agent (OPA)/Rego, sub-millisecond even across hundreds of policies. The model never sees the policy. The model cannot influence it. As I argued in The Linguistic Von Neumann Bottleneck, the policy gateway evaluates the request against deterministic rules and denies it regardless of what the model, or the injection, decided.
The risk tiers from Weighting the Switch determine how much fuel each agent carries. A Tier 1 agent with local, read-only, reversible operations gets a scoped read token. A Tier 3 agent crossing trust boundaries gets a single-operation, single-resource token with a short time-to-live (TTL), or no token at all. The classification is authored once by a security engineer and enforced thousands of times by the harness.
Starving the fuel does not eliminate the trifecta. The agent still processes untrusted input and can still communicate externally. But the fire has nothing valuable to burn. A successful injection against an agent scoped to read one SKU yields one SKU. The economics of the attack collapse.
Removing the heat #
If the agent must access sensitive data and communicate externally, the remaining option is to prevent untrusted content from reaching the privileged context. This is the hardest leg to break cleanly, because the value of many agents is precisely that they process external content. But isolation does not require elimination. It requires separation.
The IBM/ETH Zurich design patterns paper proposes six patterns for this separation. The paper’s core principle deserves direct quotation. “Once an LLM agent has ingested untrusted input, it must be constrained so that it is impossible for that input to trigger any consequential actions.” The most relevant pattern is Plan-Then-Execute, and Anthropic has shipped a production implementation of it. Claude Code’s Plan Mode separates strategy from execution. Rather than asking for approval action-by-action, Claude presents its intended plan up front. The user reviews, edits, and approves the entire sequence before anything runs. Anthropic’s trustworthy agents paper (April 2026) explains the design rationale. “This shifts the user’s level of oversight from the individual step to the overall strategy, which we find tends to be where users most want to exercise judgment.” The plan is fixed before untrusted content enters the context window. A user asks “summarize this email and reply with the key dates.” The orchestrator decomposes this into a fixed plan (read the email, extract dates, compose a reply to the sender) before the email’s content enters any context window. The plan is immutable. The email content can corrupt the body of the reply, but it cannot change the recipient, add new tool calls, or redirect the output. The ignition source is present, but it is contained in a fireproof room.
Claude Code’s auto mode takes this further with a separate classifier model that reviews each action before it runs. The classifier blocks sending sensitive data to external endpoints, production deploys, mass deletion, and IAM permission grants by default. The classifier sees user messages and tool calls but tool results are stripped, so hostile content in a file or web page cannot manipulate the classifier directly. This is defense in depth at the permission layer. The plan constrains the sequence. The classifier constrains each step. Neither relies on the model policing itself.
The Dual LLM pattern, originally proposed by Willison and formalized by Google DeepMind’s CaMeL paper, takes this further. A privileged LLM orchestrates without ever seeing untrusted content. A quarantined LLM processes the untrusted content and returns symbolic variables ($EMAIL_SUMMARY, $KEY_DATES) that the privileged LLM can route without being exposed to the tainted tokens. The quarantined LLM has no access to tools or sensitive data. The privileged LLM has no exposure to untrusted content. The two contexts never merge.
For MCP specifically, Trusting the Label identified the mechanism. Tool descriptions enter the context window as tokens and compete for attention weight against the system prompt. The defense is to pin descriptions at approval time and hash them. Any change triggers re-review. The description is treated as code, not metadata, because the LLM executes it as code.
Removing the heat is the most architecturally complex defense. It requires redesigning the agent’s processing pipeline to separate trusted and untrusted contexts. The tradeoff is real. The agent loses the ability to reason freely across all available information. But the fire cannot start without a heat source, regardless of how much fuel and oxygen are present.
Cutting the oxygen #
The most mechanically straightforward leg to break is the communication pathway. If the agent cannot send data out or modify external state, a successful injection is a fire in a sealed room. The attacker gains influence over the model’s reasoning but has no channel to extract the results.
Default-deny network egress is the bluntest instrument and often the most effective. The agent runs in a container or serverless environment with no outbound network access except to explicitly allowlisted endpoints. A prompt injection that instructs the agent to POST credentials to an attacker’s server fails at the network layer, not the model layer. The model follows the instruction. The firewall blocks the request. The deterministic control does not care what the model decided.
Output sanitization addresses the subtler exfiltration channels. EchoLeak demonstrated that a rendered markdown image tag is an exfiltration vector. The agent embeds stolen data in a URL, the client renders the image, and the HTTP request carries the data to the attacker’s server. Stripping URLs, images, and executable markup from agent responses closes this channel. The agent can still return text. It cannot return content that triggers outbound requests when rendered.
The Interlocking Rule from Weighting the Switch catches the composition attacks that individual controls miss. The session risk accumulator tracks what the agent has accessed, written, and executed. When the agent reads credentials in one step and attempts a network request in the next, the accumulator detects the pattern. Sensitive data in flight combined with a network boundary crossing. It escalates to Tier 3 automatically. The individual operations might each pass their respective tier classifications. The sequence does not.
Tool taxonomy provides the classification layer. Every tool the agent can invoke is tagged by side-effect scope. Read-only, state-modifying, or externally-communicating. Read-only tools run within the agent’s scoped identity. State-modifying tools require explicit authorization. Externally-communicating tools are denied by default in sessions that have processed untrusted content. This is the agentic-patterns.com recommendation formalized. Maintain a machine-readable capability matrix for every tool and enforce it at orchestration time.
Cutting the oxygen is the easiest leg to break and the hardest to break completely. An agent that can return text to a user can encode stolen data in that text. An agent that can create a file can write stolen data to a shared location. The oxygen is never fully absent. But reducing the pathway from “arbitrary HTTP requests” to “sanitized text responses” changes the economics of the attack from trivial to impractical.
Which element to remove #
The fire triangle does not prescribe a universal firefighting technique. It provides a framework for choosing the right technique for the right fire. The same applies here. The choice of which leg to break depends on which capabilities the deployment can sacrifice with the least impact on utility.
def assess_trifecta(session):
has_fuel = session.accesses_sensitive_data()
has_heat = session.processes_untrusted_input()
has_oxygen = session.can_communicate_externally()
if all([has_fuel, has_heat, has_oxygen]):
return apply_firebreak(session, weakest_leg(session))
return PERMITTED
The weakest_leg function is where the architectural decision lives. It returns the leg that costs the least utility to break for the current deployment pattern.
For the decision framework that follows, the three legs map to the fire triangle as defined earlier. Leg 1 is sensitive data (fuel). Leg 2 is untrusted input (heat). Leg 3 is external communication or state change (oxygen).
| Deployment Pattern | Legs Present | Break | Primary Control | Rationale |
|---|---|---|---|---|
| Customer support agent | All three | Leg 3 | Output gating, human review queue, Interlocking Rule | The agent reads customer records (fuel) and processes customer messages (heat). Restrict responses to deterministic templates or route through human review. The agent reasons; a human or template sends. |
| Code assistant | All three | Leg 3 | Ephemeral containers, default-deny egress | The agent reads the codebase (fuel) and processes web documentation or user prompts (heat). Sandbox execution in ephemeral containers with no network access. The agent can read and reason but cannot exfiltrate or execute destructively. |
| RAG knowledge assistant | 1 + 2, Leg 3 weak | Leg 3 | Output sanitization | The agent reads internal documents (fuel) and processes user queries that may reference external content (heat). It returns text, not actions. Sanitize rendered output to strip URLs, images, and markup that could trigger outbound requests. |
| Data analysis agent | 1 + 3, Leg 2 weak | Leg 1 | Cedar policies, RFC 8693 scope attenuation | If all queries come from authenticated internal users, untrusted input is minimal. Scope data access per-query so the agent sees only the tables and rows relevant to the current analysis. A compromised session yields one query’s worth of data. |
| Email/calendar agent | All three, maximally | Leg 2 | Plan-Then-Execute, Dual LLM | The hardest pattern. The agent reads emails (fuel), incoming mail is attacker-controlled (heat), and it sends replies (oxygen). Process email content in a quarantined context. Plan the response structure before the email body enters the privileged context. The email can influence the reply’s content but not its recipient, attachments, or tool calls. |
Two caveats. First, “Leg 2 weak” for the data analysis agent assumes the database contains only trusted data. If the database includes user-generated content (comments, feedback, free-text fields) then Leg 2 is present through indirect injection, and the agent needs a different break. Trace the data back to its source. If any field could have been influenced by an external actor, the heat source is live.
Second, for patterns where all three legs are maximally present and no single leg can be cleanly broken, the defense is depth across all three. Scoped identity reduces fuel. Quarantined processing contains heat. Egress controls restrict oxygen. The Interlocking Rule from Weighting the Switch serves as the safety net, catching the composition attacks that slip through individual controls.
The cost of the firebreak #
A firebreak is not free. Bulldozing a strip of forest destroys the trees in that strip. The defense has a cost, and the cost is utility.
Breaking Leg 1 limits what the agent can help with. An agent scoped to one SKU cannot answer questions about inventory trends across the catalog. Breaking Leg 2 limits what the agent can process. An agent that quarantines email content cannot reason freely about the relationship between an email and a calendar event. Breaking Leg 3 limits what the agent can do. An agent that cannot send emails requires a human to send them, reintroducing the latency the agent was supposed to eliminate.
The calibration challenge is the same one I described in Weighting the Switch. Start tight and loosen based on evidence. An agent that demonstrates safe handling of untrusted input over ten thousand interactions earns broader data access. An agent that never triggers the Interlocking Rule earns less restrictive egress controls. The earned-trust model applies to firebreaks as well as to human access. Prove you need the capability, demonstrate you use it safely, then expand.
The decision framework above is a starting point, not a complete solution. Multi-agent systems, where one agent’s output becomes another agent’s untrusted input, create trifecta compositions that the single-agent framework does not fully address. An agent that individually satisfies the Rule of Two can still participate in a chain where the trifecta emerges across agents rather than within one. That problem, lateral movement through delegated authority, is the subject of a future post.
The triangle and the four pillars #
This post ties together the first five posts in this series. The Linguistic Von Neumann Bottleneck established the thesis. Deterministic controls outside the model, not better prompts inside it. Weighting the Switch provided the risk framework and the Interlocking Rule that catches composition attacks across trifecta legs. The Architecture of Inevitability explained why guardrails cannot break the triangle. The attention mechanism processes all tokens identically, and probabilistic filters share the same architectural limitation. Trusting the Label identified MCP tool descriptions as a heat source that enters the context window from the supply chain. Stamps Without Passports defined the scoped identity that starves the fuel.
The lethal trifecta is the threat model. The blog’s four pillars (isolation as identity, deterministic policy gateways, semantic observability, and the agentic supply chain) are the firebreaks. Each pillar removes or constrains one element of the triangle. The next post will go deeper into the policy layer. Cedar for agent authorization, where formal policy composition for multi-agent delegation chains provides the deterministic enforcement that makes the fuel-starving strategy work at scale. The fire triangle does not care how sophisticated the arsonist is. It cares whether the fuel, heat, and oxygen are present. Remove one, and the fire cannot start.
Thanks for reading Probably Secure. Let’s get to work. Always be curious…all opinions are my own.