↓Skip to main content

The Architecture of Inevitability

·16 mins
A forged painting under examination, representing how prompt injection bypasses AI guardrails

In 2011, the Knoedler Gallery in New York closed its doors after 165 years in business. Over the previous fifteen years, the gallery had sold approximately $80 million in paintings attributed to Mark Rothko, Jackson Pollock, Robert Motherwell, and other Abstract Expressionists. Every painting was a forgery. The works had passed expert authentication. Trained eyes evaluated brushwork, pigment composition, canvas age, and stylistic consistency, and declared them genuine. The experts were probabilistic classifiers. They assessed what the paintings looked like, not where they came from. A provenance system, an unbroken, documented chain of custody from the artist’s studio to the gallery wall, catches every forgery on day one, because it does not evaluate appearance. It verifies origin.

The AI security industry is building higher-fidelity art experts. It calls them guardrails. To understand why they fail, we need to understand what they guard against. The standard explanation gets that wrong.

The symptom and the mechanism #

I talk to penetration testers regularly, across customers and within Amazon. They already know. Prompt injection is not a bug in the model. It is a behavior of the architecture, and it takes them minutes to prove it. A NIST-partnered red teaming competition in 2025 confirmed this at scale. Nearly 2,000 participants attacked 22 frontier AI agents from OpenAI, Anthropic, and Google, and every agent failed in every category tested. On the curated attack set, policy violations occurred 20 to 60 percent of the time on the first query; after ten queries, nearly every attack succeeded. The pen testers know this. The security architectures deployed around these models do not reflect it.

Ask an engineer what prompt injection is and you will hear a variation of “the model ignores its instructions.” This describes the symptom. It does not explain the mechanism. Saying “the model ignores its instructions” is like saying “the patient has a fever.” It tells you something is wrong without telling you why.

Prompt injection exploits a specific architectural property. The Transformer’s self-attention mechanism processes all tokens in the context window identically, with no metadata indicating origin or authority. A system prompt token and a user input token become vectors in the same high-dimensional space. The attention mechanism computes relationships between them using the same mathematical operation. The architecture has no mechanism to mark one as “instruction” and the other as “data.” An attacker who crafts input tokens that produce high relevance scores in the attention computation can shift the model’s behavior away from the system prompt and toward the attacker’s intent.

This is not a new category of error. It is the oldest category of error in computer science, wearing new clothes.

In 1998, SQL injection exploited the same structural weakness. Code and data occupied the same string. Parameterized queries fixed it. The database engine treats the parameter as data, not as code. Deterministic. Complete. As I argued in The Linguistic Von Neumann Bottleneck, the same structural problem runs deeper. Von Neumann’s stored-program architecture put instructions and data in the same memory, and the NX bit enforced a hardware boundary. Memory is writable or executable, not both. Deterministic. Complete.

Prompt injection is the third iteration. Instructions and data occupy the same context window. But this time, there is no equivalent of parameterized queries. There is no NX bit for tokens. The UK’s National Cyber Security Centre made this explicit in their December 2025 advisory. Prompt injection “may never be totally mitigated” as a category, and treating it like SQL injection is a dangerous category error. SQL databases are deterministic engines with a clear code/data boundary. LLMs are not. To understand why, you have to open the Transformer.

Inside the attention mechanism #

The Transformer processes input through a pipeline. As I described in The Linguistic Von Neumann Bottleneck, the tokenizer and embedding layer convert the entire input into a single sequence of token vectors, discarding origin in the process. A token from the system prompt and a token from user input are geometrically indistinguishable, points in the same vector space with no label indicating where they came from. Modern LLMs, including GPT, Claude, and LLaMA, use only the decoder half of the original Transformer architecture, processing the entire conversation, system prompt, user input, and model responses, as one flat token sequence with causal self-attention. The original Transformer had an encoder-decoder split that could theoretically separate input processing from output generation. The industry chose decoder-only because it scales better and generalizes across tasks. The cost of that choice is that no structural boundary exists between instructions and data.

The self-attention mechanism then computes context. For every token in the sequence, it calculates how much every other token should influence it. Most explanations name the three projections involved, Query (Q), Key (K), and Value (V), without explaining why there are three or what each one does. The answer matters for understanding why injection works. Q, K, and V are three separate linear projections of the same token embedding, each defined by a different matrix of weights learned during training. Three projections exist because the mechanism needs to decouple two operations, deciding which tokens are relevant to each other and deciding what information flows between them. Q and K handle relevance, and they are deliberately asymmetric. A token’s Query vector represents what that token needs from context, the question it asks of the sequence. A token’s Key vector represents what that token offers to other tokens looking for context. The dot product between a Query and a Key measures how well one token’s need matches another token’s offer. A high dot product means high relevance. V handles the second operation. Once the relevance scores determine which tokens to attend to, the Value vectors carry the actual content that flows between them. The three weight matrices are learned from data, not designed. The model discovers which token relationships reduce prediction error, and Q, K, V encode those relationships. The following formula gives the attention weight between any two tokens.

$$ \alpha_{ij} = \frac{\exp(Q_i \cdot K_j / \sqrt{d_k})}{\sum_k \exp(Q_i \cdot K_k / \sqrt{d_k})} $$

The numerator is the Q·K compatibility score, scaled by the square root of the key dimension to prevent large values from saturating the softmax. The denominator is the softmax normalization across every token in the sequence, forcing all attention weights to sum to exactly 1.

The math is not the point. The point is what it produces. Every token in the context window competes for influence over every other token, and the competition is determined entirely by content, not by origin. A system prompt token and a user input token enter the same computation with the same weight matrices and the same rules. Modern chat models use special tokens like <|system|> and <|user|> to mark role boundaries, and training (SFT, RLHF) teaches the model to weight system instructions more heavily. But this is a learned preference, not an architectural guarantee. It is a soft bias encoded in the weight matrices, and an adversary with gradient access can find inputs that override it. The architecture has no concept of trust, authority, or precedence. It has dot products.

This is the mechanism behind prompt injection, and the detail that the standard explanations miss. Softmax creates a zero-sum competition for attention weight. Every point of attention that one token gains is a point that every other token loses. When the model generates the next output token, it attends to the entire context window through this competition. System prompt tokens compete directly with user input tokens for influence over the output. There is no reserved allocation. There is no priority lane. The competition is purely mathematical. Whichever tokens produce the highest Q·K dot products win.

An adversarial suffix exploits this directly. The Greedy Coordinate Gradient (GCG) algorithm uses gradient-based optimization to find token sequences that maximize the probability of the model producing an attacker-desired response. The gradients flow through the entire model, including the attention mechanism, finding tokens that shift the model’s behavior away from the system prompt. The original GCG attack produces suffixes that succeed roughly 2% of the time against commercial models. Generating thousands of candidates is a one-time computational cost. Subsequent work like AmpleGCG achieves near-100% success rates on open-weight models by training generative models to produce adversarial suffixes at scale. The system prompt does not disappear. The adversarial tokens drown it out. Softmax redistributes attention weight accordingly, away from your instructions and toward the adversary’s.

Two symptoms, one architecture #

The industry treats prompt injection and hallucination as separate problems. Prompt injection is an adversarial attack. Hallucination is a reliability issue. Different teams work on them. Different tools address them. They share the same architectural root cause.

Ji and others (2025) identify what they call “Artificial Certainty.” The Transformer’s softmax function collapses ambiguous attention scores into a single probability distribution, discarding uncertainty information at each layer. When the model encounters ambiguous or conflicting evidence in the context window, softmax forces it to commit. It must allocate exactly 100% of its attention across tokens. It cannot say “I am uncertain which tokens to attend to.” It cannot abstain. The architecture demands a probability distribution, and a probability distribution demands commitment.

This forced commitment is the mechanism behind hallucination. When the model generates the next token, it selects from a probability distribution over the entire vocabulary. Even when the evidence is ambiguous, when several tokens are plausible and the model’s internal representation does not decisively favor one, softmax produces a confident-looking distribution. The model outputs a token that reads as certain even when the underlying computation was not. Research confirms that models hallucinate with high confidence even when they possess the correct knowledge. The architecture overrides the knowledge.

The same forced commitment is the mechanism behind prompt injection. When adversarial tokens produce high Q·K dot products, softmax redistributes attention weight toward them. The model does not “choose” to follow the injection. The attention mechanism mathematically shifts weight to the tokens with the highest relevance scores, and the attacker optimizes the adversarial tokens to score high. The model commits to attending to the adversarial tokens because softmax demands commitment.

Two symptoms. One architecture. In hallucination, the model commits to a token when it lacks sufficient evidence. In injection, the model commits to attending to adversarial tokens because they win the softmax competition. Both exploit the same property, forced commitment over a flat token sequence with no origin metadata. The hallucination literature identifies many contributing factors beyond attention, including training data quality, knowledge cutoffs, decoding strategy, and RLHF-induced sycophancy among them. The architectural argument here is narrower and specific. Softmax-forced commitment is a shared mechanism that neither problem can escape within the current architecture. The attention mechanism offers no way to distinguish uncertain computation from confident computation, or legitimate tokens from adversarial ones.

In agentic systems, this convergence matters practically, and the threat extends beyond direct prompt injection. Indirect prompt injection, where the malicious payload hides in a document, email, or web page the agent processes, is the more dangerous variant. The NIST competition data confirms this. Indirect attacks succeeded 27.1% of the time compared to 5.7% for direct attacks. The agent reads a poisoned document and the document’s tokens compete with the system prompt’s tokens through the same softmax competition. No adversary needs access to the prompt. They need to place content where the agent will read it.

A hallucinated tool call is functionally equivalent to a successful prompt injection. The model takes an action it did not have authorization to take. The difference is intent. No adversary is required for a hallucination. But the architectural cause is identical, and the damage is the same. If your agent hallucinates a database deletion with the same confidence it uses to format a query, the distinction between “hallucination” and “injection” is academic. The blast radius is real.

Steel and concrete #

A guardrail on a highway is steel and concrete. It does not evaluate the car’s intent or assess the probability that the car is leaving its lane. It enforces the boundary through physics. It works at 2 AM when the driver is asleep. It works against a driver who actively wants to cross the median. It works because it operates on a different plane than the driver’s decision-making.

The systems the AI industry calls “guardrails” share none of these properties. They are probabilistic classifiers, frequently LLMs or ML models themselves, that process natural language input and return a confidence score. They evaluate what the input looks like, not where it came from. They are the art experts at the Knoedler Gallery, assessing brushwork and declaring authenticity.

The architectural problem is recursive. An LLM-based guardrail processes tokens through the same self-attention mechanism as the model it guards. It suffers from the same softmax-forced commitment. The same adversarial techniques fool it. You have taken a probabilistic system and asked another probabilistic system to reliably police it.

The evidence confirms this. Mindgard and Lancaster University (2025) tested six production guardrail systems from Microsoft, Nvidia, Meta, Protect AI, and Vijil. They bypassed every one using rudimentary techniques, including character obfuscation, adversarial perturbations, and emoji smuggling, which embeds hidden payloads within emoji characters that current guardrails fail to detect. The attack success rate reached 100% against Protect AI v2 and Azure Prompt Shield. These are not theoretical attacks against research prototypes. These are production-grade defenses deployed by the largest AI companies in the world, defeated by emojis.

The strongest counterargument is training-based defense. OpenAI’s instruction hierarchy trains models to prioritize system-level instructions over user input, and their IH-Challenge benchmark shows up to 15 percentage point improvement in resistance to injection attacks on benchmarks like TensorTrust. This is a real improvement, not a marketing claim. But probability compounds over attempts. Gray Swan’s benchmark of Claude Opus 4.5, the highest-scoring model in their evaluation, found a 4.7% success rate per attempt. At ten attempts, 33.6%. At one hundred attempts, 63%. For comparison, if each attempt were independent, the Bernoulli calculation at 4.7% gives 38% at ten attempts and 99.2% at one hundred. Gray Swan’s observed numbers reflect their specific attack methodology, but the direction is the same. Probability compounds over attempts, and the curve bends toward certainty. An agent that processes two hundred requests per day faces those odds daily. Training-based defenses shift the curve. They do not change its shape. The probability of successful injection over repeated attempts still converges toward certainty, because the architecture has not changed. Softmax still forces commitment over a flat token sequence, and the adversary still gets to run gradient-based search against it.

The term “guardrail” is doing real damage. When a security team hears “we have guardrails,” they hear “we have a barrier,” not “we have a probabilistic filter that a motivated attacker can bypass with Unicode characters.” Call them what they are. A system that processes natural language and returns a confidence score is a probabilistic filter or a content classifier. Reserve “guardrail” for controls that enforce boundaries regardless of input. These include policy engines that evaluate structured data against formal logic, identity systems that cryptographically bind scope and delegation, and isolation boundaries that physically contain the impact. These are the steel and concrete. The probabilistic filters are the painted lane markings.

Probabilistic filters belong on every model invocation. They catch the bulk of unsophisticated injection attempts, enforce content policies, filter toxic output, and provide configurable controls that adapt to your application’s risk profile. In a defense-in-depth stack, they are the first layer, not the last. The mistake is treating them as the last.

The harness, not the filter #

In the 1960s, Shigeo Shingo at Toyota observed that end-of-line quality inspectors caught most defects but not all, especially novel ones the inspector had not seen before. His insight was that inspection is inherently reactive. He developed poka-yoke, or mistake-proofing. The principle is to redesign the manufacturing process so that defects cannot physically occur. A part that can be inserted in one orientation cannot be inserted wrong, regardless of the inspector’s vigilance. Do not inspect quality in. Design it in.

The same principle applies here. Since prompt injection is an architectural property of the Transformer, not a defect to be patched, the defense must also be architectural. Filtering the input does not change the mechanism that processes it. The defense must operate outside the context window, where the attention mechanism cannot reach it.

The Linguistic Von Neumann Bottleneck details the full architecture. Treat the model as an untrusted execution core. Give agents their own cryptographic identities through RFC 8693 token exchange. Enforce authorization through formally verified policy languages like Cedar, which evaluates structured requests against deterministic logic 42–60 times faster than OPA/Rego, sub-millisecond even across hundreds of policies. As I argued in Weighting the Switch, the control must match the risk. Auto-approve scoped read operations. Require contextual human review for state-modifying operations. Enforce deterministic boundaries for irreversible or high-privilege actions regardless of what the model, or the human, decides.

None of these controls prevent prompt injection. They make prompt injection survivable, which is the correct framing for a threat that cannot be eliminated. The Knoedler analogy is instructive but imprecise. A true provenance system would verify the origin of every token, marking system instructions as authoritative and user input as untrusted. The architecture cannot do this. What Cedar and RFC 8693 provide is the next best thing, containment. The injection succeeds inside the model, but the scoped identity limits what the injection can do outside it. An attacker can inject the model, but the policy gateway evaluates the request against formal rules and denies it. The scope of a compromised model shrinks from “everything the user can access” to “everything the agent’s scoped token permits for the next fifteen minutes.” This is the difference between a forgery that fools an expert and a forgery that fails a provenance check. The expert is fallible. The chain of custody is not.

The tradeoff #

These controls are not free. I covered the costs in detail in The Linguistic Von Neumann Bottleneck. They include latency on every request, operational complexity for identity management, and reduced agent flexibility when permissions are scoped tightly.

But the alternative is not “no cost.” The alternative is a 63% breach probability at one hundred attempts against the best available model, with no upper bound on the number of attempts an attacker can make. The cost of probabilistic security is not the cost of the filter. It is the expected value of the failure the filter misses. For an agent with access to production databases, customer data, or financial systems, that expected value dwarfs the cost of a sub-millisecond policy evaluation on every request.

The Transformer’s attention mechanism is elegant. It produces the capabilities that make LLMs useful. It is also the reason prompt injection exists and the reason hallucinations occur. These are not bugs to be patched in the next model release. They are properties of the architecture. The industry will continue building higher-fidelity probabilistic filters, and researchers will continue bypassing them, because the filters operate on the same mathematical foundation as the models they guard. The arms race between forger and expert has no end.

Build the provenance system instead. Upcoming posts will go deeper into the linear algebra of adversarial suffixes, Cedar policy composition for multi-agent delegation chains, and the lethal trifecta pattern that makes certain agent architectures inherently exploitable. The chain of custody does not care how good the forgery is.

Thanks for reading Probably Secure. Let’s get to work. Always be curious…all opinions are my own.