↓Skip to main content

Weighting the Switch: Why Human-in-the-Loop Needs a Harness

·14 mins
A railroad dead man's switch weighted down with a toolbox, representing bypassed AI safety controls

In the 1880s, as locomotives grew faster and the stakes of a collision rose, railroad engineers faced a fundamental safety problem: if the engineer suffered a medical emergency or fell asleep, the train became a multi-ton projectile with no one at the brakes. The solution was the dead man’s switch, a spring-loaded pedal the operator had to hold down continuously. If pressure dropped, the system assumed the operator was incapacitated and triggered the brakes. It was an elegant, fail-safe design in theory. In practice, a toolbox defeated it. By placing a heavy weight on the pedal, a tired engineer could “weight the switch,” walking away while the safety system remained technically engaged. The train was not safe. It was merely Probably Secure, a state where the safety mechanism has not failed but the human element within the loop has broken the system’s logic to reduce friction.

The Modern Dead Man’s Switch #

Today, in the world of agentic AI and autonomous coding assistants, we are repeating this history. When we use tools that propose terminal commands or infrastructure changes, we rely on a “Y/N” confirmation prompt as our modern dead man’s switch. The assumption is that keeping a Human-in-the-Loop (HITL) maintains control. In practice, we are weighting the switch. Analyzing the system beyond the UI reveals an information asymmetry that makes meaningful approval impossible. An LLM can ingest hundreds of thousands of tokens of codebase context in seconds to reason through a directory structure, while the human focuses on a localized issue. Every tool-call is a state transition, but for a transition to carry real authorization, the operator must understand both the current state and the intended final state. Because the LLM’s reasoning is a stochastic black box, the human lacks the visibility to perform a true audit in the seconds before clicking “Yes.”

This asymmetry produces a mechanical failure of human vigilance: alert fatigue. If an AI agent asks for permission fifty times an hour, the approval process migrates from analytical thinking to habitual muscle memory. Over a four-hour coding session, that is two hundred prompts. By the second hour, the developer is not reviewing. They are clicking. When a safety mechanism becomes a repetitive friction point, it ceases to be a security control and becomes an obstacle. In the context of a CLI assistant, the “Yes” click becomes a weighted pedal. The human is not in the loop. The human is providing a bypass for a stochastic process.

I learned this firsthand. When I set up my AI agent, it asked permission before modifying files. After the thirtieth approval in a session, I threw a toolbox at it: “trust” mode, granting blanket write access. Because the agent helped with research and writing spread across my machine, I set the allowed directories to my entire home folder. The approvals were slowing me down, and I weighted the switch to get the task done faster. Days later, the agent began modifying configuration files I never intended it to touch. It operated exactly within the permissions I gave it. I had not scoped the access to the folders that mattered; I had removed the friction for all of them. The fix was not to go back to approving every action. It was to define specific folders the agent could modify without asking and require approval for everything else. The architecture was the failure, not the model.

The Hallucination as an Architectural Failure #

The core danger of this stochastic bypass is the hallucination. In agentic workflows, a hallucination is not a factual error; it is a logical deviation that creates a failure domain. A hallucination by itself is inert. It becomes dangerous only when the architecture gives it the identity to act. When a model receives “clear the build cache” but its internal weights trigger a hallucinated path that results in rm -rf /, the failure belongs to the architecture, not the model. In an ambient authority model, the agent inherits the developer’s full credentials: SSH keys, AWS roles, admin rights. If the AI hallucinates a destructive command, the “Yes” click executes that command with full privileges. The blast radius of a single hallucination becomes equal to the operator’s total identity.

What the Railroad Already Learned #

The railroad industry did not solve the dead man’s switch problem by removing the human. It solved it by changing what the human was asked to do, and by adding systems that enforced safety regardless of the human’s input.

The first evolution was the Vigilance Control System. Instead of a single pedal held continuously, modern systems require the engineer to perform varied actions at random intervals: press a button, adjust the throttle, blow the horn. If no contextual input registers within 30 to 60 seconds, the system initiates a three-stage escalation: visual warning, audible alarm, emergency braking. The randomness prevents habituation. The engineer cannot weight a system that asks a different question each time. (FRA Safety Advisory 2015-06)

The second evolution was Positive Train Control (PTC). PTC knows the track ahead: speed limits, signal states, the positions of other trains. It enforces boundaries regardless of what the engineer does. The engineer cannot override a speed restriction by pressing harder on the throttle. Safety moved from the operator’s vigilance into the infrastructure itself. (Federal Railroad Administration)

These two evolutions map directly to the problem of agentic tool-approval. The Y/N prompt is the dead man’s switch: one binary mechanism for every risk level. The vigilance control system is a risk-tiered engagement model where human involvement scales with the stakes. Positive train control is a deterministic guardrail that enforces boundaries architecturally, independent of human input.

The question is not whether to keep the human in the loop. The question is which loop, and for which operations.

The Risk Framework #

Agent operations carry risk across multiple dimensions. Three of the most consequential:

Dimension Low High
Reversibility Can undo (git revert, delete from temp) Cannot undo (sent email, dropped table, pushed to prod)
Blast Radius Local (single file, container) External (production, customer-facing, network egress)
Privilege Required Read-only or scoped write Broad or admin-level access

Real implementations may weight additional factors: data sensitivity, temporal context, the semantic content of arguments, or the agent’s own uncertainty. These dimensions are inputs to a classification decision. The outputs are three control tiers. Each tier pairs a human engagement model with a scoped identity: a short-lived credential issued to the agent with only the permissions that tier requires.

Tier 1, The Natural Reset. Reversible operations with local scope and read-only or scoped-write privilege. Examples: ls, cat, git status, running tests, reading a file. The agent executes within a scoped identity token that physically constrains the blast radius to near-zero. No human prompt. Full audit log. The human’s role is policy author and post-hoc reviewer, not per-action approver. Auto-approve here is not “trust the agent.” It is “the identity makes damage impossible.”

Tier 2, The Vigilance Check. Reversible operations with shared or external scope, or irreversible operations with local scope. Examples: git commit, writing to a non-production S3 bucket, modifying configuration files, running a database migration with a rollback path. The agent receives a scoped identity token broader than Tier 1 but narrower than the developer’s ambient authority, specifically write access to the targeted repository or bucket, not to unrelated resources, credentials, or network endpoints. The system pauses for human approval, but unlike the dead man’s switch, it surfaces full context: the diff, the before-and-after state, the rollback path. The human reviews a specific change, not a generic “Y/N.” Even if the approval is subverted, the scoped identity constrains the blast radius to the resources the operation was classified against. This tier triggers only for medium-risk operations, so the human is not fatigued by fifty prompts an hour. Anthropic’s analysis of millions of Claude Code interactions found that experienced users naturally shift toward this model: auto-approve rates rise from 20% to over 40% with experience, but interrupt rates also increase. Experienced operators shift from rote approval to contextual engagement. The framework formalizes what practitioners already discover through experience.

Tier 3, Positive Train Control. Irreversible operations with shared or external scope, operations requiring broad or admin-level privilege, or any operation the system cannot classify. Examples: rm -rf, IAM policy changes, production deployments, credential access, network egress, any tool outside the agent’s allowed capability set. The agent receives the most constrained identity: a single-operation, single-resource token with a short TTL, or no token at all. The system enforces containment: block, sandbox, or deny execution entirely. The human can authorize execution within those constraints, but cannot override them. The human’s role is to design the policy, not to approve individual commands. Unknown operations default to Tier 3. When the system cannot classify the risk, it applies the brakes.

The obvious objection is that this framework requires someone to correctly classify operations, moving the human problem upstream. Two things make this tractable. First, the classification is authored once by a security engineer and enforced thousands of times by the hook, a fundamentally different cognitive load than approving each action in real time. Second, the default-to-Tier-3 policy means the system fails closed. An unclassified operation gets contained, not approved at full privilege.

The Interlocking Rule #

The tiers classify individual operations. They do not address what happens when an agent chains a Tier 1 read of ~/.aws/credentials into a Tier 2 curl to an external endpoint. Individually, both operations pass their respective tiers. Together, they constitute credential exfiltration.

The railroad industry solved this problem with the interlocking system: a mechanism that connects switches and signals so that no combination of individually-safe settings can create a conflicting route. A route can only be set for a train if the entire sequence is safe.

CrowdStrike documented three attack patterns that exploit this gap: tool poisoning (hidden instructions in tool metadata), tool shadowing (one tool’s description influencing how the agent uses a different tool), and rugpull attacks (tool behavior changing after integration). The pattern: individual tools are secure in isolation, but their composition creates emergent risk.

The framework addresses this with an Interlocking Rule. The SDK hook maintains a session risk accumulator, a running record of what the agent has accessed, written, and executed. Each new operation is classified not only by its own properties but by its relationship to what came before it. Three patterns trigger automatic tier elevation:

Trigger Detection Effect
Sensitive Data in Flight Agent read credentials, PII, or secrets in a prior step and now attempts network egress or a write to shared scope Elevate to Tier 3
Write-then-Execute Agent wrote a file and now attempts to execute it Elevate to Tier 3
Trust Boundary Crossing Sequence moves from local → external scope, or from read → write → execute Elevate one tier per boundary crossed

Unknown sequences default to Tier 3. When the accumulator cannot classify the relationship between operations, it applies the brakes.

The Deterministic Harness #

For operations that fall into Tier 3, the architecture replaces permission-based security with deterministic guardrails centered on scoped agent identity. As Marc Brooker has argued, the right way to control what agents do is to put them in a box: a strong, deterministic layer of control outside the agent. The tiers above define what goes in the box. The harness below is the box itself. Trust begins with least privilege. In practice, building agent harnesses has taught me that the single highest-leverage control is ensuring an agent never inherits the developer’s full ambient authority. Instead, the system generates task-specific credentials: cryptographically bound, short-lived, scope-restricted tokens for one specific operation. When an agent receives “refactor CSS,” its identity restricts it to read on the source and write on .css files. Zero access to the network. Zero access to .aws/credentials. The developer does not audit every command because the agent physically cannot exfiltrate secrets. Safety lives in the identity, not the approval button.

The mechanism that routes operations to the correct tier is the SDK hook. By intercepting the agent’s intent before execution, the hook performs a deterministic risk assessment: evaluating the operation against the risk framework and interlocking rule. This is not a vibe check by another LLM. It is a code-level analysis of the operation’s potential impact. If the hook classifies the command as a legitimate but high-risk operation, it automatically exchanges the agent’s broad identity for a task-specific credential using a Security Token Service (STS). This process architecturally compresses the failure domain of a hallucination. The model can hallucinate as creatively as it wants, but it lacks the identity to manifest that error into a system-wide catastrophe.

Once the hook classifies the risk and scopes the identity, the system applies tiered isolation. Low-risk read commands run in a standard container with a restricted identity. High-risk commands with destructive potential never execute in the host environment. Depending on the operation, the harness denies the action outright, routes it to an ephemeral sandbox, or evaluates it against a policy engine like Cedar before permitting execution. The specific isolation technology matters less than the architectural guarantee: the failure domain is physically contained regardless of human input. Even if the developer clicks “Yes” on a hallucinated command, the sandbox constrains the blast radius to a throwaway environment with a restricted identity.

The following pseudocode demonstrates the routing layer, how the hook classifies each operation and applies the Interlocking Rule against session history:

TIER_1, TIER_2, TIER_3 = 1, 2, 3

def route_operation(op, session):
    """
    The Routing Layer: classify the operation, then apply
    the Interlocking Rule against session history.
    """
    tier = classify(op.reversibility, op.scope, op.privilege)

    # Interlocking: sensitive data in flight
    if session.touched_secrets() and op.crosses_network_boundary():
        tier = TIER_3
    # Interlocking: write-then-execute
    if session.wrote(op.target) and op.is_execute():
        tier = TIER_3
    # Interlocking: trust boundary crossing
    if op.scope > session.max_prior_scope():
        tier = min(tier + 1, TIER_3)

    session.record(op)
    return tier

Once an operation is classified as Tier 3, the harness enforces it through credential downscoping: an STS AssumeRole call with an inline policy scoped to exactly the operation at hand, a fifteen-minute token that can write to one S3 bucket and nothing else. The implementation is a standard credential exchange; the novel architecture is the routing decision of when and why to invoke it.

When the harness needs a brain #

The controls described above are deterministic and transitional. The Interlocking Rule catches three known patterns. An attacker who discovers a fourth walks through the harness unchallenged. Static rules do not scale to the combinatorial space of tool-call sequences across multi-agent systems. At some point, risk assessment itself requires AI.

The question is when that AI has earned the trust to participate. The same earned-trust model that governs human access in well-run organizations applies here. A new AI component in the harness starts at Tier 3: sandboxed, advisory-only, every recommendation logged and validated against a formal policy engine. The function it serves is pattern recognition at scale, identifying novel attack compositions, anomalous tool-call sequences, and risk correlations that static rules cannot anticipate. The implementation varies: Automated Reasoning checks translate natural language policy into formal logic and use mathematical proofs to verify LLM outputs. Observer LLMs monitor agent behavior in real time and flag deviations. Graph-based anomaly detection and runtime behavioral modeling catch multi-step attack chains that no predefined trigger covers. What matters is not the mechanism but the constraint: the AI component perceives, the formal policy engine enforces. Trust accrues from the gap between them shrinking over time, measured in production, not promised in a whitepaper.

Once that trust is earned, the classify() function and the Interlocking Rule no longer have to remain static pattern matchers. But the architectural constraint is non-negotiable: the AI component is advisory, the formal policy engine is authoritative. The failure mode is a false positive, an unnecessary escalation, never a false negative that lets a dangerous operation through. You can put a brain in the harness as long as you put a harness around that brain.

The End of Probably Secure #

The dead man’s switch failed not because humans are unreliable, but because a single binary mechanism was applied to every mile of track regardless of the terrain. The Y/N prompt fails for the same reason. If your security architecture relies on a human never being tired, pressured, or distracted, you have built a Probably Secure system. The answer is not to remove the human from the loop. The answer is to stop asking the human the same question regardless of the stakes, and to build architecture that enforces safety where human vigilance cannot.

The focus belongs on the harness: the deterministic infrastructure that matches the control to the risk. The dead man’s switch era is over. The question is whether we replace it with a vigilance system or find a heavier toolbox.

Thanks for reading Probably Secure. Let’s get to work. Always be curious…all opinions are my own.