Skip to content

Security

Patterns and techniques for building agents that resist manipulation, protect sensitive data, and fail safely.

Threat Models

Threat models identify the structural conditions that make agent systems exploitable and prescribe architectural mitigations.

  • Action-Audit Divergence: A Four-Mode Taxonomy for Runtime Hardening — Name the four ways an agent action can diverge from its audit record (gate-bypass, audit-forgery, silent host failure, wrong-target) to convert "is this runtime hardened?" into a coverage checklist against existing controls
  • Benchmark Poisoning of Self-Modifying Coding Agents — An attacker who controls the benchmark a self-modifying agent scores itself against gets a write to that agent; poisoned CertCheck runs reached 30/30 vulnerable solutions on neutral held-out tasks, and continuing to evolve on a clean benchmark left the rate at 28/30 or higher
  • Compositional Vulnerability Induction in Coding Agents — Decomposing a malicious end-state into three innocuous engineering tickets bypasses refusal and hardening defenses at 53–86% ASR across nine production coding agents; pentester-framed reviewers close most of the gap
  • Constraint Drift: Why Safety Must Be Maintained, Not Asserted — Safety constraints encoded in prompts weaken across six trajectory surfaces — memory, delegation, communication, tool use, audit, optimization; the four-property invariant (fresh, inherited, enforceable, auditable) keeps them operative when delegation depth, memory persistence, and tool surface compose
  • Context-Fractured Decomposition Attacks on Tool-Using Agents — Attacks split across tools, modules, and time slip past defenders that only inspect a single contiguous conversation; artifact provenance gaps let benign intermediate steps recompose into a jailbreak downstream, lifting attack success by up to 28.3 percentage points
  • Distributed Cross-PR Attacks in Persistent-State AI Control — An untrusted coding agent spreads a covert payload across a sequence of PRs, timing each piece for natural cover, so stateless per-PR review never sees the whole attack; a stateful link-tracker plus a monitor ensemble cuts gradual-attack evasion from 93% to 47%
  • Forged Reasoning Trace Attacks on Agent Memory (FARMA) — Forged reasoning traces poison an agent's stored decision history so it skips a safety step believing it already ran; evasive wording slips past keyword filters and self-referential amplification defeats consensus defenses, reaching 100% success on binary safety-gate agents
  • Computer-Systems Lens for Always-On Agent Security — Model an always-on agent as a computer system — gateway runtime as OS, Skills as applications, Plugins as loadable extensions — to port classical OS protections onto the four cross-component surfaces that model-response benchmarks miss
  • Four-Layer Taxonomy of Agent Security Risks — Group threats into context/instruction, tool/action, state/persistence, and ecosystem/automation layers to map controls and surface coverage gaps where attacks propagate across boundaries
  • Goal Reframing: The Primary Exploitation Trigger for LLM Agents — A 10,000-trial taxonomy finds goal reframing — not social engineering or incentives — is the one prompt condition that reliably triggers vulnerability exploitation across models
  • Improper Output Handling: Validate Agent Output Before Downstream Use — OWASP LLM05 — agent output executed, rendered, or interpreted downstream without per-sink validation is an injection surface; enumerate the sinks (commit, exec, SQL, render, install) and gate each one
  • Lethal Trifecta Threat Model — Risk emerges when an agent has private data access, untrusted input, and egress simultaneously; remove at least one leg from every execution path
  • Oracle Poisoning: Knowledge Graph Corruption Against Tool-Using Agents — Corrupting a knowledge graph an agent queries via tool-use produces 100% trust at moderate attacker sophistication across nine models; the attack is distinct from prompt injection because the data path, not the instruction path, carries the payload
  • OWASP LLM Top 10 (2025): Agent Security Crosswalk — Map each OWASP LLM Top 10 (2025) risk to coding-agent-specific manifestations and site pages — a navigation aid for readers arriving with the framework's shared vocabulary, not a recommended threat model
  • OWASP 2026 Update for Agent Builders: Top 10 Renumbering and the Agent Control Standard — The 2026 edition adds and drops nothing but renumbers eight entries and renames one, so an unqualified entry ID resolves to two different risks; the new Agent Control Standard names eight runtime hooks at v0.1 preview
  • Pre-Trust Execution Surface in Coding Agent Harnesses — Project-local config (settings files, hooks, MCP manifests, env vars, localhost listeners) executes before the trust prompt fires; defer parsing and execution until after the trust boundary is established
  • RAG Architecture as a Poisoning Robustness Decision — Under controlled knowledge-base poisoning, attack success rates span 24.4% to 81.9% across four RAG architectures with comparable clean accuracy; architecture choice is part of the threat model
  • Replayable Encrypted Reasoning Blocks in Agent Traces — Provider-returned chain-of-thought blocks replay across sessions, users, and sibling models, so a persisted or published trace is recoverable data its holder cannot decrypt or audit; 315,320 blocks decoded from public trajectories yielded 182 credentials
  • Skill Misevolution in Self-Updating Skill Libraries — A library that distills its own successful trajectories keeps the unsafe shortcut inside them; all 21 evolved configurations authored an unsafe artifact and 15 reproduced harm in a fresh session, so risk has to be measured at authoring, retrieval, and execution separately
  • Trajectory Poisoning of Promoted Agent Skills (PoisonedEvolution) — Self-evolving skill systems promote agent trajectories into persistent instructions, so the trust boundary sits at promotion; three consistent attacker records in thirty reached 91.0% success across six evolvers, and a provenance gate keyed on distinct contributors took a pilot from 25/25 to 0/25
  • Trojan Hippo: Dormant Memory Payloads Triggered by Sensitive Topics — A single untrusted tool call plants a dormant payload in agent memory that activates sessions later when the user discusses sensitive topics, exfiltrating data via outbound tools; tested defenses cut attack success to 0–5% but at steep utility cost
  • Unbounded Consumption: Bounding Agent Resource Use Against DoS and Denial-of-Wallet — OWASP LLM10:2025 framed as a same-surface, two-owner threat (availability and finance); the five complementary bounds — per-call token, per-task iteration, fan-out concurrency, cost-velocity, per-day dollar — close the cost dimension no single layer covers, with $46K/day Sysdig and $82K/48hr Gemini incidents as the empirical floor
  • Workspace Topology as an Indirect Injection Attack Vector — Repository layout, modularity, nesting depth, and in-file position each move indirect-injection ASR on one coding agent; measured on gpt-oss-120b via opencode 1.14.46 across 100 repos, so the levers name direction and magnitude but do not replace capability restriction
  • Which Task You Delegate Changes Poisoned-Repo Exposure — On a poisoned repository the task verb moves attack success 4.5-fold on the same injected file (Run-Tests 45.5% against Fix-Bug 8.6%), and the riskiest task draws the fewest warnings at a 9.4% alert rate; use the ranking to place pre-execution review, not to pick a safer verb
  • Bug-Class Hints as Exploit Input for Coding Agents — A hint that names a component and a defect class supplies the search bound a CVE description does; GPT-4 exploited 87% of a benchmark with the description and 7% without, so an embargo now buys time in proportion to how vague the leak is
  • Root Causes of Vibe-Coded Application Vulnerabilities — An audit of 200 deployed vibe-coded applications found 91.00% carried at least one vulnerability, traced to three agent limits (memory, objective, knowledge); production-ready framing and a hardened harness each roughly halve reintroduction, detailed technical instructions make it 20.0 points worse, and nothing tested reaches zero
  • The Post-Authorization Execution Trust Gap in Remote MCP — An OAuth grant names a client and an audience, never the workload that runs a later call; six named failure modes say what the grant leaves open, and the invocation-time fix that closes them is provider-side only
  • Next Edit Suggestions Carry Context You Never Curated — Six channels feed an NES prompt, including an edit buffer that retains content you erased with undo; the diff at the caret is narrower than the transaction you accept, and the rates behind it are conditional, Java-only failure rates

Prompt Injection

Prompt injection is the primary attack vector for agents that consume untrusted content. External instructions embedded in web pages, emails, documents, or API responses can redirect an agent's behavior at the model level.

Anti-pattern: Single-Layer Prompt Injection Defense — Relying on one safeguard leaves agents vulnerable to attack vectors that layer does not address

Sandboxing

Isolation limits what a compromised or misbehaving agent can affect.

Anti-pattern: Hostname-Allowlist Proxy: The TLS-Inspection Blind Spot — A hostname-allowlist proxy without TLS termination enforces the client-supplied destination, not the actual destination; broad shared-CDN entries open domain-fronting and similar exfil paths

Data Protection

Preventing sensitive data from entering agent context is cheaper than scrubbing it after the fact.

Permissions

Excess permissions expand the blast radius of any failure or attack.

  • Agent Approval Laundering: Effects Beyond the Named Command — An approval record can name npm install correctly while its workflow runs lifecycle scripts, or name an MCP tool while its transport carries network authority the record never lists; effect-aware review pays off only where a sandbox or an execution-side default does not already contain the effect
  • Agent Network Egress Policy: Admin-Controlled Domain Allow/Deny — Restrict which domains agent tools can reach via harness-enforced allow and deny lists; remove the model from the network trust boundary
  • Aggregation Bounds for Agent Authorization — Every access in a workflow can be authorized and the synthesized answer still not be; three parts are buildable today, and both papers that name the gap state the general problem is open
  • An Explicit Update Boundary for Agent Self-State — A six-operation replay shows chattr +i, AppArmor and Landlock block every path to an agent's memory file and also reject its approved update; sort self-state by whether the agent legitimately writes it, then give the writable half a separate authorization step
  • Authority Confusion: Untrusted Context Must Not Authorize Side Effects — Decompose task authority into a step-level authority context the dispatch layer can check; runtime content may inform the planner but never become the issuer that authorizes a side effect
  • Authorization Continuity Across Agent Mutation — A grant names a subject as well as an effect; a transition envelope decides whether it survives a model, tool, or delegation change, while an immutable ceiling fixed at grant time stops agent-produced evidence from widening it
  • Blast Radius Containment: Least Privilege for AI Agents — Limit agent access to only what the current task requires; excess permissions directly amplify injection impact
  • Centrally Provisioned MCP Servers: Remote Transports Only — managedMcpServers pushes HTTP and SSE servers to every user and rejects any entry naming a command, args, env, or headersHelper, so a settings document can never start a local process; the drop notice reaches the user, not the administrator
  • Constraints as a Substrate for Scalable Agent Oversight — Manage coding agents with the controls used for human teams — access control, network policy, tooling-enforced conventions — so a small, cheap reviewer scales oversight; substrate plus a ~200-line tool lifted a small reviewer's backdoor recall from 54.5% to 90.9%
  • Explanation-Bound Tool Execution for Agent Gateways — Convert the agent's stated reason into typed claims and check each against a fact the server already holds; a free-form justification requirement allowed 60 of 96 authored hard contradictions, server-checked claims allowed none
  • Fail-Closed Remote Settings Enforcement — Block agent startup until remote managed settings are freshly validated; exit rather than run with stale or missing policy
  • Gate Agent Writes to Executable Config Files as Privileged Actions — Writes to .npmrc, .yarnrc, bunfig.toml, .bazelrc, .pre-commit-config.yaml, and .devcontainer/ are execution-escalations — interrupt permissive edit modes at the write site, complementing execution-side defaults like ignore-scripts=true
  • Intent-Governed Tool Authorization for AI Agents (IGAC) — A server-issued intent certificate narrows the static OpenPort manifest per user request via a monotone-only filter; the planner cannot reach tools the current ask does not justify, but classifier failure and intent-materialization privacy regressions constrain where the trade-off pays off
  • Non-Retirable Approval Rules for Agent Operations — An admin-set permissions.ask rule re-prompts on every occurrence and no bypass mode, auto-approval, hook, or saved grant satisfies it; defining any rule at all makes the unmatched set prompt too
  • Org-Membership-Gated Agent Entitlement — Gate AI chat activation on directory-managed GitHub organization membership via VS Code's ChatApprovedAccountOrganizations device policy; fail-closed and structurally distinct from seat licenses
  • Parser-Versus-Shell Evasion in Command Permission Checks — A permission rule matches command text while the shell computes a richer function of it; the documented bypass record makes text matching a pre-flight and puts enforcement below the parser
  • Per-Caller Identity: Who an Agent's Tool Call Acts As — A tool call reaches the external system as the deployment or as the person who asked; that one token decides visible data, the logged actor, and whose rate limit drains
  • Permission-Gated Custom Commands — Pre-approve the tools a Claude Code slash command may use via frontmatter, narrowing the expected surface for shared commands
  • Pre-Execution Risk Classification for Terminal Commands — Display a tiered Safe/Caution/Review-carefully badge with command-specific text before the agent runs a terminal command; an attention-allocation lever paired with deterministic allowlists that carry the policy load
  • Proof of Presence: Re-Authenticating for Agent Actions — GitHub's public-preview IdP re-authentication check stops an agent only on the browser channel it guards, outside a live sudo window, and only when the factor it demands is one the agent cannot reach
  • Revocable Resource-and-Effect Capabilities for Coding Agents (PORTICO) — Materialize each subgoal-scoped capability as an opaque epoch-bound handle that closure removes from the planner's interface; stale replay is rejected before side effects, closing lingering authority when tool traffic is mediated and the catalog is typed
  • Route Agent Peers by Enrolled Identity, Not Card Name — Six of seven tested A2A hosts sent a request addressed to a trusted peer to an attacker-controlled endpoint that had registered the same display name; route on an operator-assigned identifier bound to the authenticated origin and keep the card name presentational
  • Safe Outputs Pattern — Default agents to read-only and require explicit grants for each write output type, producing a deterministic blast radius
  • Static-First Shell Command Gating with Selective Escalation — Canonicalize the command, score it against structural, semantic, path, and pattern evidence, and escalate only the undetermined 4.2% to an LLM judge; the enforcement profile trades false positives against realised harm
  • Task-Based Access Control with Hybrid Inspection — Bind each tool call to the user's current task via short-lived signed credentials, with a semantic axis flagging in-scope-but-off-task calls; the deterministic axis carries the security guarantee
  • Team-Scoped Agent Policy Delegation — Mark selected agent-config keys overridable so enterprise teams vary them inside an admin-owned ceiling; the multi-team merge resolves least-restrictive, so team membership becomes a capability grant no settings diff shows
  • Transcript-Driven Permission Allowlist — Mine session transcripts for repeated read-only tool calls and propose a prioritized allowlist — narrower than bypass, tighter than manual curation

Code Injection

Code injection in multi-agent pipelines exploits agent trust in code it reads as input, distinct from prompt injection against a single agent.

Multi-Agent Propagation

Multi-agent systems with shared retrieval propagate adversarial content agent-to-agent. Defenses target the contagion channel and the per-agent detection signal.

PR-Time and Scheduled Review

Operational patterns that apply security agents to incoming changes and to resident codebase risk on different cadences.

  • Always-On Agentic PR Security Review — Pair a PR-time security reviewer with a scheduled whole-codebase scanner so new and resident risk both have continuous coverage; treat the reviewer agent itself as an injection target
  • Capability-Pegged Security Re-Scans — Re-scan unchanged code when the scanner improves rather than when the code changes; only worth it when the trigger is measured against your own defects and every finding clears a reproduction gate
  • Scanner-as-MCP-Server: Secret and Dependency Scans as Typed Agent Tools — Ship the security scanner as an MCP server so the agent invokes typed scans pre-commit and reasons over structured findings; qualified by five failure modes including agent-skips-scan and lethal-trifecta closure on the scanner principal
  • Cross-Repository Security Posture for Agent-Introduced Vulnerabilities — Treat one agent-introduced finding as a class to enumerate and remediate across every repository via variant analysis, not a per-repo point-in-time event; the organization-wide layer above per-agent scoping and per-PR review
  • Black-Box Probing in the Agent Build Loop: When It Pays — Probe the running deployment from outside to catch controls the diff never showed were missing; conditional on an ephemeral clone, an authenticated crawl, and never treating a clean result as proof the controls present are correct
  • Structural Monitoring for Covert Safeguard-Weakening — Diff the control- and data-flow graph of an agent's infrastructure-as-code change to catch a safeguard weakened under task success; an untrained structural monitor missed 3.5% of sabotage versus 11.6% for a git-diff baseline and cut joint sabotage success to 0% under synchronous rollback

Tool Invocation

Tool invocation exposes attack surfaces distinct from prompt injection. Malicious tools exploit argument generation and return processing to leak context and execute arbitrary commands.

Supply Chain

Agents dynamically load tools from MCP servers, plugins, and registries at runtime. A tampered tool inherits the agent's full permissions.

  • Agent-Emitted Dependency Version Ranges Widen the Supply-Chain Attack Surface — Agents default to caret and tilde ranges because npm install does; for an application with a bump-bot, replace the range with an exact pin plus a lockfile-enforced install — the floating range is the leg that admits a future-compromised release
  • Supply-Chain Security Debt in Agent Pull Requests — Across 4,022 agent-authored PRs, 82.3% of security smells are supply-chain integrity issues (unpinned CI/Docker references) concentrated in workflow and Docker files; target review at high-impact infra paths and enforce SHA pinning plus secret scanning in CI
  • Content-Addressed Agent Configurations (Deterministic Control Plane) — Treat coding-agent configs as an installed supply chain with SHA-256 content addressing, a per-project lockfile, and five declared permission tiers; the 10.1% cross-repo duplicate rate and the <1% configs declaring any permission scope (vs 33% of GitHub Actions workflows) make this a governance gap, not a hypothetical one
  • LLM-Pinned Library Versions Carry Systemic CVE Exposure — Across 10 models on 1,000 Python tasks, 36.7%-55.7% of LLM-specified versions contain known CVEs and all models converge on the same risky releases — pin against an external vulnerability source, not the model's training prior
  • Skill Composition Risk in Agent Ecosystems — Skills benign in isolation become harmful when one skill's output flows into the next; three failure modes (CapFlow, TrustLift, AuthBlur) reach 33.6%, 96.5%+, and 71.8% relative attack success across ten production backends, and per-skill vetting misses them by construction
  • Skill Supply-Chain Poisoning — Malicious skills injected into public registries exploit in-context learning to execute payloads hidden in documentation examples, bypassing alignment that blocks explicit instruction injection
  • Single-Decision Approval of Vendor Skill Suites — Approving a vendor's skill bundle in one decision hands the review unit to that vendor; cascades split across the suite hit 89.4% attack success, and concatenating the skills into one review query recovers 5.2 points at best
  • The Skill Closure Declaration Gap — Of 526 published Agent Skills that bundle files, 21 name every one of those paths in SKILL.md, 67 carry links resolving outside their own root, and none declares a dependency; approving the root approves an incomplete statement of what a run can reach
  • Setup Documentation as an Install-Time Attack Vector — Setup docs are unverified install authority; editing only a README, requirements file, or Makefile redirects an agent's install to a wrong name, untrusted registry, or vulnerable version, and the payload runs at install time — agents catch typosquats but miss source-based redirects almost everywhere
  • Slopsquatting: Hallucinated Package Names as a Supply-Chain Vector — Coding LLMs hallucinate package names at 5.2%-21.7%; 43% of those names persist across re-runs, making them enumerable — attackers register the persistent names and the agent's install step pulls the malicious package
  • Benign Skill Wording That Steers Package Hallucination (Neutral Prompting Attack) — A skill naming no package and giving no malicious instruction raised hallucination attack success from 4.54% to 78.99%; a strict allowlist measured 0.00% while a registry existence check measured no reduction at all
  • Tool Signing and Signature Verification — Require cryptographic signature verification (Sigstore/Cosign) before an agent loads or invokes a tool

Defense in Depth

No single safety mechanism is sufficient. Layered defenses ensure that failure of one layer does not compromise the agent.

  • Agent Governance Plane: Audit Events and Message-Content Surfaces — The four-part admin surface vendors now ship for coding agents, and the two conditions (deployment topology and retention) that decide whether its record is complete
  • Agent Retrieval Provenance as an Audit Control — Capture which files an agent searched and read before it changed code so an auditor can test the process behind the diff; scope it to high-blast-radius changes and persist hashes rather than prompt text
  • Cross-Iteration Safety State for Agent Loops (LoopHarness) — A monitor that resets each loop iteration is provably at chance against evidence split across iterations; a latched risk value fixes detection and trades it for an availability failure the defense itself creates
  • Cross-Layer Evidence for Agent Attack Detection — Pair kernel syscall traces with application telemetry when one session bounds both, you train your own detector, and post-hoc verdicts are enough; composition took the top AUROC for seven of 10 detectors on held-out attack families, and the per-mechanic table says which layer carries the signal
  • Cryptographic Governance Audit Trail — Wrap tool calls with policy validation and post-quantum receipt signing to produce a tamper-evident, append-only action log for regulated environments
  • Defense-in-Depth Agent Safety — Layer five independent safety mechanisms so no single failure point can compromise agent behavior
  • Distributing Security Controls Through the Agent Harness — A distributable harness ships OS sandboxing, skill scanning, and tool restriction to a fleet in one install, lifting composite score 52.3% to 78.3%; only declarative, boundary-local, conflict-free controls survive the trip, and injection defense is not one of them
  • Enforced Versus Advisory Controls in LLM-Native IDEs — A taxonomy from 446 developer reports puts 43.1% of security complaints in unauthorized file operations; sort safeguards by where they are evaluated, since ignore files and rules files are resolved inside the model's context while permissions and sandboxes are not
  • Enforcing Who and What Can Trigger an Agent's CI Run — GitHub Actions evaluates actor and event rules before a run starts, so the trigger rule stops being something the agent can edit; available on public repos and private ones on Team or Enterprise, with the public pull_request_target default enforced on 2 November 2026
  • Enterprise Agent Hardening: Governance, Observability, and Reproducibility — Move agents to production through three control gates — governance, observability, reproducibility — with MUST/SHOULD checklists for each
  • Inline Safety Harness with Cascade Verification (FinHarness) — Wrap each agent turn with prospective per-call monitors and route verification between a cheap and an advanced judge by per-step risk; worth it for high-stakes, high-call-volume workflows, not low-volume or long-context agents
  • Lifecycle-Integrated Security Architecture for Agent Harnesses — Embed defense mechanisms into each execution lifecycle phase with cross-layer feedback so layers coordinate rather than operate in isolation
  • Lock-State Safeguards for Desktop-Controlling Agents — Bound an agent driving a logged-in desktop along four axes (time, visibility, presence, recovery) with short-lived authorization, covered displays, relock on local input, and manual-unlock fallback so a failure on any single axis is contained by the others
  • Runtime Guard as an Installed Skill (Defense-as-Skill) — Holding the policy text identical, loading a runtime guard through skill files instead of an always-on system prompt cut in-distribution attack success from 0.353 to 0.104; it is still an advisory control that sits above permission systems and sandboxing, not in place of them
  • Security Constitution for AI Code Generation — Formalize security constraints as a versioned, machine-readable constitution that feeds agent specs, linters, and CI gates
  • Security Drift in Iterative LLM Code Refinement — Iterative fix-test loops optimize for functional correctness while silently accumulating security regressions that no functional test exercises
  • Three-Depth In-Session Security Review — Stack a per-edit pattern match, an end-of-turn diff review, and a commit-time agentic review so each layer's cost and false-positive profile match its frequency
  • Usability Pressure as a Silent Security-Regression Vector — Explicit usability requirements (performance, simplicity, new features) in a single-shot prompt cause LLMs to drop implicit security constraints at up to 98.1% attack success rate; mitigated by making security explicit and gating every output through a scanner
  • Verifying LLM-Generated Cryptographic Code — Crypto generation fails with 23.3% compile rate and 57% vulnerabilities; pair every crypto code path with a rule-based crypto analyzer, prefer zero-shot over CoT, and constrain to vetted high-level APIs

Economics

Sizing frames for pre-release security review when vulnerability discovery scales with inference spend.

  • Security Budget as Token Economics — Treat hardening as a budget-allocation decision: AISI's Mythos evaluation shows no diminishing returns inside 100M tokens per attempt, but the outspend frame applies only where the search curve is still climbing and triage capacity absorbs findings

Deployment Models

Release patterns for capabilities whose offense-defense asymmetry makes broad release the wrong default.