Tags¶
Browse the full Agent Patterns library by topic tag.
Tags are the primary topic-first index for this site. Each topic tag below gets a short intro and anchor pages; the full auto-generated tag listing follows beneath the intros. Tool-specific tags (claude, copilot, cursor, tool-agnostic) carry no prose — see the auto-list at the bottom.
Topic intros¶
context-engineering¶
Patterns for managing what enters the agent's context window: budgets, compression, attention curves, retrieval, and prompt caching. Read this if you're hitting context limits, debugging "lost in the middle" failures, or paying too much for token-heavy prompts. Start with:
- Context Engineering: The Discipline of Designing Agent Context
- Context Budget Allocation: Every Token Has a Cost
- Layered Context Architecture
- Lost in the Middle: The U-Shaped Attention Curve
- Static Content First to Maximize Cache Hits
agent-design¶
How a single agent is structured: harness, delegation, backpressure, composition, and memory. Read this when picking a harness, designing a long-running loop, or choosing between role-based and task-based delegation. Start with:
- Harness Engineering
- The Delegation Decision
- Agent Backpressure: Automated Feedback for Self-Correction
- Agent Composition Patterns: Chains, Fan-Out, Pipelines
- Agent Memory Patterns
multi-agent¶
Orchestration across multiple agents: fan-out, hand-offs, file-based coordination, and sub-agent topology. Read this when one agent isn't enough and the question shifts from "how do I prompt it" to "how do I wire them together". Start with:
- Orchestrator-Worker Pattern
- Fan-Out Synthesis Pattern
- Sub-Agents for Fan-Out Research
- Agent Handoff Protocols
- File-Based Agent Coordination
memory¶
What an agent remembers across turns, sessions, and projects — versioned vs streaming, episodic vs reinforced, and the failure modes when memory leaks into wrong contexts. Read this when designing persistence or chasing memory-induced drift bugs. Start with:
- Agent Memory Patterns
- Tiered Memory Architecture
- Episodic Memory Retrieval
- Memory-Induced Tool Drift
code-review¶
Agentic code review patterns: how agents read diffs, when committee or tiered reviews pay off, and how to wire review verdicts back into iteration. Read this when integrating agents into PRs or designing a review-then-implement loop. Start with:
- Agent-Assisted Code Review
- Diff-Based Review Over Output Review
- Committee Review Pattern
- Tiered Code Review
- Agentic Code Review Architecture
instructions¶
System prompts, CLAUDE.md / AGENTS.md files, polarity rules, and the compliance ceiling that limits how many rules an agent will honour. Read this when writing or pruning instruction files, deciding what's primacy-critical, or debugging non-compliance. Start with:
- System Prompt Altitude
- CLAUDE.md Convention
- AGENTS.md as Table of Contents, Not Encyclopedia
- The Instruction Compliance Ceiling
- Instruction Polarity: Positive Rules Over Negative
workflows¶
End-to-end agent-assisted development loops: planning, eval-driven iteration, CI integration, onboarding, and team collaboration. Read this when fitting agents into your existing process rather than starting from a blank page. Start with:
- The Plan-First Loop: Design Before Code
- Eval-Driven Development
- Headless Claude in CI
- Repository Bootstrap Checklist
- Continuous Agent Improvement
cost-performance¶
Routing, token efficiency, model selection, and tool-call cost control. Read this when token spend is hurting, you're tuning a reasoning/execution model split, or designing tool descriptions to minimize per-call overhead. Start with:
- Cost-Aware Agent Design
- Token-Efficient Tool Design
- Reasoning Budget Allocation
- Stateless Agent Loop with Prompt Caching
- Gateway Model Routing
testing-verification¶
General testing and verification discipline for agent-generated work: TDD, incremental checks, verification ledgers, and pre-completion gates. Read this for the broad "did the agent do the right thing" patterns — see code-review and evals for the narrower sub-domains. Start with:
- Test-Driven Agent Development
- Incremental Verification
- Verification Ledger
- Pre-Completion Checklists
- Risk-Based Task Sizing for Verification Depth
evals¶
Evaluation design specifically: pass@k, LLM-as-judge, golden query pairs, and turning incidents into regression cases. Read this when measuring agent quality with structure, building regression suites, or scoring rollouts before they ship. Start with:
- pass@k and pass^k Metrics
- LLM-as-Judge Evaluation
- Eval-Driven Development
- Golden Query Pairs as Regression Tests
- Incident-to-Eval Synthesis
security¶
Prompt-injection defense, blast-radius containment, secrets handling, sandboxing, and the lethal-trifecta posture. Read this before granting agents write capability, ingesting untrusted content, or wiring an MCP server. Start with:
- Prompt Injection: A First-Class Threat
- Blast Radius Containment: Least Privilege for AI Agents
- Dual-Boundary Sandboxing
- URL-Based Data Exfiltration Guard
- Defense-in-Depth Agent Safety
observability¶
Tracing, loop detection, transcript analysis, OTel for agents, and harness debugging. Read this when agents fail silently, loops run wild, or you need to diagnose what an agent actually did versus what it claims. Start with:
- Agent Debugging
- Loop Detection
- Circuit Breakers for Agent Loops
- OpenTelemetry for Agent Observability
- Analyzing Agent Evaluation Transcripts
human-factors¶
The human side of agent-driven work: cognitive load, attention management across parallel sessions, skill atrophy, adoption, and team dynamics. Read this when scaling agents to a team, fighting AI fatigue, or designing supervision rituals. Start with:
- Developer as CPU Scheduler: Attention Management
- Cognitive Load, AI Fatigue, and Sustainable Agent Use
- Cross-Tool Translation
- Skill Atrophy
- Agentic Education Persona Progression
github-actions¶
GitHub Actions-specific integration for agent workflows: CI-driven triage, closed-loop remediation, headless runs, and one-click auto-fix. Read this when wiring agents into Actions or designing the CI-side automation that triggers and consumes agent output. Start with:
- Headless Claude in CI
- Closed-Loop CI Failure Remediation
- One-Click CI Auto-Fix
- Continuous Triage
- AI Bot CI Workflow Reliability
skills¶
SKILL.md and skill-design patterns: how a named, versioned skill is authored, scoped, and invoked. Spans agent-design and tool-engineering. Read this when designing a CLI-first skill, packaging behavior as an importable module, or gating skill invocation. Start with:
- CLI-First Skill Design
- Google ADK Skills: Portable SKILL.md Across ADK Agents
- Skill as Instruction Surface and Callable API (Interpreter Skills)
- On-Demand Skill Hooks: Session-Scoped Guardrails via Skill Invocation
- Handoff Skill: Structured Context Transfer Between Agent Sessions
automation¶
Non-interactive, headless agent runs and CI/CD integrations with no human in the loop — scheduled routines, programmatic dispatch, and background pipelines. Read this when wiring agents into CI, cron, or webhook triggers rather than an interactive session. Distinct from workflows, which covers human-driven development loops. Start with:
- Headless Claude in CI: Using -p and --max-turns for Safe Pipeline Integration
- Programmatic Cloud-Agent Dispatch via REST API and Webhooks
- Issue-to-PR Delegation Pipeline for AI Agent Development
- Multi-Repo and No-Repo Coding Agent Automation Templates
- Cloud-Scheduled Routines vs Local Session Scheduling
All tags¶
agent-design¶
- A Governance Framework for Production Agents
- A2UI: Framework-Agnostic Generative UI Standard for Agents
- ACID for Agent Repository State
- AI Bot CI/CD Workflow Reliability by Agent
- AI Slop as a Process Problem: Encoding Quality Standards as Pipeline Gates
- AI-Powered Vulnerability Triage for AI Agent Development
- AOCI: Symbolic-Semantic Repository Indexing
- AST-Grounded Critic Loop for Documentation Maintenance
- AX/UX/DX Triad: Three Experience Layers in Agent Systems
- Abstraction Bloat in AI Agent-Generated Code Output
- Accumulated Behavioral Rules from Review Feedback
- Action-Audit Divergence: A Four-Mode Taxonomy for Runtime Hardening
- Action-Class Decomposition for Tool-Calling Evals
- Action-Selector Pattern: LLM as Intent Decoder with Deterministic Execution
- Adaptive Generate-Rank-Verify Under Costly Verification
- Adaptive Sandbox Fan-Out Controller
- Administrative Effort Ceilings for Reasoning Budget
- Advanced Tool Use: Scaling Agent Tool Libraries
- Adversarial Multi-Model Development Pipeline (VSDD)
- Adversarial-Only Threat Modeling for Agent Data Leakage
- Agent Backpressure: Automated Feedback for Self-Correction
- Agent Cards: Capability Discovery Standard for AI Agents
- Agent Circuit Breaker
- Agent Commit Attribution: Signed Commits and Agent Identity
- Agent Composition Patterns for Multi-Agent Workflows
- Agent Debug Log Panel: Chronological Event Inspection for Session Debugging
- Agent Debugging: Diagnosing Bad Agent Output
- Agent Definition Formats: How Tools Define Agent Behavior
- Agent Design Patterns and Architectures for AI Agents
- Agent Determinism Moves to the Last Unconstrained Axis
- Agent Development Lifecycle for Agent Products
- Agent Environment Bootstrapping for AI Agent Development
- Agent Event Streaming: Consumer Contract Above the Tokens
- Agent Extension Conflicts: When Installed Skills and MCP Servers Fight Each Other
- Agent Failure Trajectories and the Recovery Window
- Agent Governance Plane: Audit Events and Message-Content Surfaces
- Agent Governance Policies for AI Agent Development
- Agent HQ (Multi-Agent Platform) for AI Agent Development
- Agent Handoff Protocols: Passing Work Between Agents
- Agent Harness: Initializer and Coding Agent Pattern
- Agent JIT Compilation: Compile Tasks Into Executable Plans
- Agent Memory Patterns: Learning Across Conversations
- Agent Mission Control for Orchestrating Agent Tasks
- Agent Network Egress Policy: Admin-Controlled Domain Allow/Deny
- Agent Observability with OpenTelemetry and Trajectory Logging
- Agent Project State Purge: Clean-Slate Session Reset
- Agent Pushback Protocol for Managing Disagreements
- Agent Retrieval Provenance as an Audit Control
- Agent Runtime Middleware: Per-Call Interception Pipeline
- Agent Skills: A Cross-Tool Task Knowledge Standard
- Agent Sprawl: Unmanaged Sub-Agent and Skill Proliferation
- Agent Terminology Disambiguation for AI Coding Systems
- Agent View: Dispatch-Attach-Monitor Surface for Parallel Sessions
- Agent as Tool vs Handoff: Who Keeps the Conversation
- Agent-Authored Messages as a Deferred Exfiltration Channel
- Agent-Authored PR Integration: Collaboration Signals That Determine Merge Success
- Agent-Aware CLI Behavior via Environment Variable
- Agent-Client Admission Control for Agentic Traffic
- Agent-Computer Interface (ACI): Tool Design as UX Discipline
- Agent-Discoverable Slash Commands
- Agent-Driven Deployment: What to Delegate and What to Gate
- Agent-Driven Fuzzing with Human-Gated Crash Triage
- Agent-Driven Greenfield Product Development from Scratch
- Agent-First Software Design for AI Agent Development
- Agent-Generated Onboarding Guide as a Durable Artifact
- Agent-Led Dev-Environment Iteration with Validation and Rollback
- Agent-Native Filesystems: Gating Effects, Not Commands
- Agent-Operable Interface Design (Affora)
- Agent-Powered Codebase Q&A and Onboarding Workflow
- Agent-Reactive Bugs at the Model-Harness Boundary
- Agent-Ready Data Architecture for Analytics Agents
- Agent-Trace Data Layer: Storage for Hours-Long Traces
- Agent-to-Agent (A2A) Protocol for AI Agent Development
- Agentic AI Architecture: From Prompt to Goal-Directed
- Agentic Detection and Response at the MCP Boundary
- Agentic Flywheel: Self-Improving Agent Systems
- Agentic Pattern Vocabulary Crosswalk
- Agentic Resource Discovery: Federated Pre-Invocation Search
- Agentic-Agile: Adapting Agile Rituals for Agent Work
- Agentless vs Autonomous: When Simple Beats Complex
- Agents vs Commands: Separation of Role and Workflow
- Aggregation Bounds for Agent Authorization
- Always-On Agentic PR Security Review
- Ambition Scaling: Moving the Target as Model Capability Increases
- An Explicit Update Boundary for Agent Self-State
- Anthropic's Effective Agents Framework: A Pattern Map
- Anti-Reward-Hacking: Rubrics That Resist Gaming
- Applying Coding Agents to Non-Code Tasks
- Approval Gate Granularity in Agent Pipelines
- Architecting a Central Repo for Shared Agent Standards
- Artifact-Driven Workflow Compilation for Agent Execution
- Ask-Everything Permission Policies Protect Less than Per-Action Approval
- Assumption Propagation: Compounding Agent Misunderstandings
- Async Non-Blocking Subagent Dispatch
- Asynchronous Agent I/O and Speculative Tool Calling
- Attention Latch: When Agents Stay Anchored to Stale Instructions
- Auth-Isolation as the MCP-vs-CLI Selection Heuristic
- Authority Confusion: Untrusted Context Must Not Authorize Side Effects
- Authorization Continuity Across Agent Mutation
- Auto Model Selection: Harness-Driven Routing per Task
- Auto-Merging a Wiki Agent's Documentation Pull Requests
- Auto-Triage Workflow: Bug-Monitoring Agent that Connects Related Reports and Opens Fix PRs
- Autonomous Research Loops: Loops That Know When to Stop
- Background Todo Agent: Offload Plan Maintenance to a Lightweight Model
- Backlog Triage as a Named Agent Skill
- Behavior Specs: Grading the Trajectory, Not the Result
- Behavioral Drivers of Coding Agent Success and Failure
- Behavioral Firewall for Tool-Call Trajectories
- Behavioral Specification Elicitation Before Synthesis (SpecFirst)
- Behavioral Testing for Non-Deterministic AI Agents
- Belief Inertia After Tool-Map Drift in AI Agents
- Benchmark Poisoning of Self-Modifying Coding Agents
- Binding an Agent's Effect to the Approval It Claims
- Blaming the Model for Scaffolding-Driven Quality Regressions
- Blast Radius Containment: Least Privilege for AI Agents
- Bootstrapping Coding Agents: The Specification Is the Program
- Boring Technology Bias: When Agents Recommend by Popularity
- Bounded Agent Steps Inside a Deterministic Workflow
- Bounded Batch Dispatch for Parallel Agent Execution
- Bounding a Headless Codex Run Without a Turn Cap
- Bounding an Embedded Copilot SDK Agent's Tool Set
- Brownfield to Agent-First: Repo Maturity Framework
- Browser Automation as a Research Tool: Bypassing Bot Detection
- Browser as Agent Action Space
- Budgeted Verification of Inherited Agent Constraints
- Building Custom Agents from Substrate to Production (Agents All the Way Down)
- Burn the Boats — Commitment-Forcing Deprecation
- CARE: Three-Party Stage-Gated Engineering of LLM Agents
- CLI-First Skill Design
- CLI-IDE-GitHub Context Ladder for AI Agent Development
- Cache-Safe Routing Boundaries: Where a Router May Act
- Calibrated Deciders for In-Loop Agent Decisions
- Caller-Actionable Error Steps in Tool Responses
- Canary Rollout for Agent Policy Changes
- Canvas as Control Surface: Steering a Long-Running Agent Mid-Run
- Canvas as Durable Workflow State: The Four-Step Blueprint and What It Costs
- Capability Declarations for Agents That Act on Data
- Capability-Additive Code Interpreters for Untrusted Agent Code
- CausalFlow: Counterfactual Repair for Failed Agent Trajectories
- Chance-Corrected Shortlist Depth Sizing for Tool Retrieval (Bits-over-Random)
- Channels Permission Relay
- Chat-Platform Agent Delegation: Invoking Cloud Coding Agents from Team Channels
- Choosing an Agent Tool Interface: Shell or Typed Catalog
- Choosing an Integration Layer for an Embedded Agent Harness
- Choosing the Right Surface for a Coding Agent Task
- Circuit Breakers for Agent Loops
- Claim-to-Evidence Trace Graphs for Auditing Agent Runs
- Clarification Mode Amplifies Prompt Injection
- Classical SE Patterns as Agent Design Analogues
- Classifier-Gated Auto-Permission for Cloud-IDE Coding Agents
- Classifier-Subagent Run Mode for Per-Call Permission Routing
- Classifying and Auto-Correcting Coding Agent Misbehaviors (Wink)
- Claude Agent SDK: Building Custom Agentic Workflows
- Claude Code /batch and Worktrees for AI Agent Development
- Claude Code Agent Teams for Collaborative AI Workflows
- Claude Code Auto Mode: Classifier-Based Permission Gating
- Claude Code Dynamic Workflows
- Claude Code Extension Points: When to Use What
- Claude Code Feature Flags and Environment Variables
- Claude Code Hooks: Deterministic Lifecycle Automation
- Claude Code Sub-Agents for Delegating Complex Tasks
- Clock-In / Clock-Out Protocol: Bracketed Session Continuity
- Close the Attack-to-Fix Loop: Adversarially Train Agent
- Closed-Loop Agent Training from Tool Schemas
- Closed-Loop CI Failure Remediation with Cloud Coding Agents
- Closed-Loop Role-Based Refinement for Agent Systems
- Closed-World Tool Call Resolution Before the Permission Gate
- Cloud Planning with Inline-Comment Review and Execute-Anywhere Choice
- Cloud-Agent Session Bootstrap: Cached Install plus Per-Session Start
- Cloud-Agent Three-Layer State Decoupling
- Cloud-Local Agent Handoff for AI Agent Development
- CoALA Decision-Making Loop as an Orchestration Lens
- CoALA Memory Taxonomy as a Classifier for Harness Artifacts
- CoALA Structured Action Space: Internal vs External Actions
- CoT Robustness in Code Generation
- Code Injection Defense in Multi-Agent Pipelines
- Code Interpreter as a Primary Agent Tool
- Code-Native Memory Substrates for Coding Agents
- Codebase Readiness for Agents: Agent-Friendly Code
- Coding Agent Scope Expansion: When to Extend Beyond the Codebase
- Coding-Agent Misalignment Forms (Seven-Symptom Taxonomy)
- Coding-Agent Working-Set Coverage (Coherence Debt)
- Cognitive Architectures for Language Agents (CoALA): A Classifier for Agent Harnesses
- Cognitive Poisoning: Untrusted Tool Feedback as a Trajectory Attack
- Cognitive Reasoning vs Execution: A Two-Layer Agent
- Cohesion-Aware Task Partitioning for Multi-Agent Coding
- Comment-Triggered Agent Dispatch on Issues and PRs
- Committee Review Pattern for Multi-Agent Code Review
- Comparison-Only Advisor: Steering a Large Actor With a Tiny Comparator
- Compiled Specialist Agents: Muscle Memory for Recurring Intent
- Compositional Skill Routing for Large Skill Libraries
- Compositional Vulnerability Induction in Coding Agents
- Compound Engineering: Learning Loops That Make Each Feature Easier
- Computer-Systems Lens for Always-On Agent Security
- Concurrent Agent Pull Requests and Merge-Conflict Cost
- Conditional Hook Execution: Filter Hooks by Tool Pattern
- Confirmation Gates for Consequential Agent Actions
- Consistent-format customer capture
- Consolidate Agent Tools to Reduce Cognitive Overhead
- Constraint Decay in Backend Code Generation
- Constraint Drift: Why Safety Must Be Maintained, Not Asserted
- Constraint Tax: Tool Suppression Under JSON Schema Decoding
- Constraint-Evasive Fabrication in Instruction Sets
- Content Exclusion Gap: AI Security Boundaries by Mode
- Context Compiler: Deterministic Assembly Over Bigger Windows
- Context Poisoning: When Hallucinations Become Premises
- Context-Fractured Decomposition Attacks on Tool-Using Agents
- Context-Injected Error Recovery for AI Agent Development
- Contextual Capability Calibration for Multi-Agent Delegation
- Continual Learning for AI Agents: Three Layers of Knowledge Accumulation
- Continuation Authority in Agent Migration
- Continuation Dispatcher: Who Owns the Iterate-or-Stop Call
- Continuous AI (Agentic CI/CD) for AI Agent Development
- Continuous AI: A Navigation Map of Always-On Agent Workflows
- Continuous Agent Improvement: Iterating on Agent Quality
- Continuous Autonomous Task Loop
- Continuous Documentation as an Agent-Driven Practice
- Continuous Triage: Automating Issue Classification with AI Workflows
- Continuously Built Agent Environments
- Control/Data-Flow Separation for Prompt Injection Defense (CaMeL)
- Controlled Benchmark Rewriting for Agent Safety Judgment
- Convention Over Configuration for Agent Workflows
- Coordination Channel Policy for Multi-Agent Coding
- Copilot Auto Tiers: Weighting Cost Against Quality
- Copilot CLI Agentic Workflows for AI Agent Development
- Copilot Cloud Agent Organization Controls
- Copilot Cloud Agent Three-Phase Execution Model
- Copilot Memory and Cross-Agent Persistence
- Copilot Unified Sessions View and CLI Agent in JetBrains IDEs
- Cost-Aware Tracing for Skill Distillation
- Cost-Driven Model Routing Without Quality Monitoring
- Coverage-Guided Agents for Fuzz Harness Generation
- Covert Success Rate for Indirect Prompt Injection
- Credential Hygiene for Agent Skill Authorship
- Critic Agent Pattern: Dual-Model Plan Review
- Cross-Component Interference in Agent Scaffolds
- Cross-Cycle Consensus Relay
- Cross-Framework Signal Semantics: Re-Measure Borrowed Trajectory Rules
- Cross-Iteration Safety State for Agent Loops (LoopHarness)
- Cross-Layer Evidence for Agent Attack Detection
- Cross-OS Library MicroVM Sandbox for Agent Code (smolvm)
- Cross-Session Peer Messaging with a Posture-Keyed Inbox Gate
- Cross-Session Re-Implementation of Existing Agent Code
- Cross-Tool Subagent Comparison
- Cross-Vendor Competitive Routing for LLM Selection
- Cryptographic Governance Audit Trail for AI Agents
- Cue-Anchored Working Memory (Delivery, Not Storage)
- Cursor /multitask: Async Subagent Dispatch in the Editor
- Cursor 3 Agents Window: Parallel Agents and Worktree Isolation
- Cursor Automations: Event-Triggered Agents and /automate
- Cursor SDK: Programmable TypeScript Agent Runtime
- Cursor Self-Hosted Cloud Agents
- Cursor for AI Agent Development
- Customer-Hosted MCP Tunnel: Outbound-Only Connectivity to Private MCP Servers
- DSLs as a Constraining Harness for LLM Code Generation
- DSPy: Programmatic Prompt Optimization for Compound Agent Systems
- Daily-Use Skill Library: Encoding Your Process as Agent Skills
- Data Fidelity Guardrails: Preventing Agent Data Mutation
- Debugging the Tool-Call Loop Before Reaching for a Framework
- Decision Unbundling: What Moves Into the Orchestrator
- Decision-Fork Replay: Grading an Agent's Mid-Run Choices
- Declarative Multi-Agent Composition
- Declarative Multi-Agent Topology: Topology-as-Code
- Declared Peer Consensus as Context for a Reviewing Agent
- Decomposed Red-Teaming for Agent Monitors
- Decomposing Agent Output Variability by Layer (Sampling vs Orchestration State)
- Decoupled Search Grounding: A Vendor-Agnostic Grounding Boundary
- Deep Agent Runtime: The Layer Beneath the Harness
- Defense-in-Depth Against Coding Agent Fabrication (Honesty Harness)
- Defense-in-Depth Agent Safety for AI Agent Development
- Deferred Permission Pattern: Headless Agent Session Pausing
- Delegated-Autonomy Boundary Artifacts (AJR and ADP)
- Delegating Delivery Stages to GitHub Agent Apps
- Delegating Dependabot Pull Request Triage to an Agent
- Delegating Multi-Hunk Bug Repair to Coding Agents
- Delegation Threshold Calibration for Orchestrator Agents
- Deletion and Cost Rules for an Evolving Agent Harness
- Delivery-Bound Tool Authorization: When Progressive Discovery Becomes Access Control
- Delta Channels: Bounded Checkpoint Storage for Append-Only Agent State
- Demand-Driven Repository Auditing
- Demo-to-Production Gap: When Demos Hide Real Costs
- Deny-Fallback Permissions for Unattended Agent Runs
- Design Docs as the Durable Artifact
- Designing Agent Plugins to Survive Co-Installation
- Designing Agent Tools Like APIs
- Designing Agents to Resist Prompt Injection
- Destructive-Failure Mechanism Attribution by Mitigation Owner (ClayBuddy Three)
- Deterministic Fast Paths: Answer Without a Model Call
- Deterministic Orchestration for Structured Modernization
- Deterministic Precondition Gates for Tool-Using Agents
- Dev Containers for AI Coding Agents: Claude Code vs Copilot CLI
- Developer as CPU Scheduler: Attention Management with Parallel Agents
- Difficulty-Aware Topology Selection for Coding Agents
- Direct Prompt Injection via Collaboration (User as Attack Vector)
- Discovering Indirect Injection Vulnerabilities in Your Agent
- Discovery-Only Refactor Pass: Surface Candidates Before Touching Code
- Discrete Phase Separation
- Dispatch-Time Reasoning Level for Delegated Agents
- Distillation-Induced Similarity Metrics for Tool-Use Agents
- Distilled Bootstrap Contract: Agent-Authored Repo Setup
- Distributed Computing Parallels in Agent Architecture
- Distributed Cross-PR Attacks in Persistent-State AI Control
- Distributing Security Controls Through the Agent Harness
- Docker sbx Adoption for Coding Agents
- Document-Borne Prompt Injection Through Agent Read Tools
- Documentation-Guided Legacy Migration: Architecture Docs as a C-to-Rust Blueprint
- Domain-Scoped Parallel Exploration for Multi-File Change Localization
- Domain-Specific Agent Challenges
- Domain-Specific System Prompts with Concrete Examples
- Dominator-Graph Trajectory Invariants for Non-Deterministic Agents
- Dormant Memory Payloads Triggered by Sensitive Topics (Trojan Hippo)
- Downstream Disclosure Coordination for Agent-Found Defects
- Dual-Boundary Sandboxing: Filesystem and Network Isolation
- Dual-Budget Control for Search Agents: VOI Scoring Per Action
- Dual-Graph Alignment for Indirect Prompt Injection Defense (AuthGraph)
- Dual-Trace Memory Encoding: Pair Facts with the Scene They Were Learned In
- Dual-Write Append-Mirror for Agent Transcript Externalization
- Durable Interactive Artifacts: Agent Output Outside the Transcript
- Earned-Complexity Agent Maturity Ladder
- Economic Value Signaling in Multi-Agent Networks
- Editor and Manager Surface Separation in Agent IDEs
- Effective Feedback Compute (EFC) for Harness Comparison
- Elastic Context Orchestration: A Per-Turn Vocabulary for Long-Horizon Search Agents
- Embedding Inversion: Vector Stores as a Source-Text Disclosure Surface
- Embedding the Copilot SDK in a Managed Java Runtime
- Emergent Architecture in AI-Driven Codebases
- Emergent Behavior Sensitivity for AI Agent Development
- Emerging Concepts for AI Agent Development
- Empirical Baseline: Agentic AI Coding Tool Configuration
- Encoding Tacit Knowledge into Agent Improvement Loops
- Enforced Versus Advisory Controls in LLM-Native IDEs
- Enforcement Modes in Spec-First Agent Frameworks
- Enforcing Agent Behavior with Hooks
- Enforcing Who and What Can Trigger an Agent's CI Run
- Enterprise Agent Hardening: Three Production Gates
- Enterprise Skill Marketplace: Distribution and Quality
- Entity Binding Failures in Tool-Augmented Agents
- Entropy Reduction Agents: Automated Codebase Hygiene
- Environment Specification as Context: Closing the Version Gap
- Episodic Memory Retrieval for AI Coding Agent Loops
- Epistemic Working Memory for Multi-Hop Reasoning (SLEUTH)
- Error Preservation in Context for AI Agent Development
- Escalation Channels: A Reporting Tool Instead of a Reward Hack
- Escape Hatches: Unsticking Stuck Agents
- Eval Blind Spots: Structural Gaps in Measurement Methodology
- Eval Strategy by Agent Generation: A Structure-to-Eval Locator
- Eval-Driven Development: Write Evals Before Building Agent
- Evaluator-Optimizer Pattern for AI Agent Development
- Event Sourcing for Agents: Separating Cognitive Intention
- Event-Driven Agent Routing for Multi-Team AI Pipelines
- Event-Loop Contention in Async Agent Fan-Out
- Evidence-Conditioned Execution: Gate Edits on Observations
- Evidence-First Reports From Failure-Diagnosis Agents
- Evidence-Grounded Disagreement in Agentic Code Review (Adversarial Review)
- Evolving Playbooks: Incremental Context That Preserves Knowledge
- Exactly-Once Enforcement Layer: Model, Harness or Contract
- Exception Handling and Recovery Patterns for AI Coding Agents
- Executable Memory: User State as Code for Personalized Agents
- Execution Lineage: DAG of Artifacts vs Agent Loops
- Execution-First Delegation: The AI-as-Executor Pattern
- Execution-Layer Security Invariants for MCP Runtimes
- Execution-State Ledger for Long-Horizon Coding Agents
- Experience Graphs as Structured Memory for Self-Evolving Agents
- Experiential-Learning Setup Agents with Snapshot Rollback (SetupX)
- Explanation-Bound Tool Execution for Agent Gateways
- External Artifacts Treated as Data, Not Adversarial Input
- Externalization in LLM Agents
- Fact Supersession Memory for Code Assistants
- Factory Over Assistant: Orchestrating Parallel Agent Fleets
- Fail-Closed Remote Settings Enforcement for Enterprise Agents
- Failure-Driven Iteration for Improving Agent Workflows
- False-Pass Liability Decides the Next Harness Component
- Fan-Out Synthesis Pattern for AI Agent Development
- Feature List Files for Reliable AI Agent Development
- Feedback as Capability Equalizer: Iterative Feedback Outweighs Model Scale
- Field-Change Intent Instead of Model-Written Diffs
- Field-Level Source Ownership for Agent Capabilities
- File-Based Agent Coordination for AI Agent Development
- Filesystem-Based Tool Discovery for AI Agent Development
- Filter and Aggregate Data in the Execution Environment
- First-Party Agent Composition: Agent-Built Features
- First-Proposal Execution in Agent Loops
- Five Design Decisions for MCP Servers and Clients
- Five-Failure-Layers Diagnostic: Attribute Before Swapping the Model
- Five-Stage Policy Layer Typology for Generalist Agents
- Flattened Tool Specs for Agent Safety Judgment (SafeKeep)
- Fleet Harness Attribution: Pinning the Model to Compare Whole Harnesses
- Fleet-Level Irreversibility Budgets for Agent Effects
- Foresight-Guided Defense Against Infectious Jailbreaks in Multi-Agent Systems
- Forged Reasoning Trace Attacks on Agent Memory (FARMA)
- Forked vs Fresh Subagents: When to Inherit the Parent Conversation
- Formal Process Models as Prompting Scaffolds (Petri Net of Thoughts)
- Four-Layer Taxonomy of Agent Security Risks
- Four-Phase Agent Delegation with Curated Artifacts
- Framework-First Agent Development: An AI Anti-Pattern
- Frameworks
- Framing Subagent Returns So They Cannot Act as Instructions
- Frozen Task Sets for Affordable Agent A/B Testing
- Functional folder taxonomy
- Gateway Hint Headers for Routing and Budgeting Agent Calls
- Gateway Model Routing: Treat the LLM Gateway as a Discovery Source
- Generated Programs as Web Agent Action Space
- Generative Agents Memory Stream: Three-Layer Architecture for Long-Running Agent Sessions
- Generative Provenance Records for Tool-Using Agents
- Git-Bound Memory for the Agentic Development Lifecycle
- GitHub Agentic Workflows for Automating Dev Processes
- GitHub Copilot Agent Mode for AI Agent Development
- GitHub Copilot Coding Agent for AI Agent Development
- GitHub Copilot Custom Agents and Skills Extensibility Guide
- GitHub Copilot Dedicated App as Agent-First Surface
- GitHub Copilot Extensions for AI Agent Development
- GitHub Copilot MCP Integration for AI Agent Development
- GitHub Copilot SDK for AI Agent Development
- GitHub Copilot for AI Agent Development
- GitHub Copilot: Customization Primitives and Stack
- GitHub Copilot: Harness Engineering for Agent-Ready Code
- GitHub Models in Actions for AI-Driven CI Workflows
- Goal Contract: Separating the Doer from the Done-Checker
- Goal Monitoring and Progress Tracking for Long-Running Agents
- Goal Reframing: The Primary Exploitation Trigger for LLM Agents
- Golden Journeys: Restartability as a First-Class Verification Primitive
- Golden Query Pairs as Continuous Regression Tests for Agents
- Governance Layer for Agent Interoperability Protocols
- Governed Sources of Truth for Analytics Agents (Structure Over Access)
- Graceful Tool-Output Truncation: The PARTIAL Signal
- Grade Agent Outcomes, Not Execution Paths
- Graph of Thoughts: Directed Graph Reasoning for Multi-Path Problems
- Grill Me: Developer-Initiated Plan Interrogation
- Guarding Against URL-Based Data Exfiltration in Agentic Workflows
- Handoff Skill: Structured Context Transfer Between Agent Sessions
- Happy Path Bias: How AI Agents Skip Error Handling
- Harness Bug Detection Patterns
- Harness Design Dimensions and Archetypes
- Harness Engineering (Training Module)
- Harness Engineering for Building Reliable AI Agents
- Harness Hill-Climbing: Eval-Driven Iterative Improvement of Agent Harnesses
- Harness Impermanence: Build Scaffolding To Be Deleted
- Harness Preflight Doctor Command for Agent Diagnostics
- Harness-Memory Coupling as a Design Axis
- Headless Claude in CI: Using -p and --max-turns for Safe Pipeline Integration
- Headless-First Services: APIs for Agent Consumers
- Heartbeat-Bound Hierarchical Credentials for Agent Swarms
- Heuristic-Based Effort Scaling in Agent System Prompts
- History Anchors: Consistency-Cued Continuation of Unsafe Prior Actions
- Hook Catalog for Claude Code Enforcement
- Hooks and Lifecycle Events: Intercepting Agent Behavior
- How Teams Build SE Agents: A Seven-Stage Build Loop
- How the Four Agent Engineering Disciplines Compound
- Human-in-the-Loop Checkpoints as Loop Control
- Human-in-the-Loop Placement: Where and How to Supervise
- Humans and Agents in Software Engineering Loops
- Hybrid Deterministic + Semantic Authorization for Agent Tool Calls
- Hypothesis-Driven Debugging: Instrument Before You Patch
- Idempotent Agent Operations: Safe to Retry
- Idle-Time Speculative Planning for ReAct Agents
- Improper Output Handling: Validate Agent Output Before Downstream Use
- In-Agent Task Prioritization: Ranking the Next Action
- In-Place Atomic Replacement for Agent-Driven Ports
- In-Thread Side-Channel: Bounded Side Questions Without Losing the Main Task
- Incident Log Investigation Skill: Parallel Queries
- Independent Test Generation in Multi-Agent Code Systems
- Indiscriminate Structured Reasoning on Every Agent Task
- Inference-Time Tool-Call Reviewer: Pre-Execution Feedback for Tool-Calling Agents
- Informed Abstention as a Tool-Boundary Runtime Gate
- Inline Safety Harness with Cascade Verification (FinHarness)
- Install-Once Plugin Trust: Vetting That Never Re-Runs on Update
- Intent-Governed Tool Authorization for AI Agents (IGAC)
- Interactive Canvases: Agent-Generated Visual Artifacts as Outputs
- Interactive Clarification for Underspecified Tasks
- Interactive Effort Sliders: Per-Turn Reasoning-Budget Controls
- Internal Hostname Disclosure in Agent-Readable Context
- Introspective Skill Generation: Mining Agent Patterns
- Inversion Analysis: Surface Capabilities Competitors Cannot Replicate
- Isometric Harness Ablation: Rank Subsystem Investment by Removing One at a Time
- Issue Requirements Preprocessing: Structured Input Before Code Generation
- Issue-Tracker as Agent Dispatch Surface
- Issue-to-PR Delegation Pipeline for AI Agent Development
- Judging Agent Safety by Task Completion (Action-Boundary Violations)
- Judgment Relocation: Where Human Decisions Land in an Agent Factory
- Kaizen-Style Continuous Code Quality Loop (Pomona)
- Knowledge Graphs as Provenance-Carrying Agent Memory
- Knowledge-Based Pull Requests for Cross-Trust-Boundary Contributions
- L0 → L1: Making the Repo Readable
- L1 → L2: Adding Feedback Loops
- L2 → L3: Building Mechanical Enforcement
- L3 → L5: Reaching Agent-First
- LLM API Fault Injection at the HTTP Layer (AgentChaos)
- LLM API Routers as Application-Layer Man-in-the-Middle
- LLM Agent Bug Fix Taxonomy: 23 Fix Patterns from 930 Real Bugs
- LLM Code Review Overcorrection for AI Agent Development
- LLM Map-Reduce Pattern for Parallel Input Processing
- LLM-as-Code Agentic Programming for Agent Harnesses
- LLM-as-Judge Evaluation with Human Spot-Checking
- Labels as Locks: Pipelined Backlog Processing with Stage Gates
- Lane-Based Execution Queueing
- Large-Codebase Coding-Agent Failure Patterns (Sourcegraph Five)
- Late Requirement Arrival in Agent Sessions
- Lay the Architectural Foundation by Hand Before Delegating
- Layer Agent Instructions by Specificity Across Scopes
- Layered Context Architecture for AI Agent Development
- Layered Domain Architecture: A Prescriptive Default for Agent-Built Code
- Layered Mutability: Governing Persistent Self-Modifying Agents
- Lazy Worktree Isolation: Enter the Worktree on First Write, Not on Dispatch
- Lead-to-Teammate Plan-Approval Handshake for Multi-Agent Work
- Learning Execution Guardrails from Agent Failure Traces
- Legacy Code Archaeology: Reconstruct Intent Before Migrating
- Lethal Trifecta Threat Model for AI Agent Development
- Lexical-First Retrieval for Agentic Search: When BM25 Is Enough
- Lifecycle-Integrated Security Architecture for Agent Harnesses
- Live Browser as Agent Context Channel
- Local Model Viability Factors for Coding
- Local Sandboxing in the Copilot App: The Credential Axis
- Lock-State Safeguards for Desktop-Controlling Agents
- Long-Running Agents: Durability and Resumability Across Sessions
- Loop Detection for AI Agents: Stopping Micro-Loops
- MCP Client Design: Building Robust Host-Side Logic
- MCP Runtime Control Plane: Policy Evaluation Between Agent and Tool
- MCP Server Design: Building Agent-Friendly Servers
- MCP: The Open Protocol Connecting Agents to External Tools
- Machine-Readable Error Responses for AI Agents (RFC 9457)
- Magentic Orchestration: Task-Ledger-Driven Adaptive Multi-Agent Planning
- Making Application Observability Legible to Agents
- Managed vs Self-Hosted Agent Harness: Deployment Trade-offs
- Measuring the Verification Tax on Agent Output
- Memory Retrieval as a Control Decision
- Memory Synthesis: Extracting Lessons from Execution Logs
- Memory Transfer Learning: Cross-Domain Memory Reuse in Coding Agents
- Memory-Induced Tool-Drift in LLM Agents
- Mermaid as Agent Output Format: When to Ask for a Diagram Instead of Prose
- Method Map: Failure-Mode to Smallest-Artifact Triage
- Mid-Trajectory Guardrail Selection for Multi-Step Tool Calls
- Minimum-Sufficient Control Ladder: Escalate by Failure Mode
- Minimum-Sufficient Execution: Estimate Scope Before Spending Budget
- Mise en Place for Agentic Coding
- Model Deprecation Lifecycle for Agent Workloads
- Model a Single Agent Turn as Many Inference and Tool-Call
- Model-Directed Subagent Tiering: Lead Model Picks the Tier
- Model-ID-as-Dependency: Migration Protocol for Deprecation Churn
- Model-Neutral Agent Architecture: Model Portability Over Cloud Portability
- Model-Set Parity: Reading Harness Efficiency Claims
- Monitor Tool: Event Streaming from Background Scripts
- Monitor or Wait: The Supervision Choice During Agent Execution
- Monolith-to-Sub-Agents Refactor: Five Lessons from a Brittle Prototype
- Monorepo Skill and Agent Discovery: Hierarchical Configuration
- Monotonic Capability Attenuation for Composition-Safe Tool Use
- Most-Restrictive-Wins Fusion for Parallel Agent Control Returns
- Multi-Agent RAG for Spec-to-Test Automation
- Multi-Agent SE Design Patterns: A Taxonomy Across 94 Papers
- Multi-Agent Systems: Coordination and Orchestration
- Multi-Agent Topology Taxonomy: Centralized, Decentralized
- Multi-Client Session Attachment: One Session, Many Clients
- Multi-Model Plan Synthesis for System Architecture
- Multi-Repo and No-Repo Coding Agent Automation Templates
- Multi-Shape BYOK Provider: Declare API Family per Endpoint
- Multi-Turn Conversation Evaluation: Per-Turn and Trace-Level Scoring Together
- Multitenant RAG: Closing the Relevance-Authorization Gap
- Natural Language Tool Selection (NLT)
- Non-Human Event Provenance Markers to Block Fabricated Approvals
- Non-Retirable Approval Rules for Agent Operations
- Nonstandard Errors in AI Agents: Model-Family Variance
- Notebook-Documented Automation for Repeat Operational Work
- OAuth Client ID Metadata Documents (CIMD) for MCP Servers
- OWASP 2026 Update for Agent Builders: Top 10 Renumbering and the Agent Control Standard
- OWASP LLM Top 10 (2025): Agent Security Crosswalk
- Objective Drift: When Agents Lose Sight of the Goal
- Observability-Driven Harness Evolution
- Observation Contract Preservation in Tool-Augmented Agents
- Observation-Driven Coordination: CRDT-Based Parallel Agent
- Offline Evaluation as an Integration Test for LLM Features
- On-Demand Skill Hooks: Session-Scoped Guardrails via Skill Invocation
- One-Click CI Auto-Fix: Human-Triggered Cloud-Agent Remediation for Failing GitHub Actions
- One-Shot Record and Deterministic Replay for Periodic Agent Tasks
- Open Agent School Pattern Mapping for Practitioners
- OpenAI Agents SDK
- OpenAPI as the Source of Truth for Agent Tool Definitions
- OpenTelemetry for AI Agent Observability and Tracing
- Opponent Processor / Multi-Agent Debate Pattern
- Oracle Poisoning: Knowledge Graph Corruption Against Tool-Using Agents
- Oracle-Based Task Decomposition for AI Agent Development
- Oracle-Gated Delegation Beyond Your Domain Expertise
- Orchestrator-Worker Pattern for AI Agent Development
- Organizational Context Layer for Agents (Company Brain)
- Organizing Filesystem Agent Memory for Retrieval Cost
- Outcome Monitors: Recovery Affordances for Tool Failures
- Outcome Pricing as a Scope Signal
- Over-Orchestrated Agent Architecture (Prefer the Simplest That Works)
- Overeager-Behavior Elicitation: Scope + Trap Fragments as a Diagnostic for Out-of-Scope Tool Calls
- Override Pattern: Reusing Interactive Commands in Automated Pipelines
- PASS@(k,T): Evaluate RL for Agents Along Sampling and Interaction Depth
- PEEK: Orientation Cache for Recurring-Context Agents
- PR-Subscribed Agent Ownership: The Agent That Opened the PR Drives It to Green
- Parallel Agent Sessions Shift the Bottleneck from Writing
- Parallel Polyglot Ports as a Spec-Ambiguity Oracle
- Parameter-Keyed Caching and Dependency-Aware Parallelism for Plan-Execute Pipelines
- Parameter-Level Permission Rules (Tool(param:value) Syntax)
- Parser-Versus-Shell Evasion in Command Permission Checks
- Parsimonious Agent Routing for Multi-Agent Dispatch
- Path-Scoped Write Contracts for Shared Agent State
- Pattern Replication Risk in Agentic Code Generation
- Pattern Selection Map
- Patterns: Agent Design, Multi-Agent, and Anti-Patterns
- Peer Refusal as a Coordination Control
- Per-Agent Capability Stores Beat One Task-Wide Allowlist
- Per-Call Budget Hints on Tool Invocations
- Per-Caller Identity: Who an Agent's Tool Call Acts As
- Per-Layer Suppression Accounting in Acceptance Gates
- Per-Model Harness Tuning: Treating the Backing Model as a Harness Variable
- Per-Run Budget Reservation for Coding Agent Model Calls
- Per-Task Agent Routing Across Coding Harnesses
- Per-Task Verification Budget: Size the Task to Fit the Check
- Per-Tool Extended Reasoning Opt-In: Tool-Call-Scoped Budgets
- Per-User Supervisor Process for Background Agent Sessions
- Permission Framework Choice Outweighs Model Choice for Limiting Overeager Actions
- Permission Gates That Deny the Agent's Own Cleanup (Denied Remediation Path)
- Permission Modes as a Defense Against a Tampered Response Path (Response-Path Control Gap)
- Permission-Gated Custom Commands for AI Agent Development
- Permutation Frameworks for Batch Code Generation
- Persistent Shared Search Sub-Agent for Output-Token Reuse
- Persistent Teammate Workspace: Durable State for Agent Teams
- Persistent-Connection Agent Transport
- Persona-as-Code: Defining Agent Roles as Structured Docs
- Phase-Specific Context Assembly for AI Agent Development
- Plan Compliance in Agents: Measure What They Execute, Not What You Wrote
- Plan-Then-Execute as the Default for Web Agents
- Planning Stage Preconditions: Budget Headroom and Task Text
- Plugin Dependency Declaration and Disable-Chain Hints
- Plugin and Extension Packaging: Distributing Agent Capabilities
- Poka-Yoke for Agent Tools: Mistake-Proof Tool Interfaces
- Portable Agent Definitions: Full-Stack Identity as Code
- Porting a Coding-Agent Harness Beyond Engineering
- Pre-Change Impact Analysis: Dependency Maps That Prevent Agent Regressions
- Pre-Completion Checklists for AI Agent Development
- Pre-Execution Codebase Exploration for AI Coding Agents
- Pre-Install Plugin Transparency: Capability Inventory and Cost Projection
- Pre-Trust Execution Surface in Coding Agent Harnesses
- Pre-Write Change Intent Admission (Claim Plane)
- Premature Completion: Agents That Declare Success Too Early
- Pricier-Per-Token Models That Cost Less Per Task
- Prior Dominance Over Feedback in Agent Optimization Loops
- Privacy-Preserving LLM Requests: Eight Techniques and a Practical Combination
- Proactive Idle-Time Anticipation (ProAct)
- Process Amplification: Scaling Human Work with Agents
- Product-Operation Tool Surfaces for In-Product Agents
- Product-as-IDE: When the Application Becomes the Development
- Production Hosting Topology for Self-Hosted Agent SDK Runtimes
- Production MCP Agent Stack: Sequencing Six Decisions into One Deployment
- Programmatic Agent Session Export via `claude agents --json`
- Programmatic Cloud-Agent Dispatch via REST API and Webhooks
- Progressive Autonomy: Scaling Trust with Model Evolution
- Progressive Disclosure for Layered Agent Definitions
- Progressive Spend Threshold Alerting for Agent Cost Governance
- Project Writing Skill: House Style as Model-Invocable Skill
- Project-Scoped Agent Workspace: Durable Context for Clean-Context Subagents
- Prompt Caching: Architectural Discipline for Agents
- Prompt Chaining: Sequential LLM Calls for Agent Workflows
- Prompt Debt: Hand-Tuning Natural-Language Prompts as Technical Debt
- Prompt Injection: A First-Class Threat to Agentic Systems
- Prompt-Only Baseline Before a Specialized Agent Subsystem
- Prompted Uncertainty Decomposition for Clarification Routing
- Proof of Presence: Re-Authenticating for Agent Actions
- Proprietary-to-Open-Standard Tool Migration (Copilot Extensions to MCP)
- Protecting Sensitive Files from Agent Context Access
- Protecting the Test Oracle From the Agent
- Prototype Before Optimizing: Establish Quality Baselines Before Token Constraints
- Provenance-Aware Decision Auditing for LLM Agents
- Provider-Hosted Subagent Delegation: One Model Price for the Whole Tree
- Public Placeholder Credentials with a Fail-Closed Injecting Proxy
- Public-Channel Agent Work as Lehrwerkstatt for Team Learning
- QA Session to Issues Pipeline for AI Agent Development
- Quality Score Rubric and Simplification Log for Agent Harnesses
- RAG Architecture as a Poisoning Robustness Decision
- RAG over Thinking Traces: Index Reasoning Trajectories Instead of Documents
- RL-Trained Automated Red Teamers for Prompt Injection Discovery
- Rainbow Deployments for Agents: Gradual Version Migration
- Re-Run an Agent's Speed-Up Claim Before Merging
- ReAct (Reason + Act): Interleaved Reasoning-Action Loops
- Reactive Environment Hooks: CwdChanged and FileChanged
- Reader-Scoped Trace Views: When to Build Your Own UI
- Reasoning Budget Allocation: The Reasoning Sandwich
- Reasoning Effort Over Tool Scaffolding for First-Try Reliability
- Reasoning Retention and Compaction as Harness Settings
- Recoverability-Gated Self-Evolution for Agent Harnesses
- Recurring Control Belongs in Harness Code, Not Context
- Recursive Agent Harnesses (RAH)
- Recursive Best-of-N Delegation
- Recursive Sub-Agent Delegation: Depth Limits and Trade-offs in Nested Hierarchies
- Red-Green-Refactor with Agents: Tests as the Spec
- Reducing Fixed CI Overhead Before Adding Shards
- Reflective Prompt Evolution with Pareto Selection (GEPA)
- Reframed Exfiltration Defeats Wording-Based Defenses (Framing Gap)
- Reliability of an Automatically Selected Agent Harness
- Remote Agent Host Sessions over SSH and Dev Tunnels
- Remote Session Control for Local CLI Agents
- Renting a Cloud Agent's Execution Sandbox
- Replayable Encrypted Reasoning Blocks in Agent Traces
- Repository Bootstrap Checklist: Wiring Agent Support
- Repository Map Pattern: AST + PageRank for Dynamic Code
- Repository Perturbation as Context-Reasoning Diagnosis (RepoMirage)
- Repository-Level Retrieval for Code Generation
- Residual Completion for Stateful Agent Handoffs (CFRC)
- Restricting a Coding Agent to a Single execute_code Tool
- Retrieval-Augmented Agent Workflows: On-Demand Context
- Retry-Switch-Abstain: A Runtime Tool-Recovery Policy
- Review-Then-Implement Loop for AI Agent Development
- Reviewing What a Memory-Fed Autofix Taught
- Revocable Resource-and-Effect Capabilities for Coding Agents (PORTICO)
- Rigor Relocation: Engineering Discipline with AI Agents
- Risk-Based Shipping: Review by Risk Matrix, Not by Default
- Risk-Based Task Sizing for Agent Verification Depth
- Role Orchestration on a Single Model
- Role-Declared Context Mode: Where the Inherit Decision Lives
- Rollback-First Design: Every Agent Action Should Be Reversible
- Route-Parity Auditing of Agent Safeguards
- Router-Imposed Quality Ceiling: Committing Before Output
- Routing Break-Even: When a Cheaper Model Actually Pays
- Routing Dependency Updates to Repair Agents by Budget
- RubricRefine: Pre-Execution Rubric Refinement for Code-Mode Tool Use
- Runbooks as Agent Instructions: Agent-Followable Ops
- Running Several Coding Agents Behind One Harness Interface
- Runtime Guard as an Installed Skill (Defense-as-Skill)
- Runtime Scaffold Evolution: Agents That Build Tools
- Runtime Workflow Selection Across Models (Project HydraFusion)
- SDLC-Phase Skill Taxonomy: Full-Lifecycle Skill Libraries
- SKILL.md Frontmatter Reference: All Fields Explained
- Safe Command Allowlisting: Reducing Approval Fatigue
- Safe Outputs Pattern for Trustworthy Agent Responses
- Sandbox + Approvals + Auto-Review Governance Triad
- Sandbox Credential Masking: Authenticate Without Seeing the Secret
- Sandbox Forking: Branch Agent Runs From a Warm Snapshot
- Sandbox-Enforced PII Tokenization in Agent Workflows
- Sandboxed Coding Environments: Containers vs MicroVMs vs OS-Level Isolators
- Scheduled Instruction File Fact-Checker for Accuracy
- Schema-Guided Graph Retrieval
- Scope Sandbox Rules to Harness-Owned Tools, Not Third-Party
- Scope-Matched Retrieval for Persisted Agent Skills
- Scoped Browser DevTools Access for Runtime Diagnosis
- Scoped Credentials via Proxy Outside the Agent Sandbox
- Scoring Constraint Loss and Tool Reach as Separate Risks
- Scout-Then-Route: Verify the Handoff Before Routing
- Seamless Background-to-Foreground Handoff
- Secrets Management for AI Agents: Credential Injection
- Security Drift in Iterative LLM Code Refinement
- Security for AI Agent Development
- Security-Aware Tool Descriptions for MCP Servers (SpellSmith)
- Selective Autonomy from Copilot Feedback
- Selective Checkpoint Restore Across Code and Conversation State
- Selective Network Access in Agent Sandboxes: The allowNetwork Pattern
- Selective Revalidation for Pending Agent Actions
- Self-Correcting Memory: Evidence-Backed Claim Repair
- Self-Discover Reasoning: LLM-Composed Reasoning Structures
- Self-Healing Production Agent: Automated Regression Detection and Autofix PR
- Self-Healing Tool Routing
- Self-Reporting Loops: Autonomous Routines That File Their Own Backlog
- Self-Rewriting Meta-Prompt Loop
- Semantic Context Loading: Language Server Plugins for Agents
- Semantic Intent Validation for Agent Skills
- Semantic Issue Search from Chat vs Query Syntax
- Semantic Tool Output: Designing for Agent Readability
- Separation of Knowledge and Execution in Agent Systems
- Session Harness Sandbox Separation for Long-Running Agents
- Session Initialization Ritual: How Agents Orient Themselves
- Session Recap: Goal-Shaped Handoff at Context Boundaries
- Shadow Tech Debt Created by Autonomous AI Agent Commits
- Shallow Agent Test Coverage from Premature Termination (Lazy Generation)
- Shared Agent Context Store API: When to Expose Curated Context as an Endpoint
- Shared Context Bundle Registry for Agent Teams
- Shortening Old Tool Results Under Context Pressure (Half-Life Truncation)
- Silent-Failure Mechanism Taxonomy in Production Agent Runtimes
- Simulation and Replay Testing for Agent Verification
- Single-Branch Git for Agent Swarms: A Trade-Off Pattern
- Single-CLI Agent Platform: Create to Production in One CLI
- Single-Layer Prompt Injection Defense Anti-Pattern
- Situated Harness Layers: Fix at the Layer That Owns It
- Six-Shape Approval Response Taxonomy: Beyond Binary Allow/Deny
- Sizing Vendor-Emitted Agent Telemetry by Signal Tier
- Sizing an Agent Migration Fan-Out by Diff Uniformity
- Skeleton Projects as Agent Scaffolding
- Skill Authoring Patterns: Description to Deployment
- Skill Composition Risk in Agent Ecosystems
- Skill Library Evolution: Lifecycle Governance for Agents
- Skill Library Refinement Loops: Organizational Feedback for Shared Skills
- Skill Library Technical Debt: Library-Time Maintenance for Agent Skills
- Skill Loadout Curation for Coding Agents
- Skill Misevolution in Self-Updating Skill Libraries
- Skill Over-Trust: Treating Topical Relevance as Evidence a Skill Helps
- Skill Program Functions: Executable Guardrails Compiled From Past Failures
- Skill Reuse as Vendored Forking
- Skill Supply-Chain Poisoning
- Skill Tool as Enforcement: Loading Command Prompts at Runtime
- Skill as Instruction Surface and Callable API (Interpreter Skills)
- Skill as Knowledge Pattern for AI Agent Development
- Skill or MCP Server: Choosing a Capability's Delivery Mechanism
- Solver-Externalized Constraint Reasoning (MaxSAT/SMT Encoding)
- Source-Grounded Test Plan with Pre-Action Assertion Annotation
- Sparse-Checkout Worktrees for Monorepo Agent Isolation
- Spec-Anchored Drift-Gated Architecture (Spec Growth Engine)
- Spec-Driven Development with Spec Kit
- Specialist Orchestrated Queuing for Multi-Agent SE (SPOQ)
- Specialized Agent Roles for Effective AI Pipelines
- Specialized Small Language Models as Agent Sub-Tools
- Specification Authority Boundary: Agents Propose, the Runtime Commits
- Specification-First Convergence Without a Test Oracle
- Splitting the Drift Judge from the Advisor (LivePlan)
- Sprint Contracts: Pre-Coding Success Agreements for Multi-Agent Tasks
- Stacked Agent Sessions on Unmerged Feature Branches
- Stage Elision Before Summarization
- Staged Evidence Gates for Agentic Program Repair
- Staged Literal Porting with a Per-Stage Numeric Oracle
- Staggered Agent Launch: Preventing Thundering-Herd in Swarms
- Stakeholder Trust Through Evals and Observability
- State-Bound Evidence and Typed Revision Contracts for Repair Loops
- Stateful Agent Evals via State Snapshots and Transition Assertions
- Stateless MCP: One Request per Tool Call
- Static Roster vs Runtime Subagent Definition
- Static-First Shell Command Gating with Selective Escalation
- Steering Running Agents: Mid-Run Redirection and Follow-Ups
- Stochastic-Deterministic Boundary as First-Class Contract
- StopFailure Hook: Observability for API Error Termination
- Strained Coherence as a Pre-Failure Signal in Agent Trajectories
- Strategy Over Code Generation: Why AI Speed Doesn't Fix Wrong Goals
- Structure Prompts with Static Content First to Maximize Cache Hits
- Structured Agentic Software Engineering (SASE)
- Structured Domain Retrieval: Knowledge Graphs and Case-Based Reasoning
- Structured Task-State Ledger for Tool-Calling Agents (LedgerAgent)
- Stuck-Loop Recovery: Detecting and Escaping Non-Converging Agent Loops
- Sub-Agents for Fan-Out Research and Context Isolation
- Subagent OTel Trace Correlation via agent_id Attribute
- Subagent Schema-Level Tool Filtering for AI Agents
- Subagent vs In-Context Skill Execution
- Subtask-Level Memory for Software Engineering Agents
- Sufficiency-Tightness Decomposition for Agent-Authored Permissions
- Suspect the Harness Before the Model on a Regression
- Symptom-First Bug Triage for Agent Code
- System-Level Optimization Pipeline
- TDD Interaction Models: Throughput Versus Test Quality
- Tail Control for Agent Workflows: Engineering for the Failure Tail, Not the Average
- Task Completion as Tool Certification (Silent Tool Rot)
- Task Feasibility Awareness: Stop Before You Start
- Task Shape Decides What a Heavier Agent Harness Buys
- Task-Based Access Control with Hybrid Inspection
- Task-Specific Agents vs Role-Based Agents
- Task-Uniform Agent Permissions Ignore Where Failures Land
- Team Onboarding for AI Agent Workflows and Adoption
- Temporary Compensatory Mechanisms in Agent Harnesses
- Tenant Model Policy: Organization-Scoped Rules for AI Model Selection
- Terminal Tools for Agents: send_to_terminal and Background Interaction
- Terminal-First Agent Interfaces with Browser Escalation
- The 7 Phases of AI-Assisted Feature Development
- The AI Development Maturity Model: From Skeptic to Agentic
- The AX Stack: A Layered Model of an AI Coding Agent's Prompt-to-Compile Path
- The Advisor Strategy: Frontier Model as Strategic Advisor
- The Agent Stack Bet: Architectural Decisions for Production Agents
- The Bottleneck Migration When Humans Supervise Agents
- The Compliance Trap: Consuming Conflicting Agent Memory
- The Copy-Paste Agent Anti-Pattern in AI Development
- The Decoding Harness Is Part of the Agent Attack Surface
- The Delegation Decision: When to Use an Agent vs Do It Yourself
- The Error-Class Governance Loop for Instruction Libraries
- The Harness as Product: What Listed-Rate Pricing Buys
- The Orchestrator's Attention Budget: Delegating to Protect Context
- The Patchwork Problem in LLM-Generated Code
- The Plan-First Loop: Always Design Before Writing Code
- The Post-Authorization Execution Trust Gap in Remote MCP
- The Reasoning-Complexity Trade-off
- The Research-Plan-Implement Pattern
- The Software Factory Model: Industrializing Agent Loops
- The Subagent Inheritance Contract: What Crosses Down
- The Think Tool: Mid-Stream Reasoning for AI Agents
- The Yes-Man Agent: Compliance Without Verification
- Three Reasoning Spaces: Plan-Bead-Code Phase Gates
- Three-Vector Evasion Taxonomy for Agent Security Tests
- Throwaway-Prototype Skill: Build to Discard, Keep Only the Answer
- Tiered Memory Architecture: Episodic-to-Semantic Consolidation for Long-Running Agents
- Tiled Agent Layout: Supervising Parallel Agents Through Dedicated Panes
- Tool Architecture Moves Consistency, Not Resolve Rate
- Tool Calling Schema Standards for AI Agent Development
- Tool Confirmation Carousel: Batched UI for Per-Call Approvals
- Tool Description Quality for Effective Agent Guidance
- Tool Engineering (Training Module)
- Tool Engineering: Designing and Managing AI Agent Tooling
- Tool Minimalism and High-Level Prompting
- Tool Necessity Probing: Reading Tool-Call Decisions From Hidden States
- Tool Operability: Interfaces That Survive a Lost Response
- Tool Preamble: User-Visible Status Updates Before Tool Calls
- Tool Signing and Signature Verification for Agents
- Tool-Call Success as Workflow Effect Evidence
- Tool-Invocation Attack Surface in Coding Agents
- Tool-Schema Bias: Measure the Interface Before You Ship
- Tool-Use Sim-to-Real Perturbation Taxonomy
- Tools as Typed Code Stubs (Programmatic Tool Calling)
- Toolset Agentization: Wrapping Co-Used Tools as Sub-Agents
- Training Modules
- Training-Data Gravity: Agents Default to Deprecated APIs
- Trajectory Decomposition: Diagnose Where Coding Agents Fail
- Trajectory Logging via Progress Files and Git History
- Trajectory Pre-Filter for Failure Diagnosis (TrajAudit)
- Trajectory Projection: A Flattened View of Agent Traces
- Trajectory-Conditioned Model Escalation (SWE-Router)
- Treat Task Scope as a Security Boundary
- Treating Agent Safety as Uniform Across a Session (Cold-Start Safety Gap)
- Treating a Clean Final State as Boundary-Compliance Evidence
- Treating a Clean Merge as Compatibility Evidence
- Treating a Local Agent Session Trace as Audit Evidence
- Trigger-Level Gating for Autonomous Agent Intake
- Trigger-to-Function Architecture for Unprompted Codebase Maintenance
- Trusting Tool Error Messages as Implicit Authority (Error-Path Injection)
- Typed Context Buys Addressability, Not Token Savings
- Typed Memory Provenance and Assertion Release Gating
- Typed Schemas at Agent Boundaries for Multi-Agent Systems
- Typed Tool Surfaces and Out-of-Loop Correctness Gates
- Unbounded Agent Feedback Paths (Infinite Agentic Loops)
- Unix CLI as the Native Tool Interface for AI Agents
- Unsignalled Tool Failure: Returning Success With an Unusable Payload
- Usage-Reinforced Memory Decay for Long-Running Agents
- Use a Public-Web Index to Gate Automatic URL Fetching
- Using the Agent to Analyze Its Own Evaluation Transcripts
- Utility-Model Split: Background Tasks on a Cheaper Model
- VS Code Agents App: Agent-Native Parallel Task Execution
- Velocity-Quality Asymmetry: Why AI Speed Gains Fade
- Verbatim Failure Records in Small-Model Agent Transcripts
- Verification Ledger for Tracking Agent Output Quality
- Verification-Centric Development for AI-Generated Code
- Verification-Gated Agent Autonomy via Automated Review
- Verifier-Driven Parallel Coding Agents (Glite ARF)
- Verify-Gated Completion as Admission Control
- Vetting Tool Definitions for Exfiltration Signatures
- Visual-Prompt Agent Steering (Cursor Design Mode)
- Voting / Ensemble Pattern for AI Agent Development
- WIP=1 and Little's Law: Kanban Throughput Theory for Agent Task Design
- Weakest Consistent Learning: What Agent Loops Should Persist
- Web Search Agent Loop: Iterative Research Patterns
- WebMCP: Browser-Hosted Tool Contracts for In-Page AI Agents
- Which Task You Delegate Changes Poisoned-Repo Exposure
- Whole-Codebase Visibility as a Migration Prerequisite
- Wiki Memory: Agent-Maintained Compressed Knowledge Base
- Windows Sandboxing for Coding Agents
- Workflows for AI Agent Development
- Working Inside an Enterprise-Managed Agent Sandbox Policy
- Workload-Keyed Sandbox Selection for Agent-Generated Code
- Workspace Topology as an Indirect Injection Attack Vector
- Worktree Isolation: Parallel Agent Sessions in Safe Sandboxes
- Write Tool Descriptions as Agent Onboarding Documents
anti-pattern¶
- AI Agent Development Anti-Patterns and Failure Modes
- AI Agents in CI/CD with Elevated Permissions and Untrusted Content (GitInject)
- Absorbing Entry-Level Work into Senior-Agent Workflows
- Abstraction Bloat in AI Agent-Generated Code Output
- Adversarial-Only Threat Modeling for Agent Data Leakage
- Agent Extension Conflicts: When Installed Skills and MCP Servers Fight Each Other
- Agent Headcount as a Vanity Metric
- Agent Sprawl: Unmanaged Sub-Agent and Skill Proliferation
- Agent-Laundered Bug Reports
- Agentic Skill Decay: Which Capabilities Erode Under Agent Delegation
- Answer-Reachable Eval Environments
- Artifact-Only Verification Hides Skipped Skill Steps
- Ask-Everything Permission Policies Protect Less than Per-Action Approval
- Assertion-Free Test Theater in Agent-Authored Patches
- Assuming Agent Interchangeability in Long-Running Teams
- Assuming Loaded Skills Stay Enforced in Long Contexts
- Assuming a CLAUDE.md Security Rule Is Enforced
- Assumption Propagation: Compounding Agent Misunderstandings
- Belief Inertia After Tool-Map Drift in AI Agents
- Blaming the Model for Scaffolding-Driven Quality Regressions
- Blind Tool Deference: Agents Parroting Callable Tools
- Boring Technology Bias: When Agents Recommend by Popularity
- Cargo Cult Agent Setup: Copying Without Understanding
- Catastrophic Remembering: Instruction Files That Only Grow
- Cheaper-Per-Token Model Upgrades That Cost More Per Task
- CodeSlop: Search-Trajectory Residue in Agent Patches
- Coding-Agent Misalignment Forms (Seven-Symptom Taxonomy)
- Completion Summary as the Oversight Surface
- Compound Prompt Constraints Degrade More Than Their Parts Predict
- Conceptual Integrity Erosion in Agent-Built Codebases
- Configuration Smells in AGENTS.md Files (Six-Smell Catalog)
- Constraint Tax: Tool Suppression Under JSON Schema Decoding
- Constraint-Evasive Fabrication in Instruction Sets
- Context Poisoning: When Hallucinations Become Premises
- Context-Fractured Decomposition Attacks on Tool-Using Agents
- Cost-Driven Model Routing Without Quality Monitoring
- Cost-Inefficient Behaviors in Coding Agents
- Cross-Component Interference in Agent Scaffolds
- Cross-Session Re-Implementation of Existing Agent Code
- Cumulative-Best Reporting Hides Repair-Loop Security Regressions
- Declared Peer Consensus as Context for a Reviewing Agent
- Delegating Change Descriptions to the Agent
- Deletion Avoidance: Agents That Guard Code Instead of Removing It
- Deliberation-Inducing Cues That Multiply Reasoning Cost
- Demo-to-Production Gap: When Demos Hide Real Costs
- Density-Normalized Quality Metrics Mask AI-Driven Code Growth
- Dependency Inlining Erodes SBOM and License Provenance
- Destructive-Failure Mechanism Attribution by Mitigation Owner (ClayBuddy Three)
- Direct Prompt Injection via Collaboration (User as Attack Vector)
- Distractor Interference: Why Relevance Is Not Enough
- Documenting Code the Agent Can Already Read
- Dynamic Tool Fetching Destroys KV Cache Performance
- Entity Binding Failures in Tool-Augmented Agents
- External Artifacts Treated as Data, Not Adversarial Input
- First-Proposal Execution in Agent Loops
- Framework-First Agent Development: An AI Anti-Pattern
- Frozen Playbook Reuse Without Target-Side Validation
- Generating Tests From Agent-Written Code (Code-First Oracle Bias)
- Happy Path Bias: How AI Agents Skip Error Handling
- Held-Out Tasks as a Harness Shortcut Defense
- Homogeneous Debate Panels as a Groundedness Quality Lever
- Indiscriminate Structured Reasoning on Every Agent Task
- Install-Once Plugin Trust: Vetting That Never Re-Runs on Update
- Judging Agent Safety by Task Completion (Action-Boundary Violations)
- Judging a Skill's Honesty by the Validity of Its Output
- Knowledge Cutoff as a Documentation Boundary
- LLM API Routers as Application-Layer Man-in-the-Middle
- LLM Code Review Overcorrection for AI Agent Development
- LLM Self-Review Failure in Code Modernization Tasks
- LLM Support During the First Detection Pass
- Large-Codebase Coding-Agent Failure Patterns (Sourcegraph Five)
- Law of Triviality in AI PRs for AI Agent Development
- MCP Allowlist by Label, Not by Identity (serverName Trap)
- Memory-Induced Tool-Drift in LLM Agents
- Mid-Session Config Changes as Invisible Cache Invalidators
- Minimality Prompts as a Patch-Size Control
- Model Confidence as Security Verification (Security Calibration Gap)
- Multi-Agent Shared State Isolation Anomalies
- Multi-Tool Threshold Poisoning Against MCP (ShareLock)
- Objective Drift: When Agents Lose Sight of the Goal
- Over-Orchestrated Agent Architecture (Prefer the Simplest That Works)
- Overtrusting Human Sign-Off on Generated Assertions
- PR Scope Creep as a Human Review Bottleneck
- Pattern Replication Risk in Agentic Code Generation
- Patterns: Agent Design, Multi-Agent, and Anti-Patterns
- Peer Refusal as a Coordination Control
- Perceived Model Degradation: Why Vibes Are Not Evals
- Permission Framework Choice Outweighs Model Choice for Limiting Overeager Actions
- Permission Gates That Deny the Agent's Own Cleanup (Denied Remediation Path)
- Permission Modes as a Defense Against a Tampered Response Path (Response-Path Control Gap)
- Pooled-Evidence Factuality Checks for MCP Agents (Cross-Source Conflation)
- Premature Completion: Agents That Declare Success Too Early
- Prescribing TDD Inside the Agent Loop (Process Theater)
- Pressuring a Coding Agent Degrades the Code It Writes
- Prior Dominance Over Feedback in Agent Optimization Loops
- Prompt Debt: Hand-Tuning Natural-Language Prompts as Technical Debt
- Prompt as Security Knob
- Prompt-Only Tool Access Control
- Reading Visible Edge-Case Handling as a Security Check
- Refactoring Runaway: Tangled Refactorings in Agent Patches
- Repository Skill Release Drift
- Rewriting a CLI Into a JSON Payload for Agents
- Rolling Out a Team-Embedded Agent Like a Tool
- Run-Status vs Task-Status Confusion in Autonomous Agent Runs
- Scoped-Looking Permission Grants
- Scoring Constraint Loss and Tool Reach as Separate Risks
- Semantic Collapse Under Underspecified Prompts
- Serving-Stack Confounds in Tool-Call Evaluation
- Setup Documentation as an Install-Time Attack Vector
- Shadow Tech Debt Created by Autonomous AI Agent Commits
- Shallow Agent Test Coverage from Premature Termination (Lazy Generation)
- Silent Adoption of Corrupted Tool Returns by Agents
- Silent-Failure Mechanism Taxonomy in Production Agent Runtimes
- Single-Decision Approval of Vendor Skill Suites
- Single-Layer Prompt Injection Defense Anti-Pattern
- Skill Atrophy: When AI Reliance Erodes Developer Capability
- Skill Over-Trust: Treating Topical Relevance as Evidence a Skill Helps
- Skill Review Without a Token Cost Baseline
- Slopsquatting: Hallucinated Package Names as a Supply-Chain Vector
- Spec Complexity Displacement: When Specs Become Code
- Stale AI Configuration Artifacts (Context Rot)
- Symptom-Reduction-as-Root-Cause: Why Oracle Tests Alone Miss Architectural Drift
- Tab-Accept Rate as a Proxy for Critical Engagement
- Task Completion as Tool Certification (Silent Tool Rot)
- Task-Uniform Agent Permissions Ignore Where Failures Land
- Team Shared-Language Desync from Removed Review Friction
- Test Oracles That Read Their Expectation From the Code
- The Agent-on-Agent Maintenance Penalty
- The Anthropomorphized Agent for AI Agent Development
- The Compliance Trap: Consuming Conflicting Agent Memory
- The Context Ceiling -- Where AI Fails Expert Architects
- The Copy-Paste Agent Anti-Pattern in AI Development
- The Effortless AI Fallacy for AI Agent Development
- The Implicit Knowledge Problem for AI Coding Agents
- The Infinite Context Anti-Pattern in Agent Systems
- The Kitchen Sink Session Anti-Pattern in AI Agents
- The Meat Proxy: Relaying Agent Output Without Reading It
- The Patchwork Problem in LLM-Generated Code
- The Prompt Tinkerer Anti-Pattern in Agent Workflows
- The Reasoning-Complexity Trade-off
- The Recall Trap: Tuning a Code Retriever on Recall@k at a Fixed Slot Budget
- The Test Homogenization Trap: When LLM-Generated Tests Mirror Model Blind Spots
- The Yes-Man Agent: Compliance Without Verification
- Token Preservation Backfire for AI Agent Development
- Token Reduction Mistaken for Cost Reduction
- Tool-Call Success as Workflow Effect Evidence
- Training-Data Gravity: Agents Default to Deprecated APIs
- Treating Agent Delegation as Routing, Not Authorization
- Treating Agent Safety as Uniform Across a Session (Cold-Start Safety Gap)
- Treating File Secrecy as Skill Confidentiality
- Treating Memory-Injection Rate as Security Evidence
- Treating Read Denial as Confidentiality for a Build Input
- Treating a Clean Final State as Boundary-Compliance Evidence
- Treating a Clean Merge as Compatibility Evidence
- Treating a Clean Static Scan as Security Evidence
- Treating a Local Agent Session Trace as Audit Evidence
- Treating a Worktree as a Safety Boundary
- Trust Without Verify: Skipping Agent Output Checks
- Trusting Claimed Prior Approval in Agent Review Gates
- Trusting Human Review to Catch Deliberate Agent Sabotage
- Trusting Model-Level Privilege Restraint at Tool Selection
- Trusting Tool Error Messages as Implicit Authority (Error-Path Injection)
- Trusting a Skill Scanner's Verdict as a Security Judgment (Green-Check Fallacy)
- Unbounded Agent Feedback Paths (Infinite Agentic Loops)
- Unsignalled Tool Failure: Returning Success With an Unusable Payload
- Unstated-Contract Bugs: Sort Tickets by Information Gap
- Unversioned Scaffolding Commands Pull Stale Templates
- Verbatim Failure Records in Small-Model Agent Transcripts
- Vibe Coding: Outcome-Oriented Agent-Assisted Development
- When Developers Understand Less of Their Own Codebase
- bypassPermissions Silently Overrides allowedTools (The Restricted-Bypass Trap)
arxiv¶
- AI Agents in CI/CD with Elevated Permissions and Untrusted Content (GitInject)
- AI Bot CI/CD Workflow Reliability by Agent
- AI Label as Reviewer Attention Redistribution
- AIRA: Inspection Framework for AI-Generated Code
- AOCI: Symbolic-Semantic Repository Indexing
- AST-Grounded Critic Loop for Documentation Maintenance
- AX/UX/DX Triad: Three Experience Layers in Agent Systems
- Absorbing Entry-Level Work into Senior-Agent Workflows
- Accumulated Behavioral Rules from Review Feedback
- Action-Class Decomposition for Tool-Calling Evals
- Action-Gated Context Trimming for Long-Horizon Agents
- Action-Graded Severity for Agent Red-Team Outcomes
- Adapting AI Assistants to Developer Interaction Style
- Adaptive Evaluation of Out-of-Band Prompt-Injection Defenses
- Adaptive Generate-Rank-Verify Under Costly Verification
- Adaptive Validation Task Selection
- Addressable Recall Compaction: Compact to Citations, Not Summaries
- Adversarial-Only Threat Modeling for Agent Data Leakage
- Advisory Prompts Distilled from Reasoning Traces
- Against-Prior Accuracy: Score the Rules That Fight Defaults
- Agent Approval Laundering: Effects Beyond the Named Command
- Agent Config as a Managed Supply Chain: Hashing and Pinning
- Agent Determinism Moves to the Last Unconstrained Axis
- Agent Failure Trajectories and the Recovery Window
- Agent JIT Compilation: Compile Tasks Into Executable Plans
- Agent-Generated Code Maintenance Asymmetry
- Agent-Initiated Rubric-Gated Self-Compaction (SelfCompact)
- Agent-Native Filesystems: Gating Effects, Not Commands
- Agent-Operable Interface Design (Affora)
- Agent-Reactive Bugs at the Model-Harness Boundary
- Agent-Ready Bug Reports for Software Repair Agents
- Agentic AI Architecture: From Prompt to Goal-Directed
- Agentic Detection and Response at the MCP Boundary
- Agentic Education: Persona Progression for Teaching AI Coding Tools
- Agentic Review Comment Acceptance
- Aggregation Bounds for Agent Authorization
- Ambiguity Stability as a Model-Selection Criterion
- An Explicit Update Boundary for Agent Self-State
- Answer-Reachable Eval Environments
- Artifact-Driven Workflow Compilation for Agent Execution
- Artifact-Level Accountability Mapping for Agent Workflows
- Artifact-Only Verification Hides Skipped Skill Steps
- Ask-Everything Permission Policies Protect Less than Per-Action Approval
- Assertion-Free Test Theater in Agent-Authored Patches
- Assuming Agent Interchangeability in Long-Running Teams
- Assuming Loaded Skills Stay Enforced in Long Contexts
- Assuming a CLAUDE.md Security Rule Is Enforced
- Attention Latch: When Agents Stay Anchored to Stale Instructions
- Attention Sinks: Why First Tokens Always Win
- Audit Your Test Suite With an Agent, Then Certify Each Flag
- Audit the Noise Floor Before Trusting a Benchmark Gap
- Audit-Budget Allocation for Agent Fleets
- Auditing Agent Tool Chains for Silent Partial Success
- Authority Confusion: Untrusted Context Must Not Authorize Side Effects
- Authorization Continuity Across Agent Mutation
- Baseline-Aware Test Evaluation for Multi-Agent Issue Resolution (Phoenix)
- Behavior-Partitioned Security Tests as Executable Specs
- Behavioral Drivers of Coding Agent Success and Failure
- Behavioral Firewall for Tool-Call Trajectories
- Behavioral Specification Elicitation Before Synthesis (SpecFirst)
- Belief Inertia After Tool-Map Drift in AI Agents
- Benchmark Contamination as Eval Risk
- Benchmark Poisoning of Self-Modifying Coding Agents
- Benchmark-Driven Tool Selection for Code Generation
- Benign Skill Wording That Steers Package Hallucination (Neutral Prompting Attack)
- Binding an Agent's Effect to the Approval It Claims
- Black-Box Agent Risk Scoring by Domain
- Blaming the Model for Scaffolding-Driven Quality Regressions
- Blind Resampling Over Self-Repair in Small Code Models
- Blind Tool Deference: Agents Parroting Callable Tools
- Bootstrapping Coding Agents: The Specification Is the Program
- Bounded Repair-Loop Iterations
- Bounded Tool Surfaces for Code Review Agents
- Budgeted Verification of Inherited Agent Constraints
- Bug-Discriminating Validation Evidence for Repair Agents (BSG-VA)
- Building Custom Agents from Substrate to Production (Agents All the Way Down)
- CARE: Three-Party Stage-Gated Engineering of LLM Agents
- CRA-Only Review and the Merge Rate Gap
- Cache-Safe Routing Boundaries: Where a Router May Act
- Calibrated Early Termination and Warm Restart for Agent Runs (FailFast-RestartSmart)
- Caller-Actionable Error Steps in Tool Responses
- Canary Tools for Diagnosing Tool-Selection Reasoning
- Catastrophic Remembering: Instruction Files That Only Grow
- CausalFlow: Counterfactual Repair for Failed Agent Trajectories
- Chain-of-Verification for Coding Agents
- Choosing a Compression Budget for Agent Control Context
- Choosing a Skill Loading Method for Agents
- Choosing an Agent Tool Interface: Shell or Typed Catalog
- Choosing the Judge Model That Grades Your Agent Evals
- Choosing the Right Surface for a Coding Agent Task
- Chunking Strategy for RAG-Based Code Completion
- Claim-Scoped Invalidation for Agent Memory
- Claim-to-Evidence Trace Graphs for Auditing Agent Runs
- Clarification Mode Amplifies Prompt Injection
- Classification Before Repair in an Analyzer Backlog
- Classifying and Auto-Correcting Coding Agent Misbehaviors (Wink)
- Closed-Loop Agent Training from Tool Schemas
- Closed-World Tool Call Resolution Before the Permission Gate
- CoALA Decision-Making Loop as an Orchestration Lens
- CoALA Memory Taxonomy as a Classifier for Harness Artifacts
- CoALA Structured Action Space: Internal vs External Actions
- Code Cleanliness as an Agent Cost Lever
- Code Health as a Signal for Agent-Generated Test Quality
- Code Injection Defense in Multi-Agent Pipelines
- Code-Native Memory Substrates for Coding Agents
- CodeSlop: Search-Trajectory Residue in Agent Patches
- Coding-Agent Misalignment Forms (Seven-Symptom Taxonomy)
- Coding-Agent Reversibility: Platform Choice as a Two-Way Door
- Coding-Agent Working-Set Coverage (Coherence Debt)
- Cognitive Poisoning: Untrusted Tool Feedback as a Trajectory Attack
- Comment Content as Code-Generation Context
- Comparison-Only Advisor: Steering a Large Actor With a Tiny Comparator
- Compiled Specialist Agents: Muscle Memory for Recurring Intent
- Completion Failure Taxonomy: Why Code Suggestions Miss
- Completion Summary as the Oversight Surface
- ComplexMCP: Three Bottlenecks in Large Interdependent Tool Sandboxes
- Component-Isolated Memory Stress Testing for LLM Agents
- Component-Wise RAG Prioritization for Software Engineering Tasks
- Compositional Skill Routing for Large Skill Libraries
- Compositional Vulnerability Induction in Coding Agents
- Compound Prompt Constraints Degrade More Than Their Parts Predict
- Computer-Systems Lens for Always-On Agent Security
- Concurrent Agent Pull Requests and Merge-Conflict Cost
- Configuration File Structure Does Not Drive Compliance
- Configuration Smells in AGENTS.md Files (Six-Smell Catalog)
- Constraint Decay in Backend Code Generation
- Constraint Degradation in AI Code Generation
- Constraint Drift: Why Safety Must Be Maintained, Not Asserted
- Constraint Encoding Does Not Fix Constraint Compliance
- Constraint Preambles and the Gain Your Scanner Misses
- Constraint Tax: Tool Suppression Under JSON Schema Decoding
- Constraint-Evasive Fabrication in Instruction Sets
- Constraints as a Substrate for Scalable Agent Oversight
- Content-Addressed Agent Configurations (Deterministic Control Plane)
- Context Compiler: Deterministic Assembly Over Bigger Windows
- Context Lifecycle Management: Beyond Store and Retrieve
- Context Priming: Pre-Loading Files for AI Agent Tasks
- Context Quality as a Leading Indicator of Agent Reliability
- Context-Fractured Decomposition Attacks on Tool-Using Agents
- Contextual Capability Calibration for Multi-Agent Delegation
- Continuation Authority in Agent Migration
- Contract-Domain Tracing for Rubric Credit
- Contractual Skill Files: Inspectable SKILL.md for Enterprise Agents
- Control Lexical Leakage in Agent-Memory Retrieval Evals (Entity-Collision)
- Control/Data-Flow Separation for Prompt Injection Defense (CaMeL)
- Controlled Benchmark Rewriting for Agent Safety Judgment
- Coordination Channel Policy for Multi-Agent Coding
- Cost-Aware Skill Rewriting: Preserve Operational Anchors, Not Skill Tokens
- Cost-Aware Tracing for Skill Distillation
- Cost-Inefficient Behaviors in Coding Agents
- Coverage-Aware Skill Selection Under a Token Budget
- Coverage-Guided Fuzzing for Multi-Agent LLM Systems (FLARE)
- Covert Success Rate for Indirect Prompt Injection
- Critical Instruction Repetition via Primacy and Recency
- Criticality and Containment: Scoping Which Agent Code You Read
- Cross-Component Interference in Agent Scaffolds
- Cross-Framework Signal Semantics: Re-Measure Borrowed Trajectory Rules
- Cross-Iteration Safety State for Agent Loops (LoopHarness)
- Cross-Layer Evidence for Agent Attack Detection
- Cross-Lingual Prompt Preprocessing (Local-LLM Token Arbitrage)
- Cross-Session Re-Implementation of Existing Agent Code
- Cue-Anchored Working Memory (Delivery, Not Storage)
- Cumulative-Best Reporting Hides Repair-Loop Security Regressions
- Cut-Point Replay: Test a Fix Against a Recorded Agent Run
- Debugging the Tool-Call Loop Before Reaching for a Framework
- Decentralized Memory for Self-Evolving Multi-Agent Systems
- Decision-Fork Replay: Grading an Agent's Mid-Run Choices
- Declared Peer Consensus as Context for a Reviewing Agent
- Decomposed Red-Teaming for Agent Monitors
- Decomposing Agent Output Variability by Layer (Sampling vs Orchestration State)
- Decoupled Search Grounding: A Vendor-Agnostic Grounding Boundary
- Defense-in-Depth Against Coding Agent Fabrication (Honesty Harness)
- Defense-in-Depth Agent Safety for AI Agent Development
- Delegated-Autonomy Boundary Artifacts (AJR and ADP)
- Delegating Multi-Hunk Bug Repair to Coding Agents
- Deletion Avoidance: Agents That Guard Code Instead of Removing It
- Deletion and Cost Rules for an Evolving Agent Harness
- Deliberation-Inducing Cues That Multiply Reasoning Cost
- Delivery-Bound Tool Authorization: When Progressive Discovery Becomes Access Control
- Density-Normalized Quality Metrics Mask AI-Driven Code Growth
- Dependency Gap Validation for AI-Generated Code
- Deriving a Specification From Buggy Code Before Generating Tests
- Design Docs as the Durable Artifact
- Designing Agents to Resist Prompt Injection
- Destructive-Failure Mechanism Attribution by Mitigation Owner (ClayBuddy Three)
- Destyling Untrusted Input as a Prompt Injection Defense
- Detecting Memory-Poisoning Exfiltration by Tool-Call Order (Recall-Before-Send Signature)
- Detecting Self-Preference in a Single LLM Judge
- Deterministic Anchoring: Static Facts as Stable Context
- Deterministic Fast Paths: Answer Without a Model Call
- Deterministic Precondition Gates for Tool-Using Agents
- Developer Control Strategies for AI Coding Agents
- Diagram as the Shared Spec: One Artifact for the Picture and the Prompt
- Diff-Coverage Gating for Agent-Authored Pull Requests
- Difficulty-Aware Topology Selection for Coding Agents
- Discovery-Only Refactor Pass: Surface Candidates Before Touching Code
- Dispatch-Time Reasoning Level for Delegated Agents
- Distillation-Induced Similarity Metrics for Tool-Use Agents
- Distilled Bootstrap Contract: Agent-Authored Repo Setup
- Distractor Interference: Why Relevance Is Not Enough
- Distributed Cross-PR Attacks in Persistent-State AI Control
- Distributing Security Controls Through the Agent Harness
- Do Not Price the Rules in Your Agent Instruction File
- Documentation Read Counts Measure Retrievability, Not Value
- Documentation-Guided Legacy Migration: Architecture Docs as a C-to-Rust Blueprint
- Documenting Code the Agent Can Already Read
- Domain-Scoped Parallel Exploration for Multi-File Change Localization
- Dormant Memory Payloads Triggered by Sensitive Topics (Trojan Hippo)
- Dual Executable Specifications for Long-Horizon Features
- Dual-Budget Control for Search Agents: VOI Scoring Per Action
- Dual-Graph Alignment for Indirect Prompt Injection Defense (AuthGraph)
- Dual-Trace Memory Encoding: Pair Facts with the Scene They Were Learned In
- Effective Feedback Compute (EFC) for Harness Comparison
- Elastic Context Orchestration: A Per-Turn Vocabulary for Long-Horizon Search Agents
- Encoding Values in AGENTS.md: Why Prose Without Verification Fails
- Enforced Versus Advisory Controls in LLM-Native IDEs
- Enforcement Modes in Spec-First Agent Frameworks
- Enterprise Agent Hardening: Three Production Gates
- Entity Binding Failures in Tool-Augmented Agents
- Enumerate Every Visible Exit Before Scoring Containment
- Environment Specification as Context: Closing the Version Gap
- Episodic Memory Retrieval for AI Coding Agent Loops
- Epistemic Working Memory for Multi-Hop Reasoning (SLEUTH)
- Equivalence Testing for Agent Configuration Changes
- Escalation Channels: A Reporting Tool Instead of a Reward Hack
- Eval Blind Spots: Structural Gaps in Measurement Methodology
- Event-Driven System Reminders for AI Agent Development
- Evidence-Bundled Agent PRs: Sizing the Reviewer's Effort
- Evidence-Conditioned Execution: Gate Edits on Observations
- Evidence-First Reports From Failure-Diagnosis Agents
- Evidence-Gated Lifecycle Control for Coding Agents (Proof-or-Stop)
- Evidence-Grounded Disagreement in Agentic Code Review (Adversarial Review)
- Evolving Playbooks: Incremental Context That Preserves Knowledge
- Exactly-Once Enforcement Layer: Model, Harness or Contract
- Example-Driven vs Rule-Driven Instructions
- Executable Memory: User State as Code for Personalized Agents
- Execution Budgeting in Agentic Program Repair
- Execution-Layer Security Invariants for MCP Runtimes
- Execution-State Ledger for Long-Horizon Coding Agents
- Experience Graphs as Structured Memory for Self-Evolving Agents
- Experiential-Learning Setup Agents with Snapshot Rollback (SetupX)
- Explained Feedback for LLM Vulnerability Repair
- Explanation-Bound Tool Execution for Agent Gateways
- External Artifacts Treated as Data, Not Adversarial Input
- Externalization in LLM Agents
- Fact Supersession Memory for Code Assistants
- Failure-Aware Observability for Multi-Agent LLM Systems
- False-Pass Liability Decides the Next Harness Component
- Feedback as Capability Equalizer: Iterative Feedback Outweighs Model Scale
- Field-Change Intent Instead of Model-Written Diffs
- Field-Level Source Ownership for Agent Capabilities
- First-Proposal Execution in Agent Loops
- Five-Pass Blunder Hunt: Repeated Critique Passes for Plans
- Flattened Tool Specs for Agent Safety Judgment (SafeKeep)
- Fleet-Level Irreversibility Budgets for Agent Effects
- Foresight-Guided Defense Against Infectious Jailbreaks in Multi-Agent Systems
- Forged Reasoning Trace Attacks on Agent Memory (FARMA)
- Four Reporting Levels for Agent Working Memory Evaluation
- Four-Phase Agent Delegation with Curated Artifacts
- Framework-First Agent Development: An AI Anti-Pattern
- From Preventive to Reactive: Front-Loading Security in AI Coding Prompts
- Frontmatter and Body Rule Drift in Agentic Workflows
- Frozen Playbook Reuse Without Target-Side Validation
- Frozen Task Sets for Affordable Agent A/B Testing
- Frozen-Base Task Mining for Repository Instruction Files
- Frozen-Stimulus Panels for Cross-Vendor Behavior Measurement
- Function-Level Debugger Interfaces for Coding Agents
- GEO for Technical Docs: Developer Documentation Checklist
- Gate Best-of-k Selection on Compliance Before Score
- Gate Generation on Retrieval Sufficiency, Not Model Confidence
- Generated Programs as Web Agent Action Space
- Generating Tests From Agent-Written Code (Code-First Oracle Bias)
- Generative Provenance Records for Tool-Using Agents
- Git-Bound Memory for the Agentic Development Lifecycle
- Give the Model the Target's Contract, Not Similar Solutions
- Goal Reframing: The Primary Exploitation Trigger for LLM Agents
- Governance Layer for Agent Interoperability Protocols
- Handoff-Boundary Fault Injection (llmmas-otel)
- Harness Design Dimensions and Archetypes
- Harness-Controlled Token Economics (The Harness Effect)
- Heartbeat-Bound Hierarchical Credentials for Agent Swarms
- Held-Out Tasks as a Harness Shortcut Defense
- History Anchors: Consistency-Cued Continuation of Unsafe Prior Actions
- Homogeneous Debate Panels as a Groundedness Quality Lever
- How Teams Build SE Agents: A Seven-Stage Build Loop
- Human-AI Review Synergy in Agentic Code Review
- Hybrid Deterministic + Semantic Authorization for Agent Tool Calls
- Idle-Time Speculative Planning for ReAct Agents
- Independent Test Generation in Multi-Agent Code Systems
- Inference-Time Tool-Call Reviewer: Pre-Execution Feedback for Tool-Calling Agents
- Informed Abstention as a Tool-Boundary Runtime Gate
- Injected-Context Cost Attribution in Agent Workflows
- Inline Safety Harness with Cascade Verification (FinHarness)
- Inline Suggestion Attachment in Agent Code Review
- Install-Once Plugin Trust: Vetting That Never Re-Runs on Update
- Instruction-Guided Code Completion: Controlling What Models Generate
- Intent-Centric Engineering: Oversight Over Authorship
- Intent-Governed Tool Authorization for AI Agents (IGAC)
- Interaction-Pattern Evaluation for Agentic PRs
- Interactive Clarification for Underspecified Tasks
- Issue Requirements Preprocessing: Structured Input Before Code Generation
- Iterative Binary Feedback for Pattern Adherence
- Judging Agent Safety by Task Completion (Action-Boundary Violations)
- Judging a Skill's Honesty by the Validity of Its Output
- Judgment Relocation: Where Human Decisions Land in an Agent Factory
- Kaizen-Style Continuous Code Quality Loop (Pomona)
- Knowledge Cutoff as a Documentation Boundary
- Knowledge Gap or Skill Gap: Triage Before Writing Context
- LLM API Fault Injection at the HTTP Layer (AgentChaos)
- LLM Agent Bug Fix Taxonomy: 23 Fix Patterns from 930 Real Bugs
- LLM Code Review Overcorrection for AI Agent Development
- LLM Refactoring Adoption Patterns
- LLM Self-Review Failure in Code Modernization Tasks
- LLM Static Verification Against Natural-Language Requirements
- LLM Support During the First Detection Pass
- LLM-Driven Benchmark Auditing
- LLM-Driven Logical Retrieval: Boolean Queries over an Inverted Index
- LLM-Pinned Library Versions Carry Systemic CVE Exposure
- LLM-as-Code Agentic Programming for Agent Harnesses
- Language Choice as an Agent Token-Cost Lever
- Late Requirement Arrival in Agent Sessions
- Layered Mutability: Governing Persistent Self-Modifying Agents
- Layered Oracle Stack for Agent IaC Security Repair (TerraProbe)
- Learned Prefix Monitors for Agent Traces
- Learning Execution Guardrails from Agent Failure Traces
- Lexical-First Retrieval for Agentic Search: When BM25 Is Enough
- Lifecycle-Integrated Security Architecture for Agent Harnesses
- Line-Anchored Feedback: Deliver Change Requests as Inline Comments
- Lost in the Middle: The U-Shaped Attention Curve
- MCP Approval-View Fidelity Gap and Unicode Concealment
- MCP-vs-CLI Cost Ratios Are a Property of the Scaffolding
- Marking Which Artifacts Are for Humans or Agents (Landmarking)
- Mask Tools Instead of Removing Them
- Match Architecture Spec Format to Model Capability
- Measure the Judge Before You Freeze a Gate on It
- Measuring Reacquisition Cost Under Context Compaction
- Measuring Synthetic Eval Data Quality (SynAE)
- Measuring the Verification Tax on Agent Output
- Memory Retrieval as a Control Decision
- Memory Transfer Learning: Cross-Domain Memory Reuse in Coding Agents
- Memory-Induced Tool-Drift in LLM Agents
- Mid-Trajectory Guardrail Selection for Multi-Step Tool Calls
- Minimality Prompts as a Patch-Size Control
- Minimum-Cost Evidence Selection for Agent Changes (Assurance Envelopes)
- Minimum-Sufficient Execution: Estimate Scope Before Spending Budget
- Mise en Place for Agentic Coding
- Model Confidence as Security Verification (Security Calibration Gap)
- Monitor or Wait: The Supervision Choice During Agent Execution
- Monotonic Capability Attenuation for Composition-Safe Tool Use
- Multi-Agent RAG for Spec-to-Test Automation
- Multi-Agent SE Design Patterns: A Taxonomy Across 94 Papers
- Multi-Agent Shared State Isolation Anomalies
- Multi-Layer Specification Redundancy as a Robustness Budget
- Multi-Run, Shuffled-Order Evaluation for Self-Improving Agents
- Multi-Tool Threshold Poisoning Against MCP (ShareLock)
- Multitenant RAG: Closing the Relevance-Authorization Gap
- Mutation Testing as a Quality Gate for AI-Generated Test Suites
- Mutation Testing for LLM Judges: Scoring an Evaluator on Injected Defects
- Name the Check That Passed Before Accepting AI Code
- Narrative Problem Reformulation for Code Generation
- Natural Language Tool Selection (NLT)
- Natural-Language Documentation as a Code-Review Intermediate (Verifiable Literate Programming)
- Next Edit Suggestions Carry Context You Never Curated
- Non-Compensatory Readiness Gates Before Agent Release
- Nonstandard Errors in AI Agents: Model-Family Variance
- Observability-Driven Harness Evolution
- Observation Contract Preservation in Tool-Augmented Agents
- Observation-Driven Coordination: CRDT-Based Parallel Agent
- Offline Trajectory Replay for Multi-Agent Workflow Debugging
- One-Shot Record and Deterministic Replay for Periodic Agent Tasks
- OpenAPI Documentation Smells for Agent-Ready APIs
- Organizing Filesystem Agent Memory for Retrieval Cost
- Outcome Monitors: Recovery Affordances for Tool Failures
- Overeager-Behavior Elicitation: Scope + Trap Fragments as a Diagnostic for Out-of-Scope Tool Calls
- Overtrusting Human Sign-Off on Generated Assertions
- PASS@(k,T): Evaluate RL for Agents Along Sampling and Interaction Depth
- PEEK: Orientation Cache for Recurring-Context Agents
- PR Description Style as a Lever for Agent PR Merge Rates
- Parallel Polyglot Ports as a Spec-Ambiguity Oracle
- Parameter-Keyed Caching and Dependency-Aware Parallelism for Plan-Execute Pipelines
- Path-Scoped Write Contracts for Shared Agent State
- Peer Refusal as a Coordination Control
- Per-Agent Capability Stores Beat One Task-Wide Allowlist
- Per-Attempt Sandboxes for Agents That Change the Filesystem
- Per-Change Deploy Monitors: Report the Verdict, Don't Act on It
- Per-Layer Suppression Accounting in Acceptance Gates
- Per-Line Requirement Citations for Hallucination Detection
- Per-Object Context Allocation (Selective Invariance)
- Per-Reviewer Context Views for Code Review Agents
- Per-Step Preconditions and Postconditions in Skill Files
- Per-Type Retention Policy for Agent Compaction (Knowledge Triage)
- Permission Framework Choice Outweighs Model Choice for Limiting Overeager Actions
- Persistent Shared Search Sub-Agent for Output-Token Reuse
- Persistent Teammate Workspace: Durable State for Agent Teams
- Personalized vs Generic Agent Skills: Where Effort Pays
- Phantom Symbol Detection for LLM API Migration
- Plan Compliance in Agents: Measure What They Execute, Not What You Wrote
- Plan-Then-Execute as the Default for Web Agents
- Planning Stage Preconditions: Budget Headroom and Task Text
- Plugin Component Co-Change: Scripts and Their Instructions Move Together
- Policy File Validation: Catching Silent Non-Enforcement
- Policy-Graded Evaluation of Coding Agents
- Pooled-Evidence Factuality Checks for MCP Agents (Cross-Source Conflation)
- Post-Merge Fix Signals for Agent Merges
- Pre-Execution Codebase Exploration for AI Coding Agents
- Pre-Execution Failure Scoring with a Draft Model (Speculative Uncertainty)
- Pre-Generation Complexity Scoring for Code Reliability
- Pre-Write Change Intent Admission (Claim Plane)
- Precise Debugging: Measure Edit Precision, Not Just Test Pass Rate
- Predicting Reviewable Code: Pre-Flagging Functions Reviewers Will Delete
- Preempting Agentic PR Rejection by Failure-Mode Category
- Pressuring a Coding Agent Degrades the Code It Writes
- Prior Dominance Over Feedback in Agent Optimization Loops
- Privacy-Preserving LLM Requests: Eight Techniques and a Practical Combination
- Proactive Idle-Time Anticipation (ProAct)
- Probe-Run Calibration for Predicting Agent Token Spend
- Probe-and-Refine Tuning of Repository Guidance for Coding Agents
- Probing Unstated Constraints in Generated Code (Intent Violation Rate)
- Profile Your Agent Test Suite Against Measured Practice
- Profiler-Guided Optimization Loops for Coding Agents
- Programming Language Choice Still Shapes Agent Artifacts
- Prompt Cache Keepalive for Agent Pauses
- Prompt as Security Knob
- Prompt-Only Baseline Before a Specialized Agent Subsystem
- Prompted Uncertainty Decomposition for Clarification Routing
- Proof of Presence: Re-Authenticating for Agent Actions
- Proprioceptive Context Dashboard: Agent Self-Managed Context
- Provenance-Aware Decision Auditing for LLM Agents
- Public Rules-File Corpora as Evidence
- Purpose-Built Eval Suites for Model and Harness Swaps
- Query-Conditioned Reuse of Retrieved Agent Trajectories
- RAG Architecture as a Poisoning Robustness Decision
- RAG over Thinking Traces: Index Reasoning Trajectories Instead of Documents
- RAMP: Committed AI Configuration and the Quality Cost
- Rank Resolution: Reading a Converged Coding-Agent Leaderboard
- Re-Run an Agent's Speed-Up Claim Before Merging
- Re-Run the Original Test Suite After Every Refinement Turn
- ReAct (Reason + Act): Interleaved Reasoning-Action Loops
- Reading Visible Edge-Case Handling as a Security Check
- Reading a Coding-Agent Vendor's Security Certificate
- Reasoning Effort Over Tool Scaffolding for First-Try Reliability
- Recover the Six Measurement Choices Behind an Attack Success Rate
- Recoverability-Gated Self-Evolution for Agent Harnesses
- Recurring Control Belongs in Harness Code, Not Context
- Recursive Agent Harnesses (RAH)
- Recursive Best-of-N Delegation
- Red-Team Your Blocking Monitor Before You Trust It
- Refactoring Runaway: Tangled Refactorings in Agent Patches
- Reframed Exfiltration Defeats Wording-Based Defenses (Framing Gap)
- Reliability of an Automatically Selected Agent Harness
- Repairing Agent Prompts from Trace Contrast, Not Search
- Replayable Encrypted Reasoning Blocks in Agent Traces
- Repository Perturbation as Context-Reasoning Diagnosis (RepoMirage)
- Repository Skill Release Drift
- Reproducibility Artifacts as Agent Context
- Request Shaping to Cut Wasted Agent Turns
- Requirement Smells: No Category Signal, Density Unconfirmed
- Rerunnable Claim Graph as Shared Agent Memory
- Residual Completion for Stateful Agent Handoffs (CFRC)
- Restraint Rules Need External Enforcement
- Restricting a Coding Agent to a Single execute_code Tool
- Retry-Switch-Abstain: A Runtime Tool-Recovery Policy
- Reverse-Engineered Executable Specifications for Agentic Program Repair
- Review Constraint Tests as a Second Acceptance Gate
- Reviewer Habituation in Agent PR Review
- Reviewer Precision as a Pipeline Quality Proxy
- Reviewer Theme Distribution Audit for AI Code Review
- Revocable Resource-and-Effect Capabilities for Coding Agents (PORTICO)
- Risk Architecture for AI-Native Engineering Teams
- Risk-Score Threshold Calibration for Auto-Approval
- Role Orchestration on a Single Model
- Rolling Out CLI Coding Agents at Organization Scale
- Rolling Out a Team-Embedded Agent Like a Tool
- Root Causes of Vibe-Coded Application Vulnerabilities
- Route Agent Peers by Enrolled Identity, Not Card Name
- Route-Parity Auditing of Agent Safeguards
- Routing Dependency Updates to Repair Agents by Budget
- RubricRefine: Pre-Execution Rubric Refinement for Code-Mode Tool Use
- Runtime Guard as an Installed Skill (Defense-as-Skill)
- Runtime Resource Limits as Prompt Context
- Runtime Scaffold Evolution: Agents That Build Tools
- SUDP: Secret-Use Delegation Protocol for Agentic Systems
- Schema-Guided Graph Retrieval
- Scope-Matched Retrieval for Persisted Agent Skills
- Scoped-Looking Permission Grants
- Scoring Constraint Loss and Tool Reach as Separate Risks
- Scoring a Compaction Policy on Latency and Billed Cost
- Scout-Then-Route: Verify the Handoff Before Routing
- Security-Aware Tool Descriptions for MCP Servers (SpellSmith)
- Selective Autonomy from Copilot Feedback
- Selective Revalidation for Pending Agent Actions
- Self-Discover Reasoning: LLM-Composed Reasoning Structures
- Semantic Collapse Under Underspecified Prompts
- Semantic Density Optimization for Agent Codebases
- Serving-Stack Confounds in Tool-Call Evaluation
- Severity-Stratified Evaluation of Security Prompts
- Shallow Agent Test Coverage from Premature Termination (Lazy Generation)
- Shortening Old Tool Results Under Context Pressure (Half-Life Truncation)
- Silent Adoption of Corrupted Tool Returns by Agents
- Silent Handoff Failure in Delegated Code Search
- Silent-Failure Mechanism Taxonomy in Production Agent Runtimes
- Single-Decision Approval of Vendor Skill Suites
- Size Agent Comparisons by Run-to-Run Variance
- Skill Authoring as Software Engineering: What Transfers
- Skill Composition Risk in Agent Ecosystems
- Skill File Linting: Which Three Checks to Run First
- Skill Library Technical Debt: Library-Time Maintenance for Agent Skills
- Skill Lift: Measuring What a Skill Adds at Runtime
- Skill Loadout Curation for Coding Agents
- Skill Misevolution in Self-Updating Skill Libraries
- Skill Over-Trust: Treating Topical Relevance as Evidence a Skill Helps
- Skill Packs: Registry Distribution Needs Pinning Discipline
- Skill Program Functions: Executable Guardrails Compiled From Past Failures
- Skill Reuse as Vendored Forking
- Skill Review Without a Token Cost Baseline
- Skill Specification Violation Fuzzing
- Skill Test Coverage as a Release Gate
- Skill-Use Gates: Trigger, Compliance and Boundary
- Solver-Externalized Constraint Reasoning (MaxSAT/SMT Encoding)
- Source Code Minification for State-in-Context Agents
- Spec-Anchored Drift-Gated Architecture (Spec Growth Engine)
- Spec-Derived Execution as a Correctness Oracle
- Spec-Driven Test Generation: Contract Coverage Is the Lever
- Specialist Orchestrated Queuing for Multi-Agent SE (SPOQ)
- Specification Authority Boundary: Agents Propose, the Runtime Commits
- Specification Memory: What a Shared Agent Workspace Keeps
- Specification Portability Across Coding Agents
- Specification-First Convergence Without a Test Oracle
- Specification-Grounded Test Writing
- Specification-Path Testing: Same Contract, Different History
- Splitting an Agent Token Budget at the Scaling Inflection Point
- Splitting the Drift Judge from the Advisor (LivePlan)
- Stage Elision Before Summarization
- Stage-Targeted Prompt Structure for Pull Request Outcomes
- Staged Literal Porting with a Per-Stage Numeric Oracle
- Stale AI Configuration Artifacts (Context Rot)
- Standard-Grounded NFR Specs: Quality Up, Correctness Flat
- State-Bound Evidence and Typed Revision Contracts for Repair Loops
- State-Conditioned Evidence Selection for Mid-Task Retrieval
- Stateful Agent Evals via State Snapshots and Transition Assertions
- Stateful Iteration State-Carry: Typed Persistent State for Long Agent Loops
- Static Difficulty Estimation for Agent Issue Triage
- Static-First Shell Command Gating with Selective Escalation
- Step Budgets and Trust in Agent-Generated Code Tours
- Strained Coherence as a Pre-Failure Signal in Agent Trajectories
- Strategy Over Code Generation: Why AI Speed Doesn't Fix Wrong Goals
- Structural Coverage Criteria for Agent Workflows
- Structural Monitoring for Covert Safeguard-Weakening
- Structure-Aware Diff Labeling with Two-Stage LLM Pipelines
- Structured Agentic Software Engineering (SASE)
- Structured Domain Retrieval: Knowledge Graphs and Case-Based Reasoning
- Structured Task-State Ledger for Tool-Calling Agents (LedgerAgent)
- Subagent vs In-Context Skill Execution
- Subtask-Level Memory for Software Engineering Agents
- Sufficiency-Tightness Decomposition for Agent-Authored Permissions
- Suggestion Gating: Fewer Completions, Better DX
- Supply-Chain Security Debt in Agent Pull Requests
- Suspect the Harness Before the Model on a Regression
- Symptom-First Bug Triage for Agent Code
- Symptom-Reduction-as-Root-Cause: Why Oracle Tests Alone Miss Architectural Drift
- TDD Interaction Models: Throughput Versus Test Quality
- Tab-Accept Rate as a Proxy for Critical Engagement
- Task Alignment: The Selective-Compliance Gap Benchmarks Miss
- Task Category as the Security Review Routing Key
- Task Completion as Tool Certification (Silent Tool Rot)
- Task Feasibility Awareness: Stop Before You Start
- Task Shape Decides What a Heavier Agent Harness Buys
- Task-Based Access Control with Hybrid Inspection
- Task-Uniform Agent Permissions Ignore Where Failures Land
- Terminal-First Agent Interfaces with Browser Escalation
- Test Oracles That Read Their Expectation From the Code
- Test-Driven Intent Clarification: Tests as Intermediate Alignment Artifacts
- The Agent-on-Agent Maintenance Penalty
- The Compliance Trap: Consuming Conflicting Agent Memory
- The Decoding Harness Is Part of the Agent Attack Surface
- The Error-Class Governance Loop for Instruction Libraries
- The First Edit Predicts Whether an AI Completion Survives
- The Handoff Tax: What a Receiving Model Should Inherit
- The Instruction Compliance Ceiling: How Rule Count Limits AI
- The No-Op Test: Prune Agent Docs by Behavior, Not Length
- The Patchwork Problem in LLM-Generated Code
- The Post-Authorization Execution Trust Gap in Remote MCP
- The Productivity-Experience Paradox in AI-Assisted Development
- The Reasoning-Complexity Trade-off
- The Recall Trap: Tuning a Code Retriever on Recall@k at a Fixed Slot Budget
- The Security Review Gap in AI-Authored PRs
- The Skill Closure Declaration Gap
- The Test Homogenization Trap: When LLM-Generated Tests Mirror Model Blind Spots
- Three-Vector Evasion Taxonomy for Agent Security Tests
- Tiered Memory Architecture: Episodic-to-Semantic Consolidation for Long-Running Agents
- Timeout Oracles for Agent-Written Code
- Token Reduction Mistaken for Cost Reduction
- Token-Efficient Code Generation: Structural Beats Prompting
- Tool Architecture Moves Consistency, Not Resolve Rate
- Tool Cloning and Provenance Assessment in Agent Ecosystems
- Tool Necessity Probing: Reading Tool-Call Decisions From Hidden States
- Tool Operability: Interfaces That Survive a Lost Response
- Tool-Call Success as Workflow Effect Evidence
- Tool-Invocation Attack Surface in Coding Agents
- Tool-Schema Bias: Measure the Interface Before You Ship
- Tool-Use Sim-to-Real Perturbation Taxonomy
- Tools as Typed Code Stubs (Programmatic Tool Calling)
- Trajectory Attribution for Context Repair (TRACE)
- Trajectory Decomposition: Diagnose Where Coding Agents Fail
- Trajectory Poisoning of Promoted Agent Skills (PoisonedEvolution)
- Trajectory Pre-Filter for Failure Diagnosis (TrajAudit)
- Trajectory Projection: A Flattened View of Agent Traces
- Trajectory-Aware Benchmark Subset Selection for Agents
- Trajectory-Conditioned Model Escalation (SWE-Router)
- Transcript-Measured Review Coverage
- Treating Agent Delegation as Routing, Not Authorization
- Treating Agent Safety as Uniform Across a Session (Cold-Start Safety Gap)
- Treating File Secrecy as Skill Confidentiality
- Treating Memory-Injection Rate as Security Evidence
- Treating Read Denial as Confidentiality for a Build Input
- Treating a Clean Final State as Boundary-Compliance Evidence
- Treating a Clean Merge as Compatibility Evidence
- Treating a Clean Static Scan as Security Evidence
- Treating a Local Agent Session Trace as Audit Evidence
- Trusting Claimed Prior Approval in Agent Review Gates
- Trusting Model-Level Privilege Restraint at Tool Selection
- Trusting Tool Error Messages as Implicit Authority (Error-Path Injection)
- Typed Context Buys Addressability, Not Token Savings
- Typed Generation Contracts for Grounded Extraction
- Typed Memory Provenance and Assertion Release Gating
- Typed Pseudocode for Skill Libraries (Skill-as-Pseudocode)
- Unsignalled Tool Failure: Returning Success With an Unusable Payload
- Unstated-Contract Bugs: Sort Tickets by Information Gap
- Usability Pressure as a Silent Security-Regression Vector
- Validating Token-Optimized Formats Inside Agentic Loops
- Validity-Estimate Stopping for Noisy Verify-Repair Loops (VRR-Stop)
- Verbatim Failure Records in Small-Model Agent Transcripts
- Verification Surface: Match the Tool to the Failure
- Verifier-Driven Parallel Coding Agents (Glite ARF)
- Verify Agent Diagnoses and Fix Proposals Before Acting
- Verify Observability in Agent-Generated Code
- Verify-Gated Completion as Admission Control
- Version-Controlled Agent Context (Git Context Controller)
- Weakest Consistent Learning: What Agent Loops Should Persist
- What is GEO — Generative Engine Optimization Defined
- When a Skill Graph Cannot Beat the Ranker (Pre-Filter Topology Bound)
- Which Task You Delegate Changes Poisoned-Repo Exposure
- Workspace Topology as an Indirect Injection Attack Vector
- Workspace-Hosted Skills: Authorship Outside the Repo
- Write Agent Rules You Can Grade From the Transcript
- Zero Violations Is Not Evidence Your Hook Works
automation¶
- Bounding a Headless Codex Run Without a Turn Cap
- Cross-Repository Security Posture for Agent-Introduced Vulnerabilities
- Cursor Automations: Event-Triggered Agents and /automate
- Deny-Fallback Permissions for Unattended Agent Runs
- Field-Change Intent Instead of Model-Written Diffs
- GitHub Copilot Advanced Patterns: Multi-Agent and Automation
- Google Search Console Monitoring Workflow
- Hash-Pinned Command Approval for Non-Interactive Plugin Installs
- Headless Claude in CI: Using -p and --max-turns for Safe Pipeline Integration
- Issue-to-PR Delegation Pipeline for AI Agent Development
- Multi-Repo and No-Repo Coding Agent Automation Templates
- Non-Human Event Provenance Markers to Block Fabricated Approvals
- Programmatic Cloud-Agent Dispatch via REST API and Webhooks
- Trigger-to-Function Architecture for Unprompted Codebase Maintenance
claude¶
- @import Composition Pattern for Agent Instruction Files
- Administrative Effort Ceilings for Reasoning Budget
- Advanced Tool Use: Scaling Agent Tool Libraries
- Agent Observability with OpenTelemetry and Trajectory Logging
- Agent Project State Purge: Clean-Slate Session Reset
- Agent Time Estimates Are Not Schedules
- Agent View: Dispatch-Attach-Monitor Surface for Parallel Sessions
- Agent-Driven Deployment: What to Delegate and What to Gate
- Agent-Generated Onboarding Guide as a Durable Artifact
- Assuming a CLAUDE.md Security Rule Is Enforced
- CLAUDE.md Convention for Structuring Agent Instructions
- Centrally Provisioned MCP Servers: Remote Transports Only
- Channels Permission Relay
- Claude Agent SDK: Building Custom Agentic Workflows
- Claude Code --bare Flag
- Claude Code /batch and Worktrees for AI Agent Development
- Claude Code Agent Teams for Collaborative AI Workflows
- Claude Code Auto Mode: Classifier-Based Permission Gating
- Claude Code Dynamic Workflows
- Claude Code Extension Points: When to Use What
- Claude Code Feature Flags and Environment Variables
- Claude Code Hooks: Deterministic Lifecycle Automation
- Claude Code Review
- Claude Code Sub-Agents for Delegating Complex Tasks
- Claude Code for AI Agent Development
- Cloud Parallel Review Pattern
- Cloud Planning with Inline-Comment Review and Execute-Anywhere Choice
- Cloud-Scheduled Routines vs Local Session Scheduling
- Conditional Hook Execution: Filter Hooks by Tool Pattern
- Context-Window Diagnostic Tooling: Identifying Context-Heavy Tools
- Cross-Session Peer Messaging with a Posture-Keyed Inbox Gate
- Deferred Permission Pattern: Headless Agent Session Pausing
- Deny-Fallback Permissions for Unattended Agent Runs
- Directory-Aware Plugin Suggestions via `pluginSuggestionMarketplaces`
- Disable Attribution Headers to Preserve KV Cache in Local Inference
- Effort-Aware Hooks: Reading the Reasoning Tier from PreToolUse and PostToolUse
- Enforcing Agent Behavior with Hooks
- Enterprise-Managed Plugin Governance for Agent CLIs
- Evidence-Based Allowlist Auto-Discovery for Agents
- Exclude Dynamic System Prompt Sections for Cross-Machine Cache Sharing
- Fail-Closed Remote Settings Enforcement for Enterprise Agents
- Filesystem-Based Tool Discovery for AI Agent Development
- Gateway Hint Headers for Routing and Budgeting Agent Calls
- Gateway Model Routing: Treat the LLM Gateway as a Discovery Source
- Generated Questionnaires: Eliciting Someone Else's Context
- Handoff Skill: Structured Context Transfer Between Agent Sessions
- Hard-Deny Classifier Rule: Unconditional Block in Auto Mode
- Hash-Pinned Command Approval for Non-Interactive Plugin Installs
- Headless Claude in CI: Using -p and --max-turns for Safe Pipeline Integration
- Hierarchical CLAUDE.md: Structuring Context Files at Multiple Levels
- Hook Catalog for Claude Code Enforcement
- Hook Exec Form vs Shell Form: Shell-Injection-Safe Hook Commands
- Hooks Invoking MCP Tools: Closing the Loop Between Policy and Tool Execution
- In-Session Transcript Search: Navigating Long Agent Conversations
- Lazy Worktree Isolation: Enter the Worktree on First Write, Not on Dispatch
- Local Plugin Scaffolding via `claude plugin init` and Auto-Loaded `.claude/skills`
- MCP Elicitation: Servers Requesting Structured Input Mid-Task
- MCP Tool Result Persistence via _meta Annotation
- Managed Settings Drop-In Directory: Enterprise Policy Fragmentation
- MessageDisplay Hook: Transforming Assistant Text at the Display Boundary
- Model-Switch Lifecycle Hooks: Gating a Mid-Session Model Change
- Monitor Tool: Event Streaming from Background Scripts
- Multi-Tenant Isolation Knobs for Shared-Container Agent SDK Hosting
- On-Demand Skill Hooks: Session-Scoped Guardrails via Skill Invocation
- Out-of-Band Hook Notifications via terminalSequence
- Parameter-Level Permission Rules (Tool(param:value) Syntax)
- Per-Plugin Token-Cost Attribution via claude plugin details
- Per-Subagent Instruction Inheritance (omitClaudeMd)
- Permission Gates That Deny the Agent's Own Cleanup (Denied Remediation Path)
- Persistent Teammate Workspace: Durable State for Agent Teams
- Plan Mode: Read-Only Exploration Before Implementation
- Plan mode for knowledge artifacts
- Plugin Background Monitors: Declarative Supervision Auto-Armed at Session Start
- Plugin Component Co-Change: Scripts and Their Instructions Move Together
- Plugin Dependency Declaration and Disable-Chain Hints
- Plugin-Activated Main-Agent Override and Bin/ PATH Injection
- Post-Compaction Re-read Protocol for Agent Continuity
- PostToolBatch Hook: Once-Per-Decision-Cycle Injection at the Batch Boundary
- PostToolUse Hook for BSD/GNU CLI Incompatibilities
- PostToolUse Hooks: Automatic Formatting and Linting
- PostToolUse Output Replacement: Hooks That Rewrite Tool Results
- PostToolUse continueOnBlock: Refusal With a Load-Bearing Reason
- PowerShell Tool: Native Windows Shell for Claude Code
- Pre-Install Context-Cost Projection in Plugin Marketplaces
- Pre-Install Plugin Transparency: Capability Inventory and Cost Projection
- PreCompact Hook: Vetoing Compaction at Lifecycle Boundaries
- Production Hosting Topology for Self-Hosted Agent SDK Runtimes
- Production System Prompt Architecture and Techniques
- Programmatic Agent Session Export via `claude agents --json`
- Re-Auditing Context Engineering Across Model Generations
- Reactive Environment Hooks: CwdChanged and FileChanged
- Reducing System-Prompt Token Bloat in Coding Agents
- Reloading Skills Mid-Session in Claude Code
- Request Shaping to Cut Wasted Agent Turns
- Restricted-Access Defensive AI: Project Glasswing as a Deployment Model
- Safe Command Allowlisting: Reducing Approval Fatigue
- Sandbox Credential Masking: Authenticate Without Seeing the Secret
- Selective Checkpoint Restore Across Code and Conversation State
- Session Scheduling with Loop and Cron in Claude Code
- Six-Shape Approval Response Taxonomy: Beyond Binary Allow/Deny
- Skill Eval Loop
- Skill Shell Execution Gate: Disabling Inline Shell from Skills
- Skill disallowed-tools Frontmatter: Skill-Layer Tool Denial
- Sparse-Checkout Worktrees for Monorepo Agent Isolation
- Stated-Understanding Checks: Asking the Agent to Correct You
- StopFailure Hook: Observability for API Error Termination
- Subagent OTel Trace Correlation via agent_id Attribute
- Subprocess PID Namespace Sandboxing in Claude Code
- Symbol Ranking for Agent File Pickers
- System Prompt Delivery Channels on Shared Runners
- Tools: Claude Code, Cursor, and GitHub Copilot
- Turn-Level Context Decisions for AI Coding Sessions
- Verify Agent Diagnoses and Fix Proposals Before Acting
- Video Transcript Skill: Meeting Recording to Markdown
- Workload Identity Federation for Agent Runtimes
- bypassPermissions Silently Overrides allowedTools (The Restricted-Bypass Trap)
- claudeMdExcludes: Selective Ancestor Instruction-File Exclusion
code-generation¶
- Chunking Strategy for RAG-Based Code Completion
- Comment Content as Code-Generation Context
- Completion Failure Taxonomy: Why Code Suggestions Miss
- Constraint Degradation in AI Code Generation
- Constraint Encoding Does Not Fix Constraint Compliance
- Constraint Preambles and the Gain Your Scanner Misses
- Edit Format Selection: Diff vs. Search-Replace vs. Full Rewrite
- Give the Model the Target's Contract, Not Similar Solutions
- Instruction-Guided Code Completion: Controlling What Models Generate
- Iterative Binary Feedback for Pattern Adherence
- Match Architecture Spec Format to Model Capability
- Multi-Layer Specification Redundancy as a Robustness Budget
- Next Edit Suggestions Paradigm for AI Agent Development
- Per-Object Context Allocation (Selective Invariance)
- Repository-Level Retrieval for Code Generation
- Requirement Smells: No Category Signal, Density Unconfirmed
- Security Knowledge Priming for Code Generation (SPARK)
- Standard-Grounded NFR Specs: Quality Up, Correctness Flat
code-review¶
- AI Label as Reviewer Attention Redistribution
- AIRA: Inspection Framework for AI-Generated Code
- Accumulated Behavioral Rules from Review Feedback
- Agent Approval Authority in Code Review
- Agent Host Review Comments: Server-Side Feedback Transport
- Agent PR Volume vs. Value: The Productivity Paradox
- Agent Self-Review Loop for Iterative Self-Improvement
- Agent-Assisted Code Review: Agents as PR First Pass
- Agent-Authored PR Integration: Collaboration Signals That Determine Merge Success
- Agent-Driven PR Slicing
- Agent-Generated Code Maintenance Asymmetry
- Agent-Proposed Merge Resolution
- Agent-Resolved Review Threads
- Agentic Code Review Architecture With Tool-Calling
- Agentic Code Review Patterns and Review Architectures
- Agentic Review Comment Acceptance
- Always-On Agentic PR Security Review
- Author-to-Reviewer Role Inversion in AI-Assisted Teams
- Batched Suggestion Application: Bulk-Apply Agent Fixes on PRs
- Bounded Tool Surfaces for Code Review Agents
- CRA-Only Review and the Merge Rate Gap
- Capability-Pegged Security Re-Scans: Reviewing Unchanged Code When the Scanner Improves
- Capturing Dismissal Reasons for Agent Review Findings
- Classification Before Repair in an Analyzer Backlog
- Claude Code Review
- Cloud Parallel Review Pattern
- Code Cleanliness as an Agent Cost Lever
- Committee Review Pattern for Multi-Agent Code Review
- Compositional Vulnerability Induction in Coding Agents
- Configuring the Code Review Request Surface
- Copilot CLI Agentic Workflows for AI Agent Development
- Deferred Standards Enforcement via Review Agents
- Diff-Based Review Over Output Review
- Diff-Coverage Gating for Agent-Authored Pull Requests
- Ecosystem-Level Integration Friction Governance
- Engineering: Tools, Review, Verification, Security, and Observability
- Evidence-Bundled Agent PRs: Sizing the Reviewer's Effort
- Evidence-Grounded Disagreement in Agentic Code Review (Adversarial Review)
- Human-AI Review Synergy in Agentic Code Review
- Inline Suggestion Attachment in Agent Code Review
- Instruction-Aware Automated Code Review
- Interaction-Pattern Evaluation for Agentic PRs
- Interactive Canvases: Agent-Generated Visual Artifacts as Outputs
- Kaizen-Style Continuous Code Quality Loop (Pomona)
- LLM Code Review Overcorrection for AI Agent Development
- Language Selection Scored on Review Cost
- Law of Triviality in AI PRs for AI Agent Development
- Match Tool Instructions to the Agent Workflow
- Minimality Prompts as a Patch-Size Control
- PR Description Style as a Lever for Agent PR Merge Rates
- PR Scope Creep as a Human Review Bottleneck
- Per-Line Requirement Citations for Hallucination Detection
- Per-Reviewer Context Views for Code Review Agents
- Phantom Symbol Detection for LLM API Migration
- Polya Small-Steps: Using AI to Think Better, Not Think Less
- Post-Merge Fix Signals for Agent Merges
- Precise Debugging: Measure Edit Precision, Not Just Test Pass Rate
- Predicting Reviewable Code: Pre-Flagging Functions Reviewers Will Delete
- Preempting Agentic PR Rejection by Failure-Mode Category
- Reproduce-Before-Report Verification Gate
- Review-Comment-Derived Benchmarks for Code Review Agents
- Review-Feedback-to-Rule Loop: Promoting Recurring PR Comments into Harness Rules
- Review-Then-Implement Loop for AI Agent Development
- Reviewer Habituation in Agent PR Review
- Reviewer Precision as a Pipeline Quality Proxy
- Reviewer Theme Distribution Audit for AI Code Review
- Reviewer's Playbook for Agent-Authored Pull Requests
- Reviewing What a Memory-Fed Autofix Taught
- Risk-Score Threshold Calibration for Auto-Approval
- Self-Improving Code Review Agents — Learned Rules
- Signal Over Volume in AI Review for AI Agent Development
- Structure-Aware Diff Labeling with Two-Stage LLM Pipelines
- Supply-Chain Security Debt in Agent Pull Requests
- The Bottleneck Migration When Humans Supervise Agents
- The Merge-Conflict Resolution Skill: What to Encode
- The Security Review Gap in AI-Authored PRs
- Three-Depth In-Session Security Review
- Tiered Code Review: AI-First with Human Escalation
- Trusting Human Review to Catch Deliberate Agent Sabotage
- Tunable Effort Levels for Code Review Agents
- Velocity-Quality Asymmetry: Why AI Speed Gains Fade
- Verification Capacity as the Agent Quality Ceiling
- Verification-Gated Agent Autonomy via Automated Review
- Verifying Agent Changes in the Copilot App
context-engineering¶
- @import Composition Pattern for Agent Instruction Files
- ACDL: A Language for Describing Agentic LLM Contexts
- AGENTS.md as a Table of Contents, Not an Encyclopedia
- AOCI: Symbolic-Semantic Repository Indexing
- Action-Gated Context Trimming for Long-Horizon Agents
- Addressable Recall Compaction: Compact to Citations, Not Summaries
- Advanced Tool Use: Scaling Agent Tool Libraries
- Agent Memory Patterns: Learning Across Conversations
- Agent-Computer Interface (ACI): Tool Design as UX Discipline
- Agent-Initiated Rubric-Gated Self-Compaction (SelfCompact)
- Agent-Powered Codebase Q&A and Onboarding Workflow
- Agent-Ready Data Architecture for Analytics Agents
- Agent-Tuned Code Search: Retrieval Built for the Loop
- App-Window Snapshot as Agent Context
- Assuming Loaded Skills Stay Enforced in Long Contexts
- Attention Latch: When Agents Stay Anchored to Stale Instructions
- Attention Sinks: Why First Tokens Always Win
- Attributed Cache Misses: Reading Why a Prefix Diverged
- Auto-Merging a Wiki Agent's Documentation Pull Requests
- Batch File Operations via Bash Scripts for AI Agents
- Belief Inertia After Tool-Map Drift in AI Agents
- Budgeted Verification of Inherited Agent Constraints
- CLI Scripts as Agent Tools: Return Only What Matters
- CLI-IDE-GitHub Context Ladder for AI Agent Development
- Capability Declarations for Agents That Act on Data
- Catastrophic Remembering: Instruction Files That Only Grow
- Choosing a Compression Budget for Agent Control Context
- Choosing a Skill Loading Method for Agents
- Chunking Strategy for RAG-Based Code Completion
- Claim-Scoped Invalidation for Agent Memory
- Claude Code Dynamic Workflows
- CoALA Memory Taxonomy as a Classifier for Harness Artifacts
- Code-Native Memory Substrates for Coding Agents
- Codebase-Derived Pattern Libraries as Agent Context
- Coding-Agent Working-Set Coverage (Coherence Debt)
- Comment Content as Code-Generation Context
- Compiled Specialist Agents: Muscle Memory for Recurring Intent
- Component-Wise RAG Prioritization for Software Engineering Tasks
- Compositional Skill Routing for Large Skill Libraries
- Compound Engineering: Learning Loops That Make Each Feature Easier
- Configuration File Structure Does Not Drive Compliance
- Consistent-format customer capture
- Context Budget Allocation: Spending Every Token Wisely
- Context Compiler: Deterministic Assembly Over Bigger Windows
- Context Compression Strategies: Offloading and Summarization
- Context Engineering (Training Module)
- Context Engineering: Shaping AI Agent Input and Attention
- Context Engineering: The Practice of Shaping Agent Context
- Context Hub: On-Demand Versioned API Docs for Coding Agents
- Context Lifecycle Management: Beyond Store and Retrieve
- Context Poisoning: When Hallucinations Become Premises
- Context Priming: Pre-Loading Files for AI Agent Tasks
- Context Quality as a Leading Indicator of Agent Reliability
- Context Window Anxiety: Countering Premature Task Closure
- Context Window Management: Understanding the Dumb Zone
- Context-Injected Error Recovery for AI Agent Development
- Context-Usage Attribution: Per-Source Breakdown of Agent Context
- Context-Window Diagnostic Tooling: Identifying Context-Heavy Tools
- Convenience Loops and AI-Friendly Code in Your Stack
- Conversation Registers for AI Coding Sessions
- Copilot Memory and Cross-Agent Persistence
- Copilot Spaces: Curated Context Collections for Grounding
- Corpus Shape as a Retrieval Design Constraint
- Coverage-Aware Skill Selection Under a Token Budget
- Critical Instruction Repetition via Primacy and Recency
- Cross-Functional Knowledge Artifacts
- Cross-Lingual Prompt Preprocessing (Local-LLM Token Arbitrage)
- Cross-Reference Dereference Hop in Retrieval Loops
- Cross-Repo Agent Search: GitHub-API-Backed Text Search Beyond the Workspace
- Deferred Standards Enforcement via Review Agents
- Deterministic Anchoring: Static Facts as Stable Context
- Diagram as the Shared Spec: One Artifact for the Picture and the Prompt
- Disable Attribution Headers to Preserve KV Cache in Local Inference
- Discoverable vs Non-Discoverable Context for Agents
- Distractor Interference: Why Relevance Is Not Enough
- Distributed Computing Parallels in Agent Architecture
- Documentation Read Counts Measure Retrievability, Not Value
- Documentation-Grounding MCP Servers for Vendor SDKs
- Documenting Code the Agent Can Already Read
- Dynamic System Prompt Composition
- Dynamic Tool Fetching Destroys KV Cache Performance
- Elastic Context Orchestration: A Per-Turn Vocabulary for Long-Horizon Search Agents
- Encode Project Conventions in Distributed AGENTS.md Files
- Encoding Product-Design Taste into Agent Context
- Environment Specification as Context: Closing the Version Gap
- Episodic Memory Retrieval for AI Coding Agent Loops
- Epistemic Working Memory for Multi-Hop Reasoning (SLEUTH)
- Error Preservation in Context for AI Agent Development
- Evaluating AGENTS.md: When Context Files Hurt More Than Help
- Event-Driven System Reminders for AI Agent Development
- Evolving Playbooks: Incremental Context That Preserves Knowledge
- Example-Driven vs Rule-Driven Instructions
- Exclude Dynamic System Prompt Sections for Cross-Machine Cache Sharing
- Executable Memory: User State as Code for Personalized Agents
- Execution-State Ledger for Long-Horizon Coding Agents
- Exhaustive Retrieval for Listing Questions
- Fact Supersession Memory for Code Assistants
- Filesystem-Based Tool Discovery for AI Agent Development
- Filter and Aggregate Data in the Execution Environment
- Formal Process Models as Prompting Scaffolds (Petri Net of Thoughts)
- Foundations: Context Engineering and Instructions
- Four Reporting Levels for Agent Working Memory Evaluation
- Four-Phase Agent Delegation with Curated Artifacts
- Functional folder taxonomy
- Gate Generation on Retrieval Sufficiency, Not Model Confidence
- Generated Questionnaires: Eliciting Someone Else's Context
- GitHub Copilot: Context Engineering & Agent Workflows
- Give the Model the Target's Contract, Not Similar Solutions
- Goal Recitation: Countering Drift in Long Sessions
- Governed Sources of Truth for Analytics Agents (Structure Over Access)
- Graceful Tool-Output Truncation: The PARTIAL Signal
- Grounding Agents in Code the Model Has Never Seen
- Guardrails Beat Guidance: Rule Design for Coding Agents
- Handoff Skill: Structured Context Transfer Between Agent Sessions
- Hierarchical CLAUDE.md: Structuring Context Files at Multiple Levels
- Hints Over Code Samples in Agent Prompts
- How the Four Agent Engineering Disciplines Compound
- Hypothetical Classification for Large Label Vocabularies
- In-Thread Side-Channel: Bounded Side Questions Without Losing the Main Task
- Indexed Regex Search for Agent Tools
- Injected-Context Cost Attribution in Agent Workflows
- Instruction-Guided Code Completion: Controlling What Models Generate
- Knowledge Gap or Skill Gap: Triage Before Writing Context
- LLM Map-Reduce Pattern for Parallel Input Processing
- LLM-Driven Logical Retrieval: Boolean Queries over an Inverted Index
- Lay the Architectural Foundation by Hand Before Delegating
- Layer Agent Instructions by Specificity Across Scopes
- Layered Context Architecture for AI Agent Development
- Live Browser as Agent Context Channel
- Living-Docs-Grounded Agent Design Conversations
- Long Context vs Retrieval: The Break-Even Decision
- Lost in the Middle: The U-Shaped Attention Curve
- MCP Tool Result Persistence via _meta Annotation
- MCP alwaysLoad: Classifying Servers as Eager or Just-in-Time
- Manual Compaction Strategy for Dumb Zone Mitigation
- Mask Tools Instead of Removing Them
- Measuring Reacquisition Cost Under Context Compaction
- Memory Synthesis: Extracting Lessons from Execution Logs
- Mid-Session Config Changes as Invisible Cache Invalidators
- Mise en Place for Agentic Coding
- Model-Switch Lifecycle Hooks: Gating a Mid-Session Model Change
- Narrative Problem Reformulation for Code Generation
- Next Edit Suggestions Paradigm for AI Agent Development
- Objective Drift: When Agents Lose Sight of the Goal
- Observation Masking: Filter Tool Outputs from Context
- Open Agent School Pattern Mapping for Practitioners
- Organizational Context Layer for Agents (Company Brain)
- Organizing Filesystem Agent Memory for Retrieval Cost
- PEEK: Orientation Cache for Recurring-Context Agents
- Per-Object Context Allocation (Selective Invariance)
- Per-Subagent Instruction Inheritance (omitClaudeMd)
- Per-Type Retention Policy for Agent Compaction (Knowledge Triage)
- Phase-Specific Context Assembly for AI Agent Development
- Plan Mode: Read-Only Exploration Before Implementation
- Plan files as resumable artifacts
- Plan mode for knowledge artifacts
- Post-Compaction Re-read Protocol for Agent Continuity
- Pre-Execution Codebase Exploration for AI Coding Agents
- PreCompact Hook: Vetoing Compaction at Lifecycle Boundaries
- Production System Prompt Architecture and Techniques
- Project-Scoped Agent Workspace: Durable Context for Clean-Context Subagents
- Prompt Cache Keepalive for Agent Pauses
- Prompt Caching: Architectural Discipline for Agents
- Prompt Chaining: Sequential LLM Calls for Agent Workflows
- Prompt Compression: Maximizing Signal Per Token
- Prompt Injection: A First-Class Threat to Agentic Systems
- Prompt Layering: How Instructions Stack and Override
- Prompt Transpilation: Instructions as Build Artifacts
- Proprioceptive Context Dashboard: Agent Self-Managed Context
- Prototype Before Optimizing: Establish Quality Baselines Before Token Constraints
- Publishing Agent Instructions Outside the Repo (design.md)
- Query-Conditioned Reuse of Retrieved Agent Trajectories
- RAG Architecture as a Poisoning Robustness Decision
- RAG over Thinking Traces: Index Reasoning Trajectories Instead of Documents
- Re-Auditing Context Engineering Across Model Generations
- Reasoning Retention and Compaction as Harness Settings
- Reducing System-Prompt Token Bloat in Coding Agents
- Repository Map Pattern: AST + PageRank for Dynamic Code
- Repository Perturbation as Context-Reasoning Diagnosis (RepoMirage)
- Repository-Level Retrieval for Code Generation
- Reproducibility Artifacts as Agent Context
- Retrieval-Augmented Agent Workflows: On-Demand Context
- Role Orchestration on a Single Model
- Runtime Resource Limits as Prompt Context
- Sandbox-Enforced PII Tokenization in Agent Workflows
- Schema-Guided Graph Retrieval
- Scoring Constraint Loss and Tool Reach as Separate Risks
- Scoring a Compaction Policy on Latency and Billed Cost
- Seeding Agent Context: Breadcrumbs in Code
- Selective Rewind Summarization: Compress Earlier Turns, Keep Recent Ones Intact
- Self-Correcting Memory: Evidence-Backed Claim Repair
- Semantic Caching for Multi-Agent Code Systems
- Semantic Context Loading: Language Server Plugins for Agents
- Semantic Density Optimization for Agent Codebases
- Semantic Tool Output: Designing for Agent Readability
- Session Initialization Ritual: How Agents Orient Themselves
- Session Recap: Goal-Shaped Handoff at Context Boundaries
- Shared Context Bundle Registry for Agent Teams
- Shortening Old Tool Results Under Context Pressure (Half-Life Truncation)
- Silent Handoff Failure in Delegated Code Search
- Single-Layer Prompt Injection Defense Anti-Pattern
- Skill Context Isolation: Forking the Skill into a Subagent Window
- Skill Loadout Curation for Coding Agents
- Source Code Minification for State-in-Context Agents
- Spec-Anchored Drift-Gated Architecture (Spec Growth Engine)
- Specification Memory: What a Shared Agent Workspace Keeps
- Stage Elision Before Summarization
- Stale AI Configuration Artifacts (Context Rot)
- State-Conditioned Evidence Selection for Mid-Task Retrieval
- Stateful Iteration State-Carry: Typed Persistent State for Long Agent Loops
- Static-Context Tool Residence: Two Overrides on Hit Rate
- Structure Prompts with Static Content First to Maximize Cache Hits
- Structured Domain Retrieval: Knowledge Graphs and Case-Based Reasoning
- Structured Task-State Ledger for Tool-Calling Agents (LedgerAgent)
- Sub-Agents for Fan-Out Research and Context Isolation
- Subagent vs In-Context Skill Execution
- Subtask-Level Memory for Software Engineering Agents
- Symbol Ranking for Agent File Pickers
- System Prompt Altitude: Specific Without Being Brittle
- Team OS: Coding-Agent Repo as Cross-Functional Team Brain
- Terminal Tool Output Compression: Filtering Predictable Noise at the Harness
- Test Harness Design for LLM Context Windows
- The Context Ceiling -- Where AI Fails Expert Architects
- The Handoff Tax: What a Receiving Model Should Inherit
- The Infinite Context Anti-Pattern in Agent Systems
- The Instruction Compliance Ceiling: How Rule Count Limits AI
- The Kitchen Sink Session Anti-Pattern in AI Agents
- The No-Op Test: Prune Agent Docs by Behavior, Not Length
- The Orchestrator's Attention Budget: Delegating to Protect Context
- The Plan-First Loop: Always Design Before Writing Code
- The Recall Trap: Tuning a Code Retriever on Recall@k at a Fixed Slot Budget
- The Research-Plan-Implement Pattern
- The Specification as Prompt: Existing Artifacts as Agent
- The Task Framing Irrelevance Fallacy in Agent Prompting
- Three Knowledge Tiers: Sourced, Unverified, Hallucinated
- Trajectory Attribution for Context Repair (TRACE)
- Turn-Level Context Decisions for AI Coding Sessions
- Typed Context Buys Addressability, Not Token Savings
- Typed Memory Provenance and Assertion Release Gating
- Typed Pseudocode for Skill Libraries (Skill-as-Pseudocode)
- Ubiquitous Language for AI Plans
- Usage-Reinforced Memory Decay for Long-Running Agents
- Validating Token-Optimized Formats Inside Agentic Loops
- Verbatim Failure Records in Small-Model Agent Transcripts
- Version-Controlled Agent Context (Git Context Controller)
- When a Skill Graph Cannot Beat the Ranker (Pre-Filter Topology Bound)
- Why an Encoded Rule Still Fails After a Passing Eval
- Wiki Memory: Agent-Maintained Compressed Knowledge Base
- llms.txt: Making Your Project Discoverable to AI Agents
copilot¶
- AGENTS.md Design Patterns for Effective Agent Files
- Agent Approval Authority in Code Review
- Agent Environment Bootstrapping for AI Agent Development
- Agent Governance Policies for AI Agent Development
- Agent HQ (Multi-Agent Platform) for AI Agent Development
- Agent Host Review Comments: Server-Side Feedback Transport
- Agent Mission Control for Orchestrating Agent Tasks
- Agent-Resolved Review Threads
- Agentic Code Review Architecture With Tool-Calling
- Anonymized Customization Metrics in the Copilot CLI
- Auto Model Selection: Harness-Driven Routing per Task
- Bounding an Embedded Copilot SDK Agent's Tool Set
- Canvas as Control Surface: Steering a Long-Running Agent Mid-Run
- Canvas as Durable Workflow State: The Four-Step Blueprint and What It Costs
- Capturing Dismissal Reasons for Agent Review Findings
- Classification Before Repair in an Analyzer Backlog
- Cloud-Local Agent Handoff for AI Agent Development
- Cohort Segmentation in the Copilot Usage Metrics API
- Comment-Triggered Agent Dispatch on Issues and PRs
- Configuring the Code Review Request Surface
- Content Exclusion Gap: AI Security Boundaries by Mode
- Copilot App Customize Tab vs Repo Configuration Files
- Copilot Auto Tiers: Weighting Cost Against Quality
- Copilot CLI Agentic Workflows for AI Agent Development
- Copilot CLI BYOK and Local Model Support
- Copilot Cloud Agent Organization Controls
- Copilot Cloud Agent Three-Phase Execution Model
- Copilot Memory and Cross-Agent Persistence
- Copilot Spaces: Curated Context Collections for Grounding
- Copilot Unified Sessions View and CLI Agent in JetBrains IDEs
- Cross-IDE Plugin Discovery: One Install Surface, Many Consuming Agents
- Delegating Delivery Stages to GitHub Agent Apps
- Delegating Dependabot Pull Request Triage to an Agent
- Dependabot Agent Assignment for AI-Driven Vulnerability Remediation
- Dispatch-Time Reasoning Level for Delegated Agents
- Embedding the Copilot SDK in a Managed Java Runtime
- Enterprise-Managed Plugin Governance for Agent CLIs
- GitHub Agentic Workflows for Automating Dev Processes
- GitHub Copilot Advanced Patterns: Multi-Agent and Automation
- GitHub Copilot Agent Mode for AI Agent Development
- GitHub Copilot App Slash Commands and What They Change
- GitHub Copilot Coding Agent for AI Agent Development
- GitHub Copilot Custom Agents and Skills Extensibility Guide
- GitHub Copilot Dedicated App as Agent-First Surface
- GitHub Copilot Extensions for AI Agent Development
- GitHub Copilot MCP Integration for AI Agent Development
- GitHub Copilot Platform Surface Map: All Capabilities
- GitHub Copilot SDK for AI Agent Development
- GitHub Copilot Training Modules for Engineering Teams
- GitHub Copilot for AI Agent Development
- GitHub Copilot: Context Engineering & Agent Workflows
- GitHub Copilot: Customization Primitives and Stack
- GitHub Copilot: Harness Engineering for Agent-Ready Code
- GitHub Copilot: Model Selection, Routing, and Costs
- GitHub Copilot: Team Adoption and Governance Guide
- GitHub Models in Actions for AI-Driven CI Workflows
- GitHub's Copilot Cost Levers at Constant Task Quality
- Local Sandboxing in the Copilot App: The Credential Axis
- MCP LLM Sampling: Servers Requesting AI Inference Mid-Tool
- Managing Agent Skills from the GitHub CLI with gh skill
- Monorepo Skill and Agent Discovery: Hierarchical Configuration
- Next Edit Suggestions Paradigm for AI Agent Development
- Non-Retirable Approval Rules for Agent Operations
- One-Click CI Auto-Fix: Human-Triggered Cloud-Agent Remediation for Failing GitHub Actions
- Org-Membership-Gated Agent Entitlement
- Per-Agent-App Attribution in the Copilot Usage Metrics API
- Per-Surface Verification of Agent Plugin Packages
- Porting a Coding-Agent Harness Beyond Engineering
- Pre-Execution Risk Classification for Terminal Commands
- Proprietary-to-Open-Standard Tool Migration (Copilot Extensions to MCP)
- Reading Copilot Feature Engagement by Its Threshold
- Reading a Vendor-Computed AI Coding ROI Dashboard
- Reviewing What a Memory-Fed Autofix Taught
- Runtime Workflow Selection Across Models (Project HydraFusion)
- Selective Network Access in Agent Sandboxes: The allowNetwork Pattern
- Semantic Issue Search from Chat vs Query Syntax
- Shared Agent Context Store API: When to Expose Curated Context as an Endpoint
- Sizing Vendor-Emitted Agent Telemetry by Signal Tier
- Team-Scoped Agent Policy Delegation
- Tenant Model Policy: Organization-Scoped Rules for AI Model Selection
- Tools: Claude Code, Cursor, and GitHub Copilot
- VS Code Agents App: Agent-Native Parallel Task Execution
- Verifying Agent Changes in the Copilot App
- Working Inside an Enterprise-Managed Agent Sandbox Policy
- copilot-instructions.md as a Repo-Level Instruction Convention
cost-performance¶
- A Governance Framework for Production Agents
- Action-Gated Context Trimming for Long-Horizon Agents
- Adaptive Generate-Rank-Verify Under Costly Verification
- Adaptive Sandbox Fan-Out Controller
- Adaptive Validation Task Selection
- Administrative Effort Ceilings for Reasoning Budget
- Advanced Tool Use: Scaling Agent Tool Libraries
- Advisory Prompts Distilled from Reasoning Traces
- Agent Composition Patterns for Multi-Agent Workflows
- Agent JIT Compilation: Compile Tasks Into Executable Plans
- Agent Loop Go/No-Go: When Looping Earns Its Cost
- Agent Observability with OpenTelemetry and Trajectory Logging
- Agent-Client Admission Control for Agentic Traffic
- Agent-Tuned Code Search: Retrieval Built for the Loop
- Assuming Agent Interchangeability in Long-Running Teams
- Asynchronous Agent I/O and Speculative Tool Calling
- Attributed Cache Misses: Reading Why a Prefix Diverged
- Auth-Isolation as the MCP-vs-CLI Selection Heuristic
- Auto Model Selection: Harness-Driven Routing per Task
- BYOK Model Token Visibility: Closing the Observability Gap on Self-Hosted Routes
- Background Todo Agent: Offload Plan Maintenance to a Lightweight Model
- Batch File Operations via Bash Scripts for AI Agents
- Benchmark-Driven Tool Selection for Code Generation
- Blind Resampling Over Self-Repair in Small Code Models
- Bounded Batch Dispatch for Parallel Agent Execution
- Bounded Repair-Loop Iterations
- Bounded Tool Surfaces for Code Review Agents
- Bounding a Headless Codex Run Without a Turn Cap
- CLI Scripts as Agent Tools: Return Only What Matters
- Cache-Prefix Staggering for Sibling Agent Fan-Out
- Cache-Safe Routing Boundaries: Where a Router May Act
- Calibrated Early Termination and Warm Restart for Agent Runs (FailFast-RestartSmart)
- Canary Tools for Diagnosing Tool-Selection Reasoning
- Canvas as Control Surface: Steering a Long-Running Agent Mid-Run
- Centralized LLM Gateway for Per-Dimension Agent Budgets
- Chance-Corrected Shortlist Depth Sizing for Tool Retrieval (Bits-over-Random)
- Cheaper-Per-Token Model Upgrades That Cost More Per Task
- Choosing a Compression Budget for Agent Control Context
- Choosing a Skill Loading Method for Agents
- Choosing an Agent Tool Interface: Shell or Typed Catalog
- Claude Code Dynamic Workflows
- Claude Code Feature Flags and Environment Variables
- Code Cleanliness as an Agent Cost Lever
- Code Health as a Signal for Agent-Generated Test Quality
- Code Interpreter as a Primary Agent Tool
- Codified Effort and Escalation Policy in the Instruction File
- Cognitive Reasoning vs Execution: A Two-Layer Agent
- Cohesion-Aware Task Partitioning for Multi-Agent Coding
- Comparative Judging for Agent Configuration Ranking
- Comparison-Only Advisor: Steering a Large Actor With a Tiny Comparator
- Component-Wise RAG Prioritization for Software Engineering Tasks
- Configuring the Code Review Request Surface
- Consolidate Agent Tools to Reduce Cognitive Overhead
- Constraint Encoding Does Not Fix Constraint Compliance
- Context Budget Allocation: Spending Every Token Wisely
- Context Compression Strategies: Offloading and Summarization
- Context Hub: On-Demand Versioned API Docs for Coding Agents
- Context Lifecycle Management: Beyond Store and Retrieve
- Contextual Capability Calibration for Multi-Agent Delegation
- Continuous Triage: Automating Issue Classification with AI Workflows
- Coordination Channel Policy for Multi-Agent Coding
- Copilot Auto Tiers: Weighting Cost Against Quality
- Copilot CLI BYOK and Local Model Support
- Copilot vs Claude Billing Semantics for Enterprise Teams
- Cost-Aware Agent Design: Route by Complexity, Not Habit
- Cost-Aware Skill Rewriting: Preserve Operational Anchors, Not Skill Tokens
- Cost-Aware Tracing for Skill Distillation
- Cost-Driven Model Routing Without Quality Monitoring
- Cost-Inefficient Behaviors in Coding Agents
- Cost-Quality Pareto Measurement for Agent Configurations
- Cross-Component Interference in Agent Scaffolds
- Cross-Lingual Prompt Preprocessing (Local-LLM Token Arbitrage)
- Cross-Vendor Competitive Routing for LLM Selection
- DSPy: Programmatic Prompt Optimization for Compound Agent Systems
- Decision Unbundling: What Moves Into the Orchestrator
- Deliberation-Inducing Cues That Multiply Reasoning Cost
- Designing Agent Tools Like APIs
- Deterministic Fast Paths: Answer Without a Model Call
- Deterministic Orchestration for Structured Modernization
- Difficulty-Aware Topology Selection for Coding Agents
- Disable Attribution Headers to Preserve KV Cache in Local Inference
- Dispatch-Time Reasoning Level for Delegated Agents
- Dual-Budget Control for Search Agents: VOI Scoring Per Action
- Dynamic Tool Fetching Destroys KV Cache Performance
- Edit Format Selection: Diff vs. Search-Replace vs. Full Rewrite
- Effective Feedback Compute (EFC) for Harness Comparison
- Effort-Aware Hooks: Reading the Reasoning Tier from PreToolUse and PostToolUse
- Equivalence Testing for Agent Configuration Changes
- Evaluating AGENTS.md: When Context Files Hurt More Than Help
- Exclude Dynamic System Prompt Sections for Cross-Machine Cache Sharing
- Execution Budgeting in Agentic Program Repair
- Fan-Out Synthesis Pattern for AI Agent Development
- Feedback as Capability Equalizer: Iterative Feedback Outweighs Model Scale
- Filesystem-Based Tool Discovery for AI Agent Development
- Filter and Aggregate Data in the Execution Environment
- First-Party Agent Composition: Agent-Built Features
- First-Proposal Execution in Agent Loops
- Five Design Decisions for MCP Servers and Clients
- Fleet Harness Attribution: Pinning the Model to Compare Whole Harnesses
- Fleet-Level Irreversibility Budgets for Agent Effects
- Framework-First Agent Development: An AI Anti-Pattern
- Frozen Playbook Reuse Without Target-Side Validation
- Frozen Task Sets for Affordable Agent A/B Testing
- Function-Level Debugger Interfaces for Coding Agents
- Future-Based Asynchronous Function Calling
- Gateway Hint Headers for Routing and Budgeting Agent Calls
- Gateway Model Routing: Treat the LLM Gateway as a Discovery Source
- GitHub Copilot: Model Selection, Routing, and Costs
- GitHub's Copilot Cost Levers at Constant Task Quality
- Google Search Console Monitoring Workflow
- Harness-Controlled Token Economics (The Harness Effect)
- Head-to-Head Evaluation of Competing MCP Servers
- Headless Claude in CI: Using -p and --max-turns for Safe Pipeline Integration
- Heuristic-Based Effort Scaling in Agent System Prompts
- Hint-Driven Concurrency for Read-Only MCP Tools
- How the Four Agent Engineering Disciplines Compound
- Human-Equivalent Hours for Autonomous Coding Agent Productivity
- Idle-Time Speculative Planning for ReAct Agents
- In-Agent Task Prioritization: Ranking the Next Action
- Indexed Regex Search for Agent Tools
- Indiscriminate Structured Reasoning on Every Agent Task
- Injected-Context Cost Attribution in Agent Workflows
- Interactive Effort Sliders: Per-Turn Reasoning-Budget Controls
- LLM-Driven Logical Retrieval: Boolean Queries over an Inverted Index
- LLM-as-Code Agentic Programming for Agent Harnesses
- Language Choice as an Agent Token-Cost Lever
- Lexical-First Retrieval for Agentic Search: When BM25 Is Enough
- Line-Anchored Feedback: Deliver Change Requests as Inline Comments
- Local Model Viability Factors for Coding
- Long Context vs Retrieval: The Break-Even Decision
- Loop Budgeting: Allocating Iteration and Token Budget Across Turns
- MCP Client Design: Building Robust Host-Side Logic
- MCP Server Design: Building Agent-Friendly Servers
- MCP alwaysLoad: Classifying Servers as Eager or Just-in-Time
- MCP-vs-CLI Cost Ratios Are a Property of the Scaffolding
- Machine-Readable Error Responses for AI Agents (RFC 9457)
- Mask Tools Instead of Removing Them
- Match Architecture Spec Format to Model Capability
- Match Tool Instructions to the Agent Workflow
- Measuring Reacquisition Cost Under Context Compaction
- Measuring Refactoring Payback in Tokens
- Measuring the Verification Tax on Agent Output
- Mid-Session Config Changes as Invisible Cache Invalidators
- Minimum-Cost Evidence Selection for Agent Changes (Assurance Envelopes)
- Minimum-Sufficient Control Ladder: Escalate by Failure Mode
- Minimum-Sufficient Execution: Estimate Scope Before Spending Budget
- Model Deprecation Lifecycle for Agent Workloads
- Model-Directed Subagent Tiering: Lead Model Picks the Tier
- Model-ID-as-Dependency: Migration Protocol for Deprecation Churn
- Model-Neutral Agent Architecture: Model Portability Over Cloud Portability
- Model-Set Parity: Reading Harness Efficiency Claims
- Model-Switch Lifecycle Hooks: Gating a Mid-Session Model Change
- Multi-Model Plan Synthesis for System Architecture
- Multi-Shape BYOK Provider: Declare API Family per Endpoint
- Natural Language Tool Selection (NLT)
- Observation Masking: Filter Tool Outputs from Context
- Observation-Driven Coordination: CRDT-Based Parallel Agent
- One-Shot Record and Deterministic Replay for Periodic Agent Tasks
- Open Agent School Pattern Mapping for Practitioners
- OpenAPI Documentation Smells for Agent-Ready APIs
- OpenAPI as the Source of Truth for Agent Tool Definitions
- Opponent Processor / Multi-Agent Debate Pattern
- Outcome Pricing as a Scope Signal
- Parameter-Keyed Caching and Dependency-Aware Parallelism for Plan-Execute Pipelines
- Parsimonious Agent Routing for Multi-Agent Dispatch
- Pattern Selection Map
- Per-Call Budget Hints on Tool Invocations
- Per-Plugin Token-Cost Attribution via claude plugin details
- Per-Run Budget Reservation for Coding Agent Model Calls
- Per-Task Agent Routing Across Coding Harnesses
- Per-Tool Extended Reasoning Opt-In: Tool-Call-Scoped Budgets
- Perceived Model Degradation: Why Vibes Are Not Evals
- Persistent Shared Search Sub-Agent for Output-Token Reuse
- Persistent-Connection Agent Transport
- Plan Mode: Read-Only Exploration Before Implementation
- Planning Stage Preconditions: Budget Headroom and Task Text
- Policy-Graded Evaluation of Coding Agents
- Pre-Execution Failure Scoring with a Draft Model (Speculative Uncertainty)
- Pre-Install Context-Cost Projection in Plugin Marketplaces
- Pre-Install Plugin Transparency: Capability Inventory and Cost Projection
- Pricier-Per-Token Models That Cost Less Per Task
- Proactive Idle-Time Anticipation (ProAct)
- Probe-Run Calibration for Predicting Agent Token Spend
- Production MCP Agent Stack: Sequencing Six Decisions into One Deployment
- Progressive Spend Threshold Alerting for Agent Cost Governance
- Prompt Cache Keepalive for Agent Pauses
- Prompt Caching: Architectural Discipline for Agents
- Prompt Compression: Maximizing Signal Per Token
- Prototype Before Optimizing: Establish Quality Baselines Before Token Constraints
- Provider-Hosted Subagent Delegation: One Model Price for the Whole Tree
- Purpose-Built Eval Suites for Model and Harness Swaps
- Query-Conditioned Reuse of Retrieved Agent Trajectories
- Reading a Vendor-Computed AI Coding ROI Dashboard
- Reasoning Budget Allocation: The Reasoning Sandwich
- Reasoning Effort Over Tool Scaffolding for First-Try Reliability
- Recurring Control Belongs in Harness Code, Not Context
- Reducing System-Prompt Token Bloat in Coding Agents
- Reflective Prompt Evolution with Pareto Selection (GEPA)
- Rented Sandboxes for Coding-Agent Benchmark Runs
- Repairing Agent Prompts from Trace Contrast, Not Search
- Request Shaping to Cut Wasted Agent Turns
- Residual Completion for Stateful Agent Handoffs (CFRC)
- Restricting a Coding Agent to a Single execute_code Tool
- Retrieval-Augmented Agent Workflows: On-Demand Context
- Rewriting a CLI Into a JSON Payload for Agents
- Role Orchestration on a Single Model
- Rolling Out CLI Coding Agents at Organization Scale
- Router-Imposed Quality Ceiling: Committing Before Output
- Routing Break-Even: When a Cheaper Model Actually Pays
- Routing Decision Framework: Which Routing Pattern Fits Which Signal
- Running Several Coding Agents Behind One Harness Interface
- Runtime Resource Limits as Prompt Context
- Runtime Workflow Selection Across Models (Project HydraFusion)
- Scoring a Compaction Policy on Latency and Billed Cost
- Scout-Then-Route: Verify the Handoff Before Routing
- Security Budget as Token Economics
- Self-Healing Tool Routing
- Semantic Caching for Multi-Agent Code Systems
- Semantic Context Loading: Language Server Plugins for Agents
- Semantic Density Optimization for Agent Codebases
- Semantic Tool Output: Designing for Agent Readability
- Silent Handoff Failure in Delegated Code Search
- Skill Over-Trust: Treating Topical Relevance as Evidence a Skill Helps
- Skill Review Without a Token Cost Baseline
- Source Code Minification for State-in-Context Agents
- Specialist Orchestrated Queuing for Multi-Agent SE (SPOQ)
- Specialized Small Language Models as Agent Sub-Tools
- Splitting an Agent Token Budget at the Scaling Inflection Point
- Stateful Iteration State-Carry: Typed Persistent State for Long Agent Loops
- Static-Context Tool Residence: Two Overrides on Hit Rate
- Structure Prompts with Static Content First to Maximize Cache Hits
- Task Feasibility Awareness: Stop Before You Start
- Task Shape Decides What a Heavier Agent Harness Buys
- Temporal Token Routing: Batch and Flex Tiers for Non-Urgent Work
- Tenant Model Policy: Organization-Scoped Rules for AI Model Selection
- Terminal-First Agent Interfaces with Browser Escalation
- The Advisor Strategy: Frontier Model as Strategic Advisor
- The Handoff Tax: What a Receiving Model Should Inherit
- The Harness as Product: What Listed-Rate Pricing Buys
- The Infinite Context Anti-Pattern in Agent Systems
- The Kitchen Sink Session Anti-Pattern in AI Agents
- The Model Economics of Agent Swarms: Cost and Width
- The Model Preference Fallacy in Comparison Content
- The Plan-First Loop: Always Design Before Writing Code
- The Token Price Index Fallacy in Agent Cost Planning
- Token Reduction Mistaken for Cost Reduction
- Token-Cost Profiling and Reduction for Always-On Agentic Workflows
- Token-Efficient Code Generation: Structural Beats Prompting
- Token-Efficient Tool Design: Tools That Don't Eat Your Context
- Tokenizer Swap Tax: Budgeting for Model Migrations That Change Token Counts
- Tool Calling Schema Standards for AI Agent Development
- Tool Description Quality for Effective Agent Guidance
- Tool Engineering (Training Module)
- Tool Necessity Probing: Reading Tool-Call Decisions From Hidden States
- Tools as Typed Code Stubs (Programmatic Tool Calling)
- Toolset Agentization: Wrapping Co-Used Tools as Sub-Agents
- Trajectory-Aware Benchmark Subset Selection for Agents
- Trajectory-Conditioned Model Escalation (SWE-Router)
- Tunable Effort Levels for Code Review Agents
- Typed Context Buys Addressability, Not Token Savings
- Typed Tool Surfaces and Out-of-Loop Correctness Gates
- Unbounded Agent Feedback Paths (Infinite Agentic Loops)
- Unbounded Consumption: Bounding Agent Resource Use Against DoS and Denial-of-Wallet
- Unix CLI as the Native Tool Interface for AI Agents
- Utility-Model Split: Background Tasks on a Cheaper Model
- Validating Token-Optimized Formats Inside Agentic Loops
- Variance-Based RL Sample Selection
- Verification Surface: Match the Tool to the Failure
- Voting / Ensemble Pattern for AI Agent Development
- Within-Task Model Cascade: Designing the Escalation Gate
- pass@k and pass^k: Capability and Consistency Metrics
cursor¶
- Cursor /multitask: Async Subagent Dispatch in the Editor
- Cursor 3 Agents Window: Parallel Agents and Worktree Isolation
- Cursor Automations: Event-Triggered Agents and /automate
- Cursor Customize Page: Unified Surface for Agent Primitives
- Cursor Multi-Root Workspaces for Cross-Repo Agent Edits
- Cursor SDK: Programmable TypeScript Agent Runtime
- Cursor Self-Hosted Cloud Agents
- Cursor for AI Agent Development
- Enterprise-Managed Plugin Governance for Agent CLIs
- Multi-Repo and No-Repo Coding Agent Automation Templates
- PR-Subscribed Agent Ownership: The Agent That Opened the PR Drives It to Green
- Per-Change Deploy Monitors: Report the Verdict, Don't Act on It
- Project-Scoped Agent Workspace: Durable Context for Clean-Context Subagents
- Public Rules-File Corpora as Evidence
- Renting a Cloud Agent's Execution Sandbox
- Self-Improving Code Review Agents — Learned Rules
- Static-Context Tool Residence: Two Overrides on Hit Rate
- Tiled Agent Layout: Supervising Parallel Agents Through Dedicated Panes
- Tunable Effort Levels for Code Review Agents
- Visual-Prompt Agent Steering (Cursor Design Mode)
evals¶
- AX Evals: Measure the Agent-Facing Surface, Not the Model
- Action-Class Decomposition for Tool-Calling Evals
- Action-Graded Severity for Agent Red-Team Outcomes
- Adaptive Validation Task Selection
- Agent Development Lifecycle for Agent Products
- Agent Harness: Initializer and Coding Agent Pattern
- Agent-Authored Eval Suites From Repo Context and Traces
- Agent-Driven Eval Flywheel: Prove a Fix Generalizes
- Agentic-Agile: Adapting Agile Rituals for Agent Work
- Ambiguity Stability as a Model-Selection Criterion
- Answer-Reachable Eval Environments
- Anti-Reward-Hacking: Rubrics That Resist Gaming
- Audit the Noise Floor Before Trusting a Benchmark Gap
- Behavior Specs: Grading the Trajectory, Not the Result
- Behavioral Drivers of Coding Agent Success and Failure
- Behavioral Testing for Non-Deterministic AI Agents
- Benchmark Contamination as Eval Risk
- Benchmark-Driven Tool Selection for Code Generation
- Building Agent Eval Environments With a World Spec
- CARE: Three-Party Stage-Gated Engineering of LLM Agents
- Canary Tools for Diagnosing Tool-Selection Reasoning
- Choosing the Judge Model That Grades Your Agent Evals
- CoT Robustness in Code Generation
- Comparative Judging for Agent Configuration Ranking
- Completion Failure Taxonomy: Why Code Suggestions Miss
- ComplexMCP: Three Bottlenecks in Large Interdependent Tool Sandboxes
- Constraint Decay in Backend Code Generation
- Contract-Domain Tracing for Rubric Credit
- Control Lexical Leakage in Agent-Memory Retrieval Evals (Entity-Collision)
- Controlled Benchmark Rewriting for Agent Safety Judgment
- Corpus-Level Trace Diagnostics for LLM Agents
- Cost-Quality Pareto Measurement for Agent Configurations
- Coverage-Guided Fuzzing for Multi-Agent LLM Systems (FLARE)
- Cross-Framework Signal Semantics: Re-Measure Borrowed Trajectory Rules
- Decision-Fork Replay: Grading an Agent's Mid-Run Choices
- Decomposed Red-Teaming for Agent Monitors
- Decomposing Agent Output Variability by Layer (Sampling vs Orchestration State)
- Detecting Self-Preference in a Single LLM Judge
- Distillation-Induced Similarity Metrics for Tool-Use Agents
- Dominator-Graph Trajectory Invariants for Non-Deterministic Agents
- Emulate Agent-Experience Changes Before Shipping
- Emulated APIs for Agent Skill Evals
- Equivalence Testing for Agent Configuration Changes
- Eval Awareness: Designing Evals Agents Cannot Recognize
- Eval Blind Spots: Structural Gaps in Measurement Methodology
- Eval Difficulty as a Product Smell
- Eval Engineering (Training Module)
- Eval Environment Containment for Cyber-Capable Agents
- Eval Strategy by Agent Generation: A Structure-to-Eval Locator
- Eval-Driven Development Training for AI Agent Teams
- Evaluator Templates: Portable Primitives for Agent Eval Suites
- Evaluator-Optimizer Pattern for AI Agent Development
- Four Reporting Levels for Agent Working Memory Evaluation
- Frozen-Base Task Mining for Repository Instruction Files
- Frozen-Stimulus Panels for Cross-Vendor Behavior Measurement
- Gate Best-of-k Selection on Compliance Before Score
- Golden Query Pairs as Continuous Regression Tests for Agents
- Grade Agent Outcomes, Not Execution Paths
- Grading Strategies for Eval-Driven Development
- Hardening Agent Evals for Production-Grade Reliability
- Harness Hill-Climbing: Eval-Driven Iterative Improvement of Agent Harnesses
- Head-to-Head Evaluation of Competing MCP Servers
- Human-Review-Driven Curation of Golden Eval Datasets
- Incident-to-Eval Synthesis: Production Failures as Evals
- Inference-Time Tool-Call Reviewer: Pre-Execution Feedback for Tool-Calling Agents
- Inferring Agent Failure from Conversation Evidence (Perceived Error)
- Isometric Harness Ablation: Rank Subsystem Investment by Removing One at a Time
- L3 → L5: Reaching Agent-First
- LLM API Fault Injection at the HTTP Layer (AgentChaos)
- LLM-Driven Benchmark Auditing
- LLM-as-Judge Evaluation with Human Spot-Checking
- Learned Prefix Monitors for Agent Traces
- Macro Evals for Agentic Systems: Population-Level Behavior Patterns
- Markov-Chain Reliability for LLM Agents: Audit the Abstraction Before You Trust the Metric
- Measure the Judge Before You Freeze a Gate on It
- Measuring Synthetic Eval Data Quality (SynAE)
- Meta-Evaluate the LLM Judge Before Trusting Rubric Verdicts
- Multi-Run, Shuffled-Order Evaluation for Self-Improving Agents
- Multi-Turn Conversation Evaluation: Per-Turn and Trace-Level Scoring Together
- Mutation Testing as a Quality Gate for AI-Generated Test Suites
- Mutation Testing for LLM Judges: Scoring an Evaluator on Injected Defects
- Nonstandard Errors in AI Agents: Model-Family Variance
- Observability-Driven Harness Evolution
- Overeager-Behavior Elicitation: Scope + Trap Fragments as a Diagnostic for Out-of-Scope Tool Calls
- PASS@(k,T): Evaluate RL for Agents Along Sampling and Interaction Depth
- Per-Attempt Sandboxes for Agents That Change the Filesystem
- Perceived Model Degradation: Why Vibes Are Not Evals
- Plan Compliance in Agents: Measure What They Execute, Not What You Wrote
- Planted-Bug Methodology: Deliberate Bugs as Observability Calibration
- Policy-Graded Evaluation of Coding Agents
- Pre-Generation Complexity Scoring for Code Reliability
- Precise Debugging: Measure Edit Precision, Not Just Test Pass Rate
- Profile Your Agent Test Suite Against Measured Practice
- Purpose-Built Eval Suites for Model and Harness Swaps
- RAG/Agent Reliability Problem Map: 16-Domain Failure Taxonomy
- Rank Resolution: Reading a Converged Coding-Agent Leaderboard
- Reasoning Retention and Compaction as Harness Settings
- Recover the Six Measurement Choices Behind an Attack Success Rate
- Red-Team Your Blocking Monitor Before You Trust It
- Reliability of an Automatically Selected Agent Harness
- Rented Sandboxes for Coding-Agent Benchmark Runs
- Repository Perturbation as Context-Reasoning Diagnosis (RepoMirage)
- Review-Comment-Derived Benchmarks for Code Review Agents
- Seed-Variance Reporting and Measurable-Range Eval Design
- Size Agent Comparisons by Run-to-Run Variance
- Skill Eval Loop
- Skill Evals: Measuring Skill Quality as a Dataset-Graded Unit
- Skill Lift: Measuring What a Skill Adds at Runtime
- Skill Specification Violation Fuzzing
- Skill Test Coverage as a Release Gate
- Skill-Use Gates: Trigger, Compliance and Boundary
- Specification-Path Testing: Same Contract, Different History
- Stakeholder Trust Through Evals and Observability
- Stateful Agent Evals via State Snapshots and Transition Assertions
- Static Difficulty Estimation for Agent Issue Triage
- Step-by-Step: Building Your First Eval-Driven Feature
- Task Alignment: The Selective-Compliance Gap Benchmarks Miss
- Test Harness Design for LLM Context Windows
- The Consistent Capability Fallacy in LLM Agent Design
- The Eval-First Development Loop for AI Agent Features
- The Synthetic Ground Truth Fallacy in Agent Evaluation
- The Test Homogenization Trap: When LLM-Generated Tests Mirror Model Blind Spots
- Tool-Use Sim-to-Real Perturbation Taxonomy
- Traces Need Feedback to Power Learning
- Trajectory Decomposition: Diagnose Where Coding Agents Fail
- Variance-Based RL Sample Selection
- What Evals Are and Why AI Agents Need Them for Quality
- Writing Your First Agent Evaluation Suite from Scratch
- pass@k and pass^k: Capability and Consistency Metrics
fallacies¶
- Chain-of-Thought Reasoning Fallacy: Traces Are Not Truth
- LLM Comprehension Fallacy: When Models Seem to Understand
- Reference: Standards, Human Factors, Emerging, and Fallacies
- The AI Knowledge Generation Fallacy: LLMs Recombine, Not Invent
- The Consistent Capability Fallacy in LLM Agent Design
- The LLM Laziness Deficit Fallacy: Restraint Comes From Harness, Not Instruction
- The Model Preference Fallacy in Comparison Content
- The Synthetic Ground Truth Fallacy in Agent Evaluation
- The Task Framing Irrelevance Fallacy in Agent Prompting
- The Token Price Index Fallacy in Agent Cost Planning
frameworks¶
- Agentic Framework Landscape: When Each Framework Fits
- Brownfield to Agent-First: Repo Maturity Framework
- Cognitive Architectures for Language Agents (CoALA): A Classifier for Agent Harnesses
- Consistent-format customer capture
- Cross-Functional Knowledge Artifacts
- Frameworks
- Functional folder taxonomy
- L0 → L1: Making the Repo Readable
- L1 → L2: Adding Feedback Loops
- L2 → L3: Building Mechanical Enforcement
- L3 → L5: Reaching Agent-First
- Natural-language git
- Plan files as resumable artifacts
- Plan mode for knowledge artifacts
- Self-Explanation Loop
- Team OS: Coding-Agent Repo as Cross-Functional Team Brain
geo¶
- AI Crawler Policy: robots.txt for the Three-Tier Crawler Landscape
- Agent-Readiness Discovery Surfaces for Docs Sites
- Answer-First Writing: Structure Content for AI Retrieval
- Assertion Density — Stats and Quotes Over Vague Claims
- Atomic Pages and Chunking — One Concept Per Page for RAG
- GEO for Technical Docs: Developer Documentation Checklist
- Generative Engine Optimization for Developer Sites
- Google Search Console Monitoring Workflow
- How AI Engines Cite — ChatGPT, Perplexity, Claude, Gemini
- Measuring GEO Performance for AI Search Visibility
- SEO vs GEO — How Signals and Metrics Differ
- Schema and Structured Data for GEO — AI Citation Guide
- Separating Exposure From Selection in GEO Measurement
- Topical Authority — Entity Coverage for AI Citation
- What is GEO — Generative Engine Optimization Defined
- llms.txt: Full Specification, Adoption, and Limitations
github-actions¶
- AI Bot CI/CD Workflow Reliability by Agent
- Agent Environment Bootstrapping for AI Agent Development
- Claude Code --bare Flag
- Closed-Loop CI Failure Remediation with Cloud Coding Agents
- Continuous Triage: Automating Issue Classification with AI Workflows
- GitHub Agentic Workflows for Automating Dev Processes
- GitHub Models in Actions for AI-Driven CI Workflows
- Headless Claude in CI: Using -p and --max-turns for Safe Pipeline Integration
- One-Click CI Auto-Fix: Human-Triggered Cloud-Agent Remediation for Failing GitHub Actions
harness-engineering¶
- AX/UX/DX Triad: Three Experience Layers in Agent Systems
- Agent Harness: Initializer and Coding Agent Pattern
- Choosing an Integration Layer for an Embedded Agent Harness
- DSLs as a Constraining Harness for LLM Code Generation
- Fleet Harness Attribution: Pinning the Model to Compare Whole Harnesses
- GitHub Copilot: Harness Engineering for Agent-Ready Code
- Goal Monitoring and Progress Tracking for Long-Running Agents
- Golden Journeys: Restartability as a First-Class Verification Primitive
- Harness Bug Detection Patterns
- Harness Composition for Scaled Security Audits
- Harness Design Dimensions and Archetypes
- Harness Engineering (Training Module)
- Harness Engineering for Building Reliable AI Agents
- Harness Preflight Doctor Command for Agent Diagnostics
- Isometric Harness Ablation: Rank Subsystem Investment by Removing One at a Time
- Method Map: Failure-Mode to Smallest-Artifact Triage
- Observability-Driven Harness Evolution
- Per-Model Harness Tuning: Treating the Backing Model as a Harness Variable
- Prompt-Only Baseline Before a Specialized Agent Subsystem
- Quality Score Rubric and Simplification Log for Agent Harnesses
- Review-Feedback-to-Rule Loop: Promoting Recurring PR Comments into Harness Rules
- Rigor Relocation: Engineering Discipline with AI Agents
- Running Several Coding Agents Behind One Harness Interface
- Separation of Knowledge and Execution in Agent Systems
- Situated Harness Layers: Fix at the Layer That Owns It
- Suspect the Harness Before the Model on a Regression
- Temporary Compensatory Mechanisms in Agent Harnesses
- The AX Stack: A Layered Model of an AI Coding Agent's Prompt-to-Compile Path
human-factors¶
- AI Abundance Reshapes Software Engineering Identity
- AI Adoption Footprint: The Segmented Shape of Engineering Orgs
- AI Label as Reviewer Attention Redistribution
- Absorbing Entry-Level Work into Senior-Agent Workflows
- Adapting AI Assistants to Developer Interaction Style
- Agent Approval Authority in Code Review
- Agent Context File Evolution: Treating ACFs as Configuration Code
- Agent Governance Policies for AI Agent Development
- Agent Headcount as a Vanity Metric
- Agent PR Volume vs. Value: The Productivity Paradox
- Agent Rewrites Lose Meaning: The Ownership Rule for AI-Assisted Writing
- Agent Time Estimates Are Not Schedules
- Agent-Authored PR Integration: Collaboration Signals That Determine Merge Success
- Agent-Driven Greenfield Product Development from Scratch
- Agent-First Software Design for AI Agent Development
- Agent-Generated Code Maintenance Asymmetry
- Agent-Generated Onboarding Guide as a Durable Artifact
- Agent-Generated Verification Reports: A Structured Round-Trip for Human Review
- Agent-Laundered Bug Reports
- Agent-Operable Interface Design (Affora)
- Agent-Recorded Video Demos as a Verification Artifact
- Agentic Education: Persona Progression for Teaching AI Coding Tools
- Agentic Review Comment Acceptance
- Agentic Skill Decay: Which Capabilities Erode Under Agent Delegation
- Agentic-Agile: Adapting Agile Rituals for Agent Work
- Ambition Scaling: Moving the Target as Model Capability Increases
- Anonymized Customization Metrics in the Copilot CLI
- Approval Gate Granularity in Agent Pipelines
- Artifact-Level Accountability Mapping for Agent Workflows
- Ask-Everything Permission Policies Protect Less than Per-Action Approval
- Audit-Budget Allocation for Agent Fleets
- Author-to-Reviewer Role Inversion in AI-Assisted Teams
- Blaming the Model for Scaffolding-Driven Quality Regressions
- Brownfield to Agent-First: Repo Maturity Framework
- Canvas as Durable Workflow State: The Four-Step Blueprint and What It Costs
- Capturing Dismissal Reasons for Agent Review Findings
- Cargo Cult Agent Setup: Copying Without Understanding
- Chain-of-Thought Reasoning Fallacy: Traces Are Not Truth
- Channels Permission Relay
- Classical SE Patterns as Agent Design Analogues
- Claude Code Auto Mode: Classifier-Based Permission Gating
- Coding-Agent Reversibility: Platform Choice as a Two-Way Door
- Cohort Segmentation in the Copilot Usage Metrics API
- Completion Summary as the Oversight Surface
- Conceptual Integrity Erosion in Agent-Built Codebases
- Concurrent Agent Pull Requests and Merge-Conflict Cost
- Confirmation Gates for Consequential Agent Actions
- Convenience Loops and AI-Friendly Code in Your Stack
- Conversation Registers for AI Coding Sessions
- Copilot vs Claude Billing Semantics for Enterprise Teams
- Criticality and Containment: Scoping Which Agent Code You Read
- Cross-Functional Knowledge Artifacts
- Cross-Tool Translation: Learning from Multiple AI Assistants
- Delegated-Autonomy Boundary Artifacts (AJR and ADP)
- Delegating Change Descriptions to the Agent
- Deliberate AI-Assisted Learning: Accelerating Skill Acquisition
- Density-Normalized Quality Metrics Mask AI-Driven Code Growth
- Developer Control Strategies for AI Coding Agents
- Developer as CPU Scheduler: Attention Management with Parallel Agents
- Direct Prompt Injection via Collaboration (User as Attack Vector)
- Do Not Price the Rules in Your Agent Instruction File
- Earned-Complexity Agent Maturity Ladder
- Ecosystem-Level Integration Friction Governance
- Editor and Manager Surface Separation in Agent IDEs
- Empowerment Over Automation for AI Agent Development
- Encoding Tacit Knowledge into Agent Improvement Loops
- Encoding Values in AGENTS.md: Why Prose Without Verification Fails
- Enforced Versus Advisory Controls in LLM-Native IDEs
- Enterprise Skill Marketplace: Distribution and Quality
- Evaluating Agent Patterns Catalog as a Source
- Evidence-First Reports From Failure-Diagnosis Agents
- Factory Over Assistant: Orchestrating Parallel Agent Fleets
- Fallacies for AI Agent Development
- From Preventive to Reactive: Front-Loading Security in AI Coding Prompts
- GitHub Copilot: Team Adoption and Governance Guide
- How Teams Build SE Agents: A Seven-Stage Build Loop
- Human Impact of AI Agents on Developer Teams and Workflows
- Human-Equivalent Hours for Autonomous Coding Agent Productivity
- Human-Facing Docs in the Agent Era: Mental Models Over Reference
- Human-in-the-Loop Placement: Where and How to Supervise
- Humans and Agents in Software Engineering Loops
- Hyper-Personalized Software: The Return of RAD
- Initiatives and Community: Tracking the Agentic Engineering Landscape
- Inline Suggestion Attachment in Agent Code Review
- Intent-Centric Engineering: Oversight Over Authorship
- Intervention Rate as a Diagnostic North Star, Not a Target
- Judgment Relocation: Where Human Decisions Land in an Agent Factory
- LLM Comprehension Fallacy: When Models Seem to Understand
- LLM Refactoring Adoption Patterns
- LLM Support During the First Detection Pass
- Language Selection Scored on Review Cost
- Late Requirement Arrival in Agent Sessions
- Law of Triviality in AI PRs for AI Agent Development
- Lay the Architectural Foundation by Hand Before Delegating
- Managing Cognitive Load and AI Fatigue for Sustainable Agent Use
- Marking Which Artifacts Are for Humans or Agents (Landmarking)
- Monitor or Wait: The Supervision Choice During Agent Execution
- Name the Check That Passed Before Accepting AI Code
- Natural-Language Documentation as a Code-Review Intermediate (Verifiable Literate Programming)
- Natural-language git
- Next Edit Suggestions Carry Context You Never Curated
- Nonstandard Errors in AI Agents: Model-Family Variance
- Org-Membership-Gated Agent Entitlement
- Overtrusting Human Sign-Off on Generated Assertions
- PM on the AI Exponential
- PR Description Style as a Lever for Agent PR Merge Rates
- PR Scope Creep as a Human Review Bottleneck
- Parallel Agent Sessions Shift the Bottleneck from Writing
- Per-Agent-App Attribution in the Copilot Usage Metrics API
- Per-Task Verification Budget: Size the Task to Fit the Check
- Plan Mode: Read-Only Exploration Before Implementation
- Plan files as resumable artifacts
- Plan mode for knowledge artifacts
- Plugin Component Co-Change: Scripts and Their Instructions Move Together
- Polya Small-Steps: Using AI to Think Better, Not Think Less
- Porting a Coding-Agent Harness Beyond Engineering
- Pre-Execution Risk Classification for Terminal Commands
- Predicting Reviewable Code: Pre-Flagging Functions Reviewers Will Delete
- Pressuring a Coding Agent Degrades the Code It Writes
- Process Amplification: Scaling Human Work with Agents
- Programming Language Choice Still Shapes Agent Artifacts
- Progressive Autonomy: Scaling Trust with Model Evolution
- Proof of Presence: Re-Authenticating for Agent Actions
- Public-Channel Agent Work as Lehrwerkstatt for Team Learning
- RAMP: Committed AI Configuration and the Quality Cost
- Re-Run an Agent's Speed-Up Claim Before Merging
- Reading Copilot Feature Engagement by Its Threshold
- Reading Visible Edge-Case Handling as a Security Check
- Reading a Coding-Agent Vendor's Security Certificate
- Reading a Vendor-Computed AI Coding ROI Dashboard
- Reference: Standards, Human Factors, Emerging, and Fallacies
- Reviewer Habituation in Agent PR Review
- Rigor Relocation: Engineering Discipline with AI Agents
- Risk Architecture for AI-Native Engineering Teams
- Rolling Out CLI Coding Agents at Organization Scale
- Rolling Out a Team-Embedded Agent Like a Tool
- Seamless Background-to-Foreground Handoff
- Selective Autonomy from Copilot Feedback
- Self-Explanation Loop
- Skill Atrophy: When AI Reliance Erodes Developer Capability
- Skill Library Refinement Loops: Organizational Feedback for Shared Skills
- Stakeholder Trust Through Evals and Observability
- Stated-Understanding Checks: Asking the Agent to Correct You
- Steering Running Agents: Mid-Run Redirection and Follow-Ups
- Step Budgets and Trust in Agent-Generated Code Tours
- Strategy Over Code Generation: Why AI Speed Doesn't Fix Wrong Goals
- Suggestion Gating: Fewer Completions, Better DX
- Tab-Accept Rate as a Proxy for Critical Engagement
- Team OS: Coding-Agent Repo as Cross-Functional Team Brain
- Team Onboarding for AI Agent Workflows and Adoption
- Team Shared-Language Desync from Removed Review Friction
- The AI Development Maturity Model: From Skeptic to Agentic
- The AI Knowledge Generation Fallacy: LLMs Recombine, Not Invent
- The AX Stack: A Layered Model of an AI Coding Agent's Prompt-to-Compile Path
- The Addictive Flow State of Agent-Assisted Development
- The Anthropomorphized Agent for AI Agent Development
- The Bottleneck Migration When Humans Supervise Agents
- The Citizen-Agent-Expert Operating Model for AI Coding
- The Consistent Capability Fallacy in LLM Agent Design
- The Context Ceiling -- Where AI Fails Expert Architects
- The Effortless AI Fallacy for AI Agent Development
- The First Edit Predicts Whether an AI Completion Survives
- The Meat Proxy: Relaying Agent Output Without Reading It
- The Model Preference Fallacy in Comparison Content
- The Productivity-Experience Paradox in AI-Assisted Development
- The Prompt Tinkerer Anti-Pattern in Agent Workflows
- The Software Factory Model: Industrializing Agent Loops
- The Synthetic Ground Truth Fallacy in Agent Evaluation
- The Task Framing Irrelevance Fallacy in Agent Prompting
- Tiled Agent Layout: Supervising Parallel Agents Through Dedicated Panes
- Tool Confirmation Carousel: Batched UI for Per-Call Approvals
- Tool Preamble: User-Visible Status Updates Before Tool Calls
- Transcript-Measured Review Coverage
- Velocity-Quality Asymmetry: Why AI Speed Gains Fade
- Verification-Centric Development for AI-Generated Code
- Verify Agent Diagnoses and Fix Proposals Before Acting
- Vibe Coding: Outcome-Oriented Agent-Assisted Development
- Visible Thinking in AI-Assisted Development
- When Developers Understand Less of Their Own Codebase
index¶
- AI Agent Development Anti-Patterns and Failure Modes
- Agent Design Patterns and Architectures for AI Agents
- Agent Patterns for AI Agent Development
- Agentic Code Review Patterns and Review Architectures
- Brownfield to Agent-First: Repo Maturity Framework
- Claude Code for AI Agent Development
- Context Engineering: Shaping AI Agent Input and Attention
- Continuous AI: A Navigation Map of Always-On Agent Workflows
- Cursor for AI Agent Development
- Emerging Concepts for AI Agent Development
- Eval-Driven Development Training for AI Agent Teams
- Fallacies for AI Agent Development
- Foundational Disciplines for AI-Assisted Development
- Frameworks
- Generative Engine Optimization for Developer Sites
- GitHub Copilot Training Modules for Engineering Teams
- GitHub Copilot for AI Agent Development
- Human Impact of AI Agents on Developer Teams and Workflows
- Instructions: System Prompts, Rules, and Agent Configuration
- Loop Engineering: Designing Agent Loops That Converge
- Multi-Agent Systems: Coordination and Orchestration
- Observability for AI Agents: Tracing and Debugging
- Open Standards and Protocols for AI Agent Development
- Patterns: Agent Design, Multi-Agent, and Anti-Patterns
- Security for AI Agent Development
- Team OS: Coding-Agent Repo as Cross-Functional Team Brain
- Token Engineering: Fewer, Cheaper Tokens Without Losing Quality
- Tool Engineering: Designing and Managing AI Agent Tooling
- Tools: Claude Code, Cursor, and GitHub Copilot
- Training Modules
- Verification: Testing, Evals, and Guardrails for Agents
- Workflows for AI Agent Development
instructions¶
- @import Composition Pattern for Agent Instruction Files
- AGENTS.md Design Patterns for Effective Agent Files
- AGENTS.md as a Table of Contents, Not an Encyclopedia
- AGENTS.md: Project-Level README for AI Coding Agents
- Accumulated Behavioral Rules from Review Feedback
- Acknowledged-Debt Ledger with Next-Trigger Conditions
- Advisory Prompts Distilled from Reasoning Traces
- Against-Prior Accuracy: Score the Rules That Fight Defaults
- Agent Config as a Managed Supply Chain: Hashing and Pinning
- Agent Context File Evolution: Treating ACFs as Configuration Code
- Agent Debugging: Diagnosing Bad Agent Output
- Agent Pushback Protocol for Managing Disagreements
- Agent-Ready Bug Reports for Software Repair Agents
- Architecting a Central Repo for Shared Agent Standards
- Assuming a CLAUDE.md Security Rule Is Enforced
- Authority Confusion: Untrusted Context Must Not Authorize Side Effects
- Bootstrapping Coding Agents: The Specification Is the Program
- Boring Technology Bias: When Agents Recommend by Popularity
- CLAUDE.md Convention for Structuring Agent Instructions
- Cargo Cult Agent Setup: Copying Without Understanding
- Catastrophic Remembering: Instruction Files That Only Grow
- Classifier-Gated Auto-Permission for Cloud-IDE Coding Agents
- Claude Code Extension Points: When to Use What
- Claude Code Hooks: Deterministic Lifecycle Automation
- Close the Attack-to-Fix Loop: Adversarially Train Agent
- Codified Effort and Escalation Policy in the Instruction File
- Compound Prompt Constraints Degrade More Than Their Parts Predict
- Configuration File Structure Does Not Drive Compliance
- Configuration Smells in AGENTS.md Files (Six-Smell Catalog)
- Constraint Degradation in AI Code Generation
- Constraint Encoding Does Not Fix Constraint Compliance
- Constraint Preambles and the Gain Your Scanner Misses
- Constraint-Evasive Fabrication in Instruction Sets
- Content Exclusion Gap: AI Security Boundaries by Mode
- Content-Addressed Agent Configurations (Deterministic Control Plane)
- Context Priming: Pre-Loading Files for AI Agent Tasks
- Continuous Agent Improvement: Iterating on Agent Quality
- Contractual Skill Files: Inspectable SKILL.md for Enterprise Agents
- Controlling Agent Output: Concise Answers, Not Essays
- Convention Over Configuration for Agent Workflows
- Copilot App Customize Tab vs Repo Configuration Files
- Cost-Aware Skill Rewriting: Preserve Operational Anchors, Not Skill Tokens
- Critical Instruction Repetition via Primacy and Recency
- Cursor Customize Page: Unified Surface for Agent Primitives
- Daily-Use Skill Library: Encoding Your Process as Agent Skills
- Deferred Standards Enforcement via Review Agents
- Designing Agent Tools Like APIs
- Designing Agents to Resist Prompt Injection
- Destyling Untrusted Input as a Prompt Injection Defense
- Diagram as the Shared Spec: One Artifact for the Picture and the Prompt
- Discoverable vs Non-Discoverable Context for Agents
- Do Not Price the Rules in Your Agent Instruction File
- Documentation Read Counts Measure Retrievability, Not Value
- Domain-Specific System Prompts with Concrete Examples
- Dynamic System Prompt Composition
- Empirical Baseline: Agentic AI Coding Tool Configuration
- Encode Project Conventions in Distributed AGENTS.md Files
- Encoding AI Writing Tells as a Prose Style Contract
- Encoding Product-Design Taste into Agent Context
- Encoding Values in AGENTS.md: Why Prose Without Verification Fails
- Enforcing Agent Behavior with Hooks
- Evaluating AGENTS.md: When Context Files Hurt More Than Help
- Event-Driven System Reminders for AI Agent Development
- Example-Driven vs Rule-Driven Instructions
- Feature List Files for Reliable AI Agent Development
- Five-Stage Policy Layer Typology for Generalist Agents
- Foundations: Context Engineering and Instructions
- Framing Subagent Returns So They Cannot Act as Instructions
- Frontmatter and Body Rule Drift in Agentic Workflows
- Frozen Playbook Reuse Without Target-Side Validation
- Frozen Spec File: Preserving Intent in AI Agent Sessions
- Frozen-Base Task Mining for Repository Instruction Files
- Functional folder taxonomy
- GROUNDING.md: Field-Scoped Hard Constraints and Convention Parameters
- Getting Started: Setting Up Your Instruction File
- GitHub Copilot App Slash Commands and What They Change
- GitHub Copilot Custom Agents and Skills Extensibility Guide
- Goal Recitation: Countering Drift in Long Sessions
- Goal Reframing: The Primary Exploitation Trigger for LLM Agents
- Google ADK Skills: Portable SKILL.md Across ADK Agents
- Grill Me: Developer-Initiated Plan Interrogation
- Grounding Agents in Code the Model Has Never Seen
- Guardrails Beat Guidance: Rule Design for Coding Agents
- HTML as Agent Output Format: When to Ask for HTML Instead of Markdown
- Hard-Deny Classifier Rule: Unconditional Block in Auto Mode
- Heuristic-Based Effort Scaling in Agent System Prompts
- Hierarchical CLAUDE.md: Structuring Context Files at Multiple Levels
- Hints Over Code Samples in Agent Prompts
- Hook Catalog for Claude Code Enforcement
- Hooks for Enforcement vs Prompts for Guidance: When to Use Each
- How the Four Agent Engineering Disciplines Compound
- Instruction Polarity: Positive Rules Over Negative
- Instruction-Aware Automated Code Review
- Instructions: System Prompts, Rules, and Agent Configuration
- Interactive Clarification for Underspecified Tasks
- Iterative Binary Feedback for Pattern Adherence
- Knowledge Gap or Skill Gap: Triage Before Writing Context
- Layer Agent Instructions by Specificity Across Scopes
- Listener-State Naming for User-Invoked Agent Skills
- Living-Docs-Grounded Agent Design Conversations
- Managed Settings Drop-In Directory: Enterprise Policy Fragmentation
- Marking Which Artifacts Are for Humans or Agents (Landmarking)
- Match Architecture Spec Format to Model Capability
- Match Tool Instructions to the Agent Workflow
- Mermaid as Agent Output Format: When to Ask for a Diagram Instead of Prose
- MessageDisplay Hook: Transforming Assistant Text at the Display Boundary
- Method Map: Failure-Mode to Smallest-Artifact Triage
- Multi-Layer Specification Redundancy as a Robustness Budget
- Natural-Language Customization Bootstrap
- Negative Space Instructions: What NOT to Do in Agent Prompts
- Override Pattern: Reusing Interactive Commands in Automated Pipelines
- Per-Step Preconditions and Postconditions in Skill Files
- Per-Subagent Instruction Inheritance (omitClaudeMd)
- Permission-Gated Custom Commands for AI Agent Development
- Permutation Frameworks for Batch Code Generation
- Persona-as-Code: Defining Agent Roles as Structured Docs
- Personalized vs Generic Agent Skills: Where Effort Pays
- Plan Compliance in Agents: Measure What They Execute, Not What You Wrote
- Plugin Component Co-Change: Scripts and Their Instructions Move Together
- Policy File Validation: Catching Silent Non-Enforcement
- Portable Agent Definitions: Full-Stack Identity as Code
- Post-Compaction Re-read Protocol for Agent Continuity
- PostToolBatch Hook: Once-Per-Decision-Cycle Injection at the Batch Boundary
- PostToolUse continueOnBlock: Refusal With a Load-Bearing Reason
- Pre-Trust Execution Surface in Coding Agent Harnesses
- Probe-and-Refine Tuning of Repository Guidance for Coding Agents
- Production System Prompt Architecture and Techniques
- Project Instruction File Ecosystem
- Project Writing Skill: House Style as Model-Invocable Skill
- Prompt Debt: Hand-Tuning Natural-Language Prompts as Technical Debt
- Prompt Engineering for Agent Instructions and Systems
- Prompt File Libraries for Reusable Agent Instructions
- Prompt Governance via PRs: Reviewable AI Behavior
- Prompt Layering: How Instructions Stack and Override
- Prompt Transpilation: Instructions as Build Artifacts
- Prompt-Only Tool Access Control
- Prompt-Rewrite Discipline on Cross-Generation Model Migration
- Protecting Sensitive Files from Agent Context Access
- Public Rules-File Corpora as Evidence
- Publishing Agent Instructions Outside the Repo (design.md)
- RAMP: Committed AI Configuration and the Quality Cost
- Re-Auditing Context Engineering Across Model Generations
- Reflective Prompt Evolution with Pareto Selection (GEPA)
- Repairing Agent Prompts from Trace Contrast, Not Search
- Repository Bootstrap Checklist: Wiring Agent Support
- Repository Skill Release Drift
- Requirement Smells: No Category Signal, Density Unconfirmed
- Restraint Rules Need External Enforcement
- Review-Feedback-to-Rule Loop: Promoting Recurring PR Comments into Harness Rules
- Rule Lifecycle Metadata for Prunable Instruction Surfaces
- Runtime Guard as an Installed Skill (Defense-as-Skill)
- SKILL.md Frontmatter Reference: All Fields Explained
- Scheduled Instruction File Fact-Checker for Accuracy
- Scoped Credentials via Proxy Outside the Agent Sandbox
- Security Constitution for AI Code Generation
- Security Knowledge Priming for Code Generation (SPARK)
- Seeding Agent Context: Breadcrumbs in Code
- Self-Explanation Loop
- Semantic Collapse Under Underspecified Prompts
- Shared Context Bundle Registry for Agent Teams
- Skill Authoring Patterns: Description to Deployment
- Skill Authoring as Software Engineering: What Transfers
- Skill File Linting: Which Three Checks to Run First
- Skill Packs: Registry Distribution Needs Pinning Discipline
- Skill Program Functions: Executable Guardrails Compiled From Past Failures
- Skill Tool as Enforcement: Loading Command Prompts at Runtime
- Skill as Instruction Surface and Callable API (Interpreter Skills)
- Skill as Knowledge Pattern for AI Agent Development
- Skill-Use Gates: Trigger, Compliance and Boundary
- Spec Complexity Displacement: When Specs Become Code
- Spec-Driven Development with Spec Kit
- Spec-Driven Test Generation: Contract Coverage Is the Lever
- Specialized Agent Roles for Effective AI Pipelines
- Specification Portability Across Coding Agents
- Specification-Grounded Test Writing
- Stage-Targeted Prompt Structure for Pull Request Outcomes
- Stale AI Configuration Artifacts (Context Rot)
- Standard-Grounded NFR Specs: Quality Up, Correctness Flat
- Standards as Agent Instructions for AI Agent Development
- System Prompt Altitude: Specific Without Being Brittle
- System Prompt Delivery Channels on Shared Runners
- System Prompt Replacement for Domain-Specific Agent Personas
- System Prompt as Secret Store (OWASP LLM07)
- Task List Divergence as Instruction Quality Diagnostic
- Team-Scoped Agent Policy Delegation
- The Error-Class Governance Loop for Instruction Libraries
- The Implicit Knowledge Problem for AI Coding Agents
- The Instruction Compliance Ceiling: How Rule Count Limits AI
- The LLM Laziness Deficit Fallacy: Restraint Comes From Harness, Not Instruction
- The No-Op Test: Prune Agent Docs by Behavior, Not Length
- The Prompt Tinkerer Anti-Pattern in Agent Workflows
- The Specification as Prompt: Existing Artifacts as Agent
- Three Knowledge Tiers: Sourced, Unverified, Hallucinated
- Throwaway-Prototype Skill: Build to Discard, Keep Only the Answer
- Token Preservation Backfire for AI Agent Development
- Tool Minimalism and High-Level Prompting
- Treat Task Scope as a Security Boundary
- Typed Pseudocode for Skill Libraries (Skill-as-Pseudocode)
- Ubiquitous Language for AI Plans
- Unversioned Scaffolding Commands Pull Stale Templates
- Usability Pressure as a Silent Security-Regression Vector
- Use a Public-Web Index to Gate Automatic URL Fetching
- WRAP Framework for Writing Agent-Ready Issue Descriptions
- Which Task You Delegate Changes Poisoned-Repo Exposure
- Why an Encoded Rule Still Fails After a Passing Eval
- Workspace-Hosted Skills: Authorship Outside the Repo
- Write Agent Rules You Can Grade From the Transcript
- Write Tool Descriptions as Agent Onboarding Documents
- claudeMdExcludes: Selective Ancestor Instruction-File Exclusion
- copilot-instructions.md as a Repo-Level Instruction Convention
long-form¶
- Advanced Tool Use: Scaling Agent Tool Libraries
- Agent-Authored Messages as a Deferred Exfiltration Channel
- Agentic Pattern Vocabulary Crosswalk
- Auto Model Selection: Harness-Driven Routing per Task
- Classifier-Gated Auto-Permission for Cloud-IDE Coding Agents
- Cloud-Agent Session Bootstrap: Cached Install plus Per-Session Start
- Code-Native Memory Substrates for Coding Agents
- Copilot Memory and Cross-Agent Persistence
- Cost-Aware Agent Design: Route by Complexity, Not Habit
- Cost-Inefficient Behaviors in Coding Agents
- Episodic Memory Retrieval for AI Coding Agent Loops
- Eval Blind Spots: Structural Gaps in Measurement Methodology
- Eval-Driven Development: Write Evals Before Building Agent
- Fleet Harness Attribution: Pinning the Model to Compare Whole Harnesses
- GEO for Technical Docs: Developer Documentation Checklist
- Harness Design Dimensions and Archetypes
- Harness Engineering for Building Reliable AI Agents
- Loop Strategy Spectrum: Accumulated vs Fresh Context
- Memory Retrieval as a Control Decision
- Minimum-Sufficient Control Ladder: Escalate by Failure Mode
- Open Agent School Pattern Mapping for Practitioners
- Parameter-Keyed Caching and Dependency-Aware Parallelism for Plan-Execute Pipelines
- Pattern Selection Map
- Proactive Idle-Time Anticipation (ProAct)
- Production Hosting Topology for Self-Hosted Agent SDK Runtimes
- Prompt Caching: Architectural Discipline for Agents
- Repository-Level Retrieval for Code Generation
- Review-Then-Implement Loop for AI Agent Development
- Six-Shape Approval Response Taxonomy: Beyond Binary Allow/Deny
- Skill Authoring Patterns: Description to Deployment
- Specialized Agent Roles for Effective AI Pipelines
- Stacking Outer Loops Around the Agent
- Symptom-Reduction-as-Root-Cause: Why Oracle Tests Alone Miss Architectural Drift
- Tenant Model Policy: Organization-Scoped Rules for AI Model Selection
- The Context Ceiling -- Where AI Fails Expert Architects
- Verification-Gated Agent Autonomy via Automated Review
loop-engineering¶
- Agent Loop Go/No-Go: When Looping Earns Its Cost
- Agent Loop Middleware — Safety Nets and Message Injection
- Blind Resampling Over Self-Repair in Small Code Models
- Calibrated Early Termination and Warm Restart for Agent Runs (FailFast-RestartSmart)
- Comparison-Only Advisor: Steering a Large Actor With a Tiny Comparator
- Continuation Dispatcher: Who Owns the Iterate-or-Stop Call
- Convergence Detection in Iterative Agent Refinement
- Goal-Driven Autonomous Loop with Budget Cap
- Human-in-the-Loop Checkpoints as Loop Control
- In-Loop Interception: Custom Logic Between the Model Call and the Tool Call
- Loop Budgeting: Allocating Iteration and Token Budget Across Turns
- Loop Engineering: Designing Agent Loops That Converge
- Loop Strategy Spectrum: Accumulated vs Fresh Context
- Loop Trigger Selection: Pairing the Start with the Stop
- One-Shot Record and Deterministic Replay for Periodic Agent Tasks
- Stacking Outer Loops Around the Agent
- Stuck-Loop Recovery: Detecting and Escaping Non-Converging Agent Loops
- The Ralph Wiggum Loop: Fresh-Context Iteration Pattern
- The Three Loops of Agentic Coding: A Diagnostic Vocabulary
- Within-Task Model Cascade: Designing the Escalation Gate
mcp¶
- Agent Approval Laundering: Effects Beyond the Named Command
- Agentic Detection and Response at the MCP Boundary
- Auth-Isolation as the MCP-vs-CLI Selection Heuristic
- Canary Tools for Diagnosing Tool-Selection Reasoning
- Centrally Provisioned MCP Servers: Remote Transports Only
- ComplexMCP: Three Bottlenecks in Large Interdependent Tool Sandboxes
- Customer-Hosted MCP Tunnel: Outbound-Only Connectivity to Private MCP Servers
- Documentation-Grounding MCP Servers for Vendor SDKs
- Five Design Decisions for MCP Servers and Clients
- GitHub Copilot Extensions for AI Agent Development
- GitHub Copilot MCP Integration for AI Agent Development
- Hint-Driven Concurrency for Read-Only MCP Tools
- Hooks Invoking MCP Tools: Closing the Loop Between Policy and Tool Execution
- Judging MCP Capability Readiness by Client Adoption
- MCP Allowlist by Label, Not by Identity (serverName Trap)
- MCP Approval-View Fidelity Gap and Unicode Concealment
- MCP Client Design: Building Robust Host-Side Logic
- MCP Elicitation: Servers Requesting Structured Input Mid-Task
- MCP LLM Sampling: Servers Requesting AI Inference Mid-Tool
- MCP Runtime Control Plane: Policy Evaluation Between Agent and Tool
- MCP Server Design: Building Agent-Friendly Servers
- MCP Tool Result Persistence via _meta Annotation
- MCP alwaysLoad: Classifying Servers as Eager or Just-in-Time
- MCP-vs-CLI Cost Ratios Are a Property of the Scaffolding
- MCP: The Open Protocol Connecting Agents to External Tools
- Multi-Tool Threshold Poisoning Against MCP (ShareLock)
- OAuth Client ID Metadata Documents (CIMD) for MCP Servers
- Per-Server MCP Environment Scoping for Credential Isolation
- Pooled-Evidence Factuality Checks for MCP Agents (Cross-Source Conflation)
- Production MCP Agent Stack: Sequencing Six Decisions into One Deployment
- Proprietary-to-Open-Standard Tool Migration (Copilot Extensions to MCP)
- Push-Event MCP Channels: Inverting the Pull-Tool Polarity
- Scanner-as-MCP-Server: Secret and Dependency Scans as Typed Agent Tools
- Scoped MCP Server Discovery: Most-Specific-Wins Resolution
- Security-Aware Tool Descriptions for MCP Servers (SpellSmith)
- Skill or MCP Server: Choosing a Capability's Delivery Mechanism
- Stateless MCP: One Request per Tool Call
- Tool Signing and Signature Verification for Agents
- Tool-Invocation Attack Surface in Coding Agents
- Tool-Use Sim-to-Real Perturbation Taxonomy
- Vetting Tool Definitions for Exfiltration Signatures
- WebMCP: Browser-Hosted Tool Contracts for In-Page AI Agents
memory¶
- ACID for Agent Repository State
- Agent Memory Patterns: Learning Across Conversations
- Agent Project State Purge: Clean-Slate Session Reset
- Agentic Framework Landscape: When Each Framework Fits
- Auto-Merging a Wiki Agent's Documentation Pull Requests
- Belief Inertia After Tool-Map Drift in AI Agents
- Budgeted Verification of Inherited Agent Constraints
- Claim-Scoped Invalidation for Agent Memory
- Clock-In / Clock-Out Protocol: Bracketed Session Continuity
- CoALA Memory Taxonomy as a Classifier for Harness Artifacts
- Code-Native Memory Substrates for Coding Agents
- Compiled Specialist Agents: Muscle Memory for Recurring Intent
- Component-Isolated Memory Stress Testing for LLM Agents
- Context Lifecycle Management: Beyond Store and Retrieve
- Context-Graph Shared Memory for Multi-Agent Systems
- Continual Learning for AI Agents: Three Layers of Knowledge Accumulation
- Continuation Authority in Agent Migration
- Control Lexical Leakage in Agent-Memory Retrieval Evals (Entity-Collision)
- Copilot Memory and Cross-Agent Persistence
- Cost-Aware Tracing for Skill Distillation
- Cross-Cycle Consensus Relay
- Cue-Anchored Working Memory (Delivery, Not Storage)
- Decentralized Memory for Self-Evolving Multi-Agent Systems
- Detecting Memory-Poisoning Exfiltration by Tool-Call Order (Recall-Before-Send Signature)
- Dormant Memory Payloads Triggered by Sensitive Topics (Trojan Hippo)
- Dual-Trace Memory Encoding: Pair Facts with the Scene They Were Learned In
- Durable Interactive Artifacts: Agent Output Outside the Transcript
- Episodic Memory Retrieval for AI Coding Agent Loops
- Evolving Playbooks: Incremental Context That Preserves Knowledge
- Executable Memory: User State as Code for Personalized Agents
- Experience Graphs as Structured Memory for Self-Evolving Agents
- Experiential-Learning Setup Agents with Snapshot Rollback (SetupX)
- Fact Supersession Memory for Code Assistants
- Forged Reasoning Trace Attacks on Agent Memory (FARMA)
- Generative Agents Memory Stream: Three-Layer Architecture for Long-Running Agent Sessions
- Git-Bound Memory for the Agentic Development Lifecycle
- Harness-Memory Coupling as a Design Axis
- Knowledge Graphs as Provenance-Carrying Agent Memory
- Layered Mutability: Governing Persistent Self-Modifying Agents
- Memory Retrieval as a Control Decision
- Memory Synthesis: Extracting Lessons from Execution Logs
- Memory Transfer Learning: Cross-Domain Memory Reuse in Coding Agents
- Memory-Induced Tool-Drift in LLM Agents
- OpenAI Agents SDK
- Organizing Filesystem Agent Memory for Retrieval Cost
- PEEK: Orientation Cache for Recurring-Context Agents
- Per-Type Retention Policy for Agent Compaction (Knowledge Triage)
- Persistent Teammate Workspace: Durable State for Agent Teams
- Proactive Idle-Time Anticipation (ProAct)
- Project-Scoped Agent Workspace: Durable Context for Clean-Context Subagents
- Query-Conditioned Reuse of Retrieved Agent Trajectories
- RAG over Thinking Traces: Index Reasoning Trajectories Instead of Documents
- Rerunnable Claim Graph as Shared Agent Memory
- Reviewing What a Memory-Fed Autofix Taught
- Scope-Matched Retrieval for Persisted Agent Skills
- Self-Correcting Memory: Evidence-Backed Claim Repair
- Shared Agent Context Store API: When to Expose Curated Context as an Endpoint
- Specification Memory: What a Shared Agent Workspace Keeps
- Subtask-Level Memory for Software Engineering Agents
- The Compliance Trap: Consuming Conflicting Agent Memory
- Tiered Memory Architecture: Episodic-to-Semantic Consolidation for Long-Running Agents
- Trajectory Poisoning of Promoted Agent Skills (PoisonedEvolution)
- Treating Memory-Injection Rate as Security Evidence
- Typed Memory Provenance and Assertion Release Gating
- Usage-Reinforced Memory Decay for Long-Running Agents
- Wiki Memory: Agent-Maintained Compressed Knowledge Base
multi-agent¶
- Adaptive Sandbox Fan-Out Controller
- Adversarial Multi-Model Development Pipeline (VSDD)
- Agent Cards: Capability Discovery Standard for AI Agents
- Agent Composition Patterns for Multi-Agent Workflows
- Agent HQ (Multi-Agent Platform) for AI Agent Development
- Agent Handoff Protocols: Passing Work Between Agents
- Agent Headcount as a Vanity Metric
- Agent as Tool vs Handoff: Who Keeps the Conversation
- Agent-to-Agent (A2A) Protocol for AI Agent Development
- Agentic AI Architecture: From Prompt to Goal-Directed
- Agentic Framework Landscape: When Each Framework Fits
- Assuming Agent Interchangeability in Long-Running Teams
- Async Non-Blocking Subagent Dispatch
- Bounded Batch Dispatch for Parallel Agent Execution
- Cache-Prefix Staggering for Sibling Agent Fan-Out
- Claude Code /batch and Worktrees for AI Agent Development
- Claude Code Agent Teams for Collaborative AI Workflows
- Claude Code Dynamic Workflows
- Claude Code Sub-Agents for Delegating Complex Tasks
- Closed-Loop Role-Based Refinement for Agent Systems
- Cloud Parallel Review Pattern
- Code Injection Defense in Multi-Agent Pipelines
- Cognitive Reasoning vs Execution: A Two-Layer Agent
- Cohesion-Aware Task Partitioning for Multi-Agent Coding
- Committee Review Pattern for Multi-Agent Code Review
- Constraint Drift: Why Safety Must Be Maintained, Not Asserted
- Context-Graph Shared Memory for Multi-Agent Systems
- Contextual Capability Calibration for Multi-Agent Delegation
- Coordination Channel Policy for Multi-Agent Coding
- Coverage-Guided Fuzzing for Multi-Agent LLM Systems (FLARE)
- Cross-Tool Subagent Comparison
- Cursor /multitask: Async Subagent Dispatch in the Editor
- Cursor 3 Agents Window: Parallel Agents and Worktree Isolation
- Decentralized Memory for Self-Evolving Multi-Agent Systems
- Declarative Multi-Agent Composition
- Declarative Multi-Agent Topology: Topology-as-Code
- Declared Peer Consensus as Context for a Reviewing Agent
- Delegation Threshold Calibration for Orchestrator Agents
- Developer as CPU Scheduler: Attention Management with Parallel Agents
- Difficulty-Aware Topology Selection for Coding Agents
- Distributed Computing Parallels in Agent Architecture
- Economic Value Signaling in Multi-Agent Networks
- Emergent Behavior Sensitivity for AI Agent Development
- Event Sourcing for Agents: Separating Cognitive Intention
- Event-Loop Contention in Async Agent Fan-Out
- Factory Over Assistant: Orchestrating Parallel Agent Fleets
- Failure-Aware Observability for Multi-Agent LLM Systems
- Fan-Out Synthesis Pattern for AI Agent Development
- File-Based Agent Coordination for AI Agent Development
- Foresight-Guided Defense Against Infectious Jailbreaks in Multi-Agent Systems
- Forked vs Fresh Subagents: When to Inherit the Parent Conversation
- GitHub Copilot Advanced Patterns: Multi-Agent and Automation
- Governance Layer for Agent Interoperability Protocols
- Handoff-Boundary Fault Injection (llmmas-otel)
- Heartbeat-Bound Hierarchical Credentials for Agent Swarms
- Homogeneous Debate Panels as a Groundedness Quality Lever
- Independent Test Generation in Multi-Agent Code Systems
- LLM Map-Reduce Pattern for Parallel Input Processing
- Lane-Based Execution Queueing
- Lead-to-Teammate Plan-Approval Handshake for Multi-Agent Work
- Magentic Orchestration: Task-Ledger-Driven Adaptive Multi-Agent Planning
- Monolith-to-Sub-Agents Refactor: Five Lessons from a Brittle Prototype
- Multi-Agent RAG for Spec-to-Test Automation
- Multi-Agent SE Design Patterns: A Taxonomy Across 94 Papers
- Multi-Agent Shared State Isolation Anomalies
- Multi-Agent Systems: Coordination and Orchestration
- Multi-Agent Topology Taxonomy: Centralized, Decentralized
- Multi-Model Plan Synthesis for System Architecture
- Observation-Driven Coordination: CRDT-Based Parallel Agent
- Offline Trajectory Replay for Multi-Agent Workflow Debugging
- Opponent Processor / Multi-Agent Debate Pattern
- Oracle-Based Task Decomposition for AI Agent Development
- Orchestrator-Worker Pattern for AI Agent Development
- Over-Orchestrated Agent Architecture (Prefer the Simplest That Works)
- Parallel Agent Sessions Shift the Bottleneck from Writing
- Parsimonious Agent Routing for Multi-Agent Dispatch
- Path-Scoped Write Contracts for Shared Agent State
- Patterns: Agent Design, Multi-Agent, and Anti-Patterns
- Peer Refusal as a Coordination Control
- Per-Reviewer Context Views for Code Review Agents
- Persistent Shared Search Sub-Agent for Output-Token Reuse
- Pre-Write Change Intent Admission (Claim Plane)
- Rainbow Deployments for Agents: Gradual Version Migration
- Recursive Agent Harnesses (RAH)
- Recursive Best-of-N Delegation
- Recursive Sub-Agent Delegation: Depth Limits and Trade-offs in Nested Hierarchies
- Rerunnable Claim Graph as Shared Agent Memory
- Reverse-Engineered Executable Specifications for Agentic Program Repair
- Reviewer Precision as a Pipeline Quality Proxy
- Role-Declared Context Mode: Where the Inherit Decision Lives
- Route Agent Peers by Enrolled Identity, Not Card Name
- Runtime Workflow Selection Across Models (Project HydraFusion)
- Semantic Caching for Multi-Agent Code Systems
- Silent Handoff Failure in Delegated Code Search
- Specialist Orchestrated Queuing for Multi-Agent SE (SPOQ)
- Specialized Agent Roles for Effective AI Pipelines
- Sprint Contracts: Pre-Coding Success Agreements for Multi-Agent Tasks
- Staggered Agent Launch: Preventing Thundering-Herd in Swarms
- Static Roster vs Runtime Subagent Definition
- Structural Coverage Criteria for Agent Workflows
- Sub-Agents for Fan-Out Research and Context Isolation
- Subagent OTel Trace Correlation via agent_id Attribute
- Subagent Schema-Level Tool Filtering for AI Agents
- Swarm Migration Pattern
- Swarm Skills: Multi-Agent Extension of the Agent Skills Standard
- Symphony: Open Spec for Issue-Tracker-Driven Coding Agent Orchestration
- System-Level Optimization Pipeline
- The Model Economics of Agent Swarms: Cost and Width
- The Orchestrator's Attention Budget: Delegating to Protect Context
- The Subagent Inheritance Contract: What Crosses Down
- Tiled Agent Layout: Supervising Parallel Agents Through Dedicated Panes
- Toolset Agentization: Wrapping Co-Used Tools as Sub-Agents
- Treating Agent Delegation as Routing, Not Authorization
- Treating a Clean Merge as Compatibility Evidence
- Treating a Worktree as a Safety Boundary
- Typed Schemas at Agent Boundaries for Multi-Agent Systems
- Verifier-Driven Parallel Coding Agents (Glite ARF)
- Verify-Gated Completion as Admission Control
- Voting / Ensemble Pattern for AI Agent Development
observability¶
- Action-Audit Divergence: A Four-Mode Taxonomy for Runtime Hardening
- Agent Chat History as a First-Class Artifact
- Agent Debug Log Panel: Chronological Event Inspection for Session Debugging
- Agent Debugging: Diagnosing Bad Agent Output
- Agent Development Lifecycle for Agent Products
- Agent Event Streaming: Consumer Contract Above the Tokens
- Agent Failure Trajectories and the Recovery Window
- Agent Governance Plane: Audit Events and Message-Content Surfaces
- Agent Harness: Initializer and Coding Agent Pattern
- Agent Headcount as a Vanity Metric
- Agent Observability with OpenTelemetry and Trajectory Logging
- Agent View: Dispatch-Attach-Monitor Surface for Parallel Sessions
- Agent-Reactive Bugs at the Model-Harness Boundary
- Agent-Trace Data Layer: Storage for Hours-Long Traces
- Agentic AI Architecture: From Prompt to Goal-Directed
- Agentic Detection and Response at the MCP Boundary
- Agentic-Agile: Adapting Agile Rituals for Agent Work
- Anonymized Customization Metrics in the Copilot CLI
- Attributed Cache Misses: Reading Why a Prefix Diverged
- BYOK Model Token Visibility: Closing the Observability Gap on Self-Hosted Routes
- Behavioral Drivers of Coding Agent Success and Failure
- CausalFlow: Counterfactual Repair for Failed Agent Trajectories
- Centralized LLM Gateway for Per-Dimension Agent Budgets
- Circuit Breakers for Agent Loops
- Claim-to-Evidence Trace Graphs for Auditing Agent Runs
- Coding-Agent Misalignment Forms (Seven-Symptom Taxonomy)
- Cohort Segmentation in the Copilot Usage Metrics API
- Completion Summary as the Oversight Surface
- Context-Usage Attribution: Per-Source Breakdown of Agent Context
- Context-Window Diagnostic Tooling: Identifying Context-Heavy Tools
- Corpus-Level Trace Diagnostics for LLM Agents
- Cost-Aware Tracing for Skill Distillation
- Cost-Driven Model Routing Without Quality Monitoring
- Cross-Layer Evidence for Agent Attack Detection
- Debugging the Tool-Call Loop Before Reaching for a Framework
- Declarative Multi-Agent Composition
- Delta Channels: Bounded Checkpoint Storage for Append-Only Agent State
- Detecting Memory-Poisoning Exfiltration by Tool-Call Order (Recall-Before-Send Signature)
- Dominator-Graph Trajectory Invariants for Non-Deterministic Agents
- Dual-Write Append-Mirror for Agent Transcript Externalization
- Engineering: Tools, Review, Verification, Security, and Observability
- Enterprise Agent Hardening: Three Production Gates
- Event Sourcing for Agents: Separating Cognitive Intention
- Event-Loop Contention in Async Agent Fan-Out
- Evidence-Chain Run Logs: Bracket the Reported Symptom
- Evidence-First Reports From Failure-Diagnosis Agents
- Failure-Aware Observability for Multi-Agent LLM Systems
- Five-Failure-Layers Diagnostic: Attribute Before Swapping the Model
- Handoff-Boundary Fault Injection (llmmas-otel)
- Harness Bug Detection Patterns
- Harness Preflight Doctor Command for Agent Diagnostics
- In-Session Transcript Search: Navigating Long Agent Conversations
- Inferring Agent Failure from Conversation Evidence (Perceived Error)
- Intervention Rate as a Diagnostic North Star, Not a Target
- LLM Agent Bug Fix Taxonomy: 23 Fix Patterns from 930 Real Bugs
- Learned Prefix Monitors for Agent Traces
- Loop Detection for AI Agents: Stopping Micro-Loops
- Macro Evals for Agentic Systems: Population-Level Behavior Patterns
- Making Application Observability Legible to Agents
- Markov-Chain Reliability for LLM Agents: Audit the Abstraction Before You Trust the Metric
- Monitor Tool: Event Streaming from Background Scripts
- Monolith-to-Sub-Agents Refactor: Five Lessons from a Brittle Prototype
- Multi-Turn Conversation Evaluation: Per-Turn and Trace-Level Scoring Together
- Observability Feedback Loop: A 7-Step Debug Runbook for Agents
- Observability for AI Agents: Tracing and Debugging
- Observability-Driven Harness Evolution
- Offline Trajectory Replay for Multi-Agent Workflow Debugging
- OpenTelemetry for AI Agent Observability and Tracing
- Out-of-Band Hook Notifications via terminalSequence
- Per-Agent-App Attribution in the Copilot Usage Metrics API
- Per-Plugin Token-Cost Attribution via claude plugin details
- Persistent-Connection Agent Transport
- Planted-Bug Methodology: Deliberate Bugs as Observability Calibration
- Plugin Background Monitors: Declarative Supervision Auto-Armed at Session Start
- Prebuilt Agent Monitoring Dashboard
- Programmatic Agent Session Export via `claude agents --json`
- Reader-Scoped Trace Views: When to Build Your Own UI
- Reading Copilot Feature Engagement by Its Threshold
- Run-Status vs Task-Status Confusion in Autonomous Agent Runs
- Self-Reporting Loops: Autonomous Routines That File Their Own Backlog
- Session Harness Sandbox Separation for Long-Running Agents
- Silent-Failure Mechanism Taxonomy in Production Agent Runtimes
- Sizing Vendor-Emitted Agent Telemetry by Signal Tier
- Splitting the Drift Judge from the Advisor (LivePlan)
- Stakeholder Trust Through Evals and Observability
- Stateful Agent Evals via State Snapshots and Transition Assertions
- StopFailure Hook: Observability for API Error Termination
- Strained Coherence as a Pre-Failure Signal in Agent Trajectories
- Structural Monitoring for Covert Safeguard-Weakening
- Subagent OTel Trace Correlation via agent_id Attribute
- Suspect the Harness Before the Model on a Regression
- Symptom-First Bug Triage for Agent Code
- Traces Need Feedback to Power Learning
- Trajectory Attribution for Context Repair (TRACE)
- Trajectory Decomposition: Diagnose Where Coding Agents Fail
- Trajectory Logging via Progress Files and Git History
- Trajectory Pre-Filter for Failure Diagnosis (TrajAudit)
- Trajectory Projection: A Flattened View of Agent Traces
- Trajectory as the Monitoring Unit for Production Agents
- Transcript-Driven Permission Allowlist
- Transcript-Measured Review Coverage
- Using the Agent to Analyze Its Own Evaluation Transcripts
- Verification Ledger for Tracking Agent Output Quality
- Verify Observability in Agent-Generated Code
pattern¶
- Addressable Recall Compaction: Compact to Citations, Not Summaries
- Agent Terminology Disambiguation for AI Coding Systems
- Agent-Initiated Rubric-Gated Self-Compaction (SelfCompact)
- Agentic Pattern Vocabulary Crosswalk
- Anthropic's Effective Agents Framework: A Pattern Map
- Canvas as Control Surface: Steering a Long-Running Agent Mid-Run
- Canvas as Durable Workflow State: The Four-Step Blueprint and What It Costs
- Classical SE Patterns as Agent Design Analogues
- Coding-Agent Working-Set Coverage (Coherence Debt)
- Continuation Dispatcher: Who Owns the Iterate-or-Stop Call
- Delegated-Autonomy Boundary Artifacts (AJR and ADP)
- Durable Interactive Artifacts: Agent Output Outside the Transcript
- Managing Cognitive Load and AI Fatigue for Sustainable Agent Use
- Minimum-Sufficient Control Ladder: Escalate by Failure Mode
- Patterns: Agent Design, Multi-Agent, and Anti-Patterns
- Prompt-Only Baseline Before a Specialized Agent Subsystem
- Recursive Agent Harnesses (RAH)
- Selective Rewind Summarization: Compress Earlier Turns, Keep Recent Ones Intact
- Session Recap: Goal-Shaped Handoff at Context Boundaries
- Suggestion Gating: Fewer Completions, Better DX
- Version-Controlled Agent Context (Git Context Controller)
rag¶
- AOCI: Symbolic-Semantic Repository Indexing
- Chunking Strategy for RAG-Based Code Completion
- Codebase-Derived Pattern Libraries as Agent Context
- Component-Wise RAG Prioritization for Software Engineering Tasks
- Compositional Skill Routing for Large Skill Libraries
- Context Hub: On-Demand Versioned API Docs for Coding Agents
- Corpus Shape as a Retrieval Design Constraint
- Cross-Reference Dereference Hop in Retrieval Loops
- Decoupled Search Grounding: A Vendor-Agnostic Grounding Boundary
- Embedding Inversion: Vector Stores as a Source-Text Disclosure Surface
- Fact Supersession Memory for Code Assistants
- Gate Generation on Retrieval Sufficiency, Not Model Confidence
- Hypothetical Classification for Large Label Vocabularies
- LLM-Driven Logical Retrieval: Boolean Queries over an Inverted Index
- Lexical-First Retrieval for Agentic Search: When BM25 Is Enough
- Multi-Agent RAG for Spec-to-Test Automation
- Multitenant RAG: Closing the Relevance-Authorization Gap
- Organizational Context Layer for Agents (Company Brain)
- Per-Object Context Allocation (Selective Invariance)
- RAG Architecture as a Poisoning Robustness Decision
- RAG over Thinking Traces: Index Reasoning Trajectories Instead of Documents
- RAG/Agent Reliability Problem Map: 16-Domain Failure Taxonomy
- Repository Map Pattern: AST + PageRank for Dynamic Code
- Repository-Level Retrieval for Code Generation
- Retrieval-Augmented Agent Workflows: On-Demand Context
- Schema-Guided Graph Retrieval
- Semantic Context Loading: Language Server Plugins for Agents
- Silent Handoff Failure in Delegated Code Search
- State-Conditioned Evidence Selection for Mid-Task Retrieval
- Structured Domain Retrieval: Knowledge Graphs and Case-Based Reasoning
- Typed Generation Contracts for Grounded Extraction
- When a Skill Graph Cannot Beat the Ranker (Pre-Filter Topology Bound)
reliability¶
- Agent Backpressure: Automated Feedback for Self-Correction
- Agent Circuit Breaker
- Agent Failure Trajectories and the Recovery Window
- Agent-Client Admission Control for Agentic Traffic
- Behavioral Drivers of Coding Agent Success and Failure
- Caller-Actionable Error Steps in Tool Responses
- Classifying and Auto-Correcting Coding Agent Misbehaviors (Wink)
- Cross-Vendor Competitive Routing for LLM Selection
- Decoupled Search Grounding: A Vendor-Agnostic Grounding Boundary
- Dispatch-Time Reasoning Level for Delegated Agents
- Dual-Budget Control for Search Agents: VOI Scoring Per Action
- Effective Feedback Compute (EFC) for Harness Comparison
- Error Preservation in Context for AI Agent Development
- Exception Handling and Recovery Patterns for AI Coding Agents
- Feedback as Capability Equalizer: Iterative Feedback Outweighs Model Scale
- Five-Failure-Layers Diagnostic: Attribute Before Swapping the Model
- Gateway Model Routing: Treat the LLM Gateway as a Discovery Source
- Heuristic-Based Effort Scaling in Agent System Prompts
- Idempotent Agent Operations: Safe to Retry
- Interactive Effort Sliders: Per-Turn Reasoning-Budget Controls
- Long-Running Agents: Durability and Resumability Across Sessions
- Multi-Client Session Attachment: One Session, Many Clients
- Observation Contract Preservation in Tool-Augmented Agents
- Outcome Monitors: Recovery Affordances for Tool Failures
- Per-Call Budget Hints on Tool Invocations
- Per-Run Budget Reservation for Coding Agent Model Calls
- Per-Tool Extended Reasoning Opt-In: Tool-Call-Scoped Budgets
- Per-User Supervisor Process for Background Agent Sessions
- Progressive Spend Threshold Alerting for Agent Cost Governance
- Reasoning Budget Allocation: The Reasoning Sandwich
- Reasoning Effort Over Tool Scaffolding for First-Try Reliability
- Remote Agent Host Sessions over SSH and Dev Tunnels
- Residual Completion for Stateful Agent Handoffs (CFRC)
- Retry-Switch-Abstain: A Runtime Tool-Recovery Policy
- Rollback-First Design: Every Agent Action Should Be Reversible
- RubricRefine: Pre-Execution Rubric Refinement for Code-Mode Tool Use
- Selective Checkpoint Restore Across Code and Conversation State
- Selective Revalidation for Pending Agent Actions
- Self-Healing Production Agent: Automated Regression Detection and Autofix PR
- Specialized Small Language Models as Agent Sub-Tools
- Splitting the Drift Judge from the Advisor (LivePlan)
- Tail Control for Agent Workflows: Engineering for the Failure Tail, Not the Average
- Task Feasibility Awareness: Stop Before You Start
- Tenant Model Policy: Organization-Scoped Rules for AI Model Selection
- The Advisor Strategy: Frontier Model as Strategic Advisor
- Trajectory-Conditioned Model Escalation (SWE-Router)
- WIP=1 and Little's Law: Kanban Throughput Theory for Agent Task Design
sandboxing¶
- Browser Sandbox for Agent-Generated HTML (Sandboxed Iframe + Immutable CSP)
- Cross-OS Library MicroVM Sandbox for Agent Code (smolvm)
- Docker sbx Adoption for Coding Agents
- Dual-Boundary Sandboxing: Filesystem and Network Isolation
- In-Process WebAssembly Sandboxes for Agent-Generated Code
- Sandboxed Coding Environments: Containers vs MicroVMs vs OS-Level Isolators
- Scope Sandbox Rules to Harness-Owned Tools, Not Third-Party
- Subprocess PID Namespace Sandboxing in Claude Code
- Windows Sandboxing for Coding Agents
- Workload-Keyed Sandbox Selection for Agent-Generated Code
security¶
- A Governance Framework for Production Agents
- AI Agents in CI/CD with Elevated Permissions and Untrusted Content (GitInject)
- AI-Powered Vulnerability Triage for AI Agent Development
- Action-Audit Divergence: A Four-Mode Taxonomy for Runtime Hardening
- Action-Graded Severity for Agent Red-Team Outcomes
- Action-Selector Pattern: LLM as Intent Decoder with Deterministic Execution
- Adaptive Evaluation of Out-of-Band Prompt-Injection Defenses
- Adversarial-Only Threat Modeling for Agent Data Leakage
- Agent Approval Laundering: Effects Beyond the Named Command
- Agent Commit Attribution: Signed Commits and Agent Identity
- Agent Governance Plane: Audit Events and Message-Content Surfaces
- Agent Network Egress Policy: Admin-Controlled Domain Allow/Deny
- Agent Retrieval Provenance as an Audit Control
- Agent Runtime Middleware: Per-Call Interception Pipeline
- Agent-Authored Messages as a Deferred Exfiltration Channel
- Agent-Driven Fuzzing with Human-Gated Crash Triage
- Agent-Emitted Dependency Version Ranges Widen the Supply-Chain Attack Surface
- Agent-Native Filesystems: Gating Effects, Not Commands
- Agentic Detection and Response at the MCP Boundary
- Aggregation Bounds for Agent Authorization
- Always-On Agentic PR Security Review
- An Explicit Update Boundary for Agent Self-State
- Assuming a CLAUDE.md Security Rule Is Enforced
- Auth-Isolation as the MCP-vs-CLI Selection Heuristic
- Authority Confusion: Untrusted Context Must Not Authorize Side Effects
- Authorization Continuity Across Agent Mutation
- Behavior-Partitioned Security Tests as Executable Specs
- Behavioral Firewall for Tool-Call Trajectories
- Benchmark Poisoning of Self-Modifying Coding Agents
- Benign Skill Wording That Steers Package Hallucination (Neutral Prompting Attack)
- Binding an Agent's Effect to the Approval It Claims
- Black-Box Agent Risk Scoring by Domain
- Black-Box Probing in the Agent Build Loop: When It Pays
- Blast Radius Containment: Least Privilege for AI Agents
- Browser Sandbox for Agent-Generated HTML (Sandboxed Iframe + Immutable CSP)
- Browser as Agent Action Space
- Bug-Class Hints as Exploit Input for Coding Agents
- Calibrated Deciders for In-Loop Agent Decisions
- Capability-Additive Code Interpreters for Untrusted Agent Code
- Capability-Pegged Security Re-Scans: Reviewing Unchanged Code When the Scanner Improves
- Centrally Provisioned MCP Servers: Remote Transports Only
- Chat-Platform Agent Delegation: Invoking Cloud Coding Agents from Team Channels
- Clarification Mode Amplifies Prompt Injection
- Classifier-Gated Auto-Permission for Cloud-IDE Coding Agents
- Classifier-Subagent Run Mode for Per-Call Permission Routing
- Claude Code Auto Mode: Classifier-Based Permission Gating
- Close the Attack-to-Fix Loop: Adversarially Train Agent
- Closed-World Tool Call Resolution Before the Permission Gate
- Code Injection Defense in Multi-Agent Pipelines
- Code Interpreter as a Primary Agent Tool
- Coding Agent Scope Expansion: When to Extend Beyond the Codebase
- Cognitive Poisoning: Untrusted Tool Feedback as a Trajectory Attack
- Comment-Triggered Agent Dispatch on Issues and PRs
- Compositional Vulnerability Induction in Coding Agents
- Computer-Systems Lens for Always-On Agent Security
- Confirmation Gates for Consequential Agent Actions
- Constraint Drift: Why Safety Must Be Maintained, Not Asserted
- Constraint Preambles and the Gain Your Scanner Misses
- Constraints as a Substrate for Scalable Agent Oversight
- Containment Playbook: npm-to-Signing-Channel Compromise
- Content Exclusion Gap: AI Security Boundaries by Mode
- Content-Addressed Agent Configurations (Deterministic Control Plane)
- Context-Fractured Decomposition Attacks on Tool-Using Agents
- Control/Data-Flow Separation for Prompt Injection Defense (CaMeL)
- Controlled Benchmark Rewriting for Agent Safety Judgment
- Copilot Cloud Agent Organization Controls
- Covert Success Rate for Indirect Prompt Injection
- Credential Hygiene for Agent Skill Authorship
- Cross-Iteration Safety State for Agent Loops (LoopHarness)
- Cross-Layer Evidence for Agent Attack Detection
- Cross-OS Library MicroVM Sandbox for Agent Code (smolvm)
- Cross-Repo Agent Search: GitHub-API-Backed Text Search Beyond the Workspace
- Cross-Repository Security Posture for Agent-Introduced Vulnerabilities
- Cryptographic Governance Audit Trail for AI Agents
- Cumulative-Best Reporting Hides Repair-Loop Security Regressions
- Cursor Automations: Event-Triggered Agents and /automate
- Cursor Self-Hosted Cloud Agents
- Customer-Hosted MCP Tunnel: Outbound-Only Connectivity to Private MCP Servers
- Data Fidelity Guardrails: Preventing Agent Data Mutation
- Decomposed Red-Teaming for Agent Monitors
- Defense-in-Depth Agent Safety for AI Agent Development
- Deferred Permission Pattern: Headless Agent Session Pausing
- Delegating Delivery Stages to GitHub Agent Apps
- Delegating Dependabot Pull Request Triage to an Agent
- Delivery-Bound Tool Authorization: When Progressive Discovery Becomes Access Control
- Deny-Fallback Permissions for Unattended Agent Runs
- Dependabot Agent Assignment for AI-Driven Vulnerability Remediation
- Designing Agents to Resist Prompt Injection
- Destyling Untrusted Input as a Prompt Injection Defense
- Detecting Memory-Poisoning Exfiltration by Tool-Call Order (Recall-Before-Send Signature)
- Direct Prompt Injection via Collaboration (User as Attack Vector)
- Directory-Aware Plugin Suggestions via `pluginSuggestionMarketplaces`
- Discovering Indirect Injection Vulnerabilities in Your Agent
- Distributed Cross-PR Attacks in Persistent-State AI Control
- Distributing Security Controls Through the Agent Harness
- Docker sbx Adoption for Coding Agents
- Document-Borne Prompt Injection Through Agent Read Tools
- Dormant Memory Payloads Triggered by Sensitive Topics (Trojan Hippo)
- Downstream Disclosure Coordination for Agent-Found Defects
- Dual-Boundary Sandboxing: Filesystem and Network Isolation
- Dual-Graph Alignment for Indirect Prompt Injection Defense (AuthGraph)
- Embedding Inversion: Vector Stores as a Source-Text Disclosure Surface
- Enforced Versus Advisory Controls in LLM-Native IDEs
- Enforcing Who and What Can Trigger an Agent's CI Run
- Engineering: Tools, Review, Verification, Security, and Observability
- Enterprise Agent Hardening: Three Production Gates
- Enterprise-Managed Plugin Governance for Agent CLIs
- Enumerate Every Visible Exit Before Scoring Containment
- Eval Environment Containment for Cyber-Capable Agents
- Evidence-Based Allowlist Auto-Discovery for Agents
- Execution-Layer Security Invariants for MCP Runtimes
- Explained Feedback for LLM Vulnerability Repair
- Explanation-Bound Tool Execution for Agent Gateways
- External Artifacts Treated as Data, Not Adversarial Input
- Fail-Closed Remote Settings Enforcement for Enterprise Agents
- Field-Level Source Ownership for Agent Capabilities
- Five-Stage Policy Layer Typology for Generalist Agents
- Flattened Tool Specs for Agent Safety Judgment (SafeKeep)
- Foresight-Guided Defense Against Infectious Jailbreaks in Multi-Agent Systems
- Forged Reasoning Trace Attacks on Agent Memory (FARMA)
- Four-Layer Taxonomy of Agent Security Risks
- Framing Subagent Returns So They Cannot Act as Instructions
- From Preventive to Reactive: Front-Loading Security in AI Coding Prompts
- Gate Agent Writes to Executable Config Files as Privileged Actions
- Generated Programs as Web Agent Action Space
- GitHub Agentic Workflows for Automating Dev Processes
- Goal Reframing: The Primary Exploitation Trigger for LLM Agents
- Guarding Against URL-Based Data Exfiltration in Agentic Workflows
- Hard-Deny Classifier Rule: Unconditional Block in Auto Mode
- Harness Composition for Scaled Security Audits
- Heartbeat-Bound Hierarchical Credentials for Agent Swarms
- History Anchors: Consistency-Cued Continuation of Unsafe Prior Actions
- Hook Exec Form vs Shell Form: Shell-Injection-Safe Hook Commands
- Hooks Invoking MCP Tools: Closing the Loop Between Policy and Tool Execution
- Hostname-Allowlist Proxy: The TLS-Inspection Blind Spot
- Hybrid Deterministic + Semantic Authorization for Agent Tool Calls
- Improper Output Handling: Validate Agent Output Before Downstream Use
- In-Process WebAssembly Sandboxes for Agent-Generated Code
- Inline Safety Harness with Cascade Verification (FinHarness)
- Install-Once Plugin Trust: Vetting That Never Re-Runs on Update
- Intent-Governed Tool Authorization for AI Agents (IGAC)
- Internal Hostname Disclosure in Agent-Readable Context
- Judging Agent Safety by Task Completion (Action-Boundary Violations)
- Judging a Skill's Honesty by the Validity of Its Output
- Knowledge-Based Pull Requests for Cross-Trust-Boundary Contributions
- LLM API Routers as Application-Layer Man-in-the-Middle
- LLM-Pinned Library Versions Carry Systemic CVE Exposure
- Layered Oracle Stack for Agent IaC Security Repair (TerraProbe)
- Lethal Trifecta Threat Model for AI Agent Development
- Lifecycle-Integrated Security Architecture for Agent Harnesses
- Live Browser as Agent Context Channel
- Local Sandboxing in the Copilot App: The Credential Axis
- Lock-State Safeguards for Desktop-Controlling Agents
- MCP Allowlist by Label, Not by Identity (serverName Trap)
- MCP Approval-View Fidelity Gap and Unicode Concealment
- MCP Runtime Control Plane: Policy Evaluation Between Agent and Tool
- Managed Settings Drop-In Directory: Enterprise Policy Fragmentation
- Mid-Trajectory Guardrail Selection for Multi-Step Tool Calls
- Model Confidence as Security Verification (Security Calibration Gap)
- Monotonic Capability Attenuation for Composition-Safe Tool Use
- Most-Restrictive-Wins Fusion for Parallel Agent Control Returns
- Multi-Repo and No-Repo Coding Agent Automation Templates
- Multi-Tenant Isolation Knobs for Shared-Container Agent SDK Hosting
- Multi-Tool Threshold Poisoning Against MCP (ShareLock)
- Multitenant RAG: Closing the Relevance-Authorization Gap
- Network-less Container + Unix-Socket Egress Proxy for Agent Sandboxes
- Next Edit Suggestions Carry Context You Never Curated
- Non-Human Event Provenance Markers to Block Fabricated Approvals
- Non-Retirable Approval Rules for Agent Operations
- OAuth Client ID Metadata Documents (CIMD) for MCP Servers
- OWASP 2026 Update for Agent Builders: Top 10 Renumbering and the Agent Control Standard
- OWASP LLM Top 10 (2025): Agent Security Crosswalk
- On-Demand Skill Hooks: Session-Scoped Guardrails via Skill Invocation
- OpenAI Agents SDK
- Oracle Poisoning: Knowledge Graph Corruption Against Tool-Using Agents
- Org-Membership-Gated Agent Entitlement
- Overeager-Behavior Elicitation: Scope + Trap Fragments as a Diagnostic for Out-of-Scope Tool Calls
- Parameter-Level Permission Rules (Tool(param:value) Syntax)
- Parser-Versus-Shell Evasion in Command Permission Checks
- Per-Agent Capability Stores Beat One Task-Wide Allowlist
- Per-Caller Identity: Who an Agent's Tool Call Acts As
- Per-Server MCP Environment Scoping for Credential Isolation
- Permission Framework Choice Outweighs Model Choice for Limiting Overeager Actions
- Permission Gates That Deny the Agent's Own Cleanup (Denied Remediation Path)
- Permission Modes as a Defense Against a Tampered Response Path (Response-Path Control Gap)
- Permission-Gated Custom Commands for AI Agent Development
- Permitted Egress Routes as Agent Sandbox Attack Surface
- Plan-Then-Execute as the Default for Web Agents
- Plugin-Activated Main-Agent Override and Bin/ PATH Injection
- Policy-Graded Evaluation of Coding Agents
- Pooled-Evidence Factuality Checks for MCP Agents (Cross-Source Conflation)
- PostToolUse Output Replacement: Hooks That Rewrite Tool Results
- Pre-Execution Risk Classification for Terminal Commands
- Pre-Trust Execution Surface in Coding Agent Harnesses
- Privacy-Preserving LLM Requests: Eight Techniques and a Practical Combination
- Programmatic Cloud-Agent Dispatch via REST API and Webhooks
- Prompt Injection: A First-Class Threat to Agentic Systems
- Prompt as Security Knob
- Prompt-Only Tool Access Control
- Proof of Presence: Re-Authenticating for Agent Actions
- Protecting Sensitive Files from Agent Context Access
- Provenance-Aware Decision Auditing for LLM Agents
- Public Placeholder Credentials with a Fail-Closed Injecting Proxy
- RAG Architecture as a Poisoning Robustness Decision
- RL-Trained Automated Red Teamers for Prompt Injection Discovery
- Reading Visible Edge-Case Handling as a Security Check
- Reading a Coding-Agent Vendor's Security Certificate
- Recover the Six Measurement Choices Behind an Attack Success Rate
- Red-Team Your Blocking Monitor Before You Trust It
- Reframed Exfiltration Defeats Wording-Based Defenses (Framing Gap)
- Replayable Encrypted Reasoning Blocks in Agent Traces
- Restricted-Access Defensive AI: Project Glasswing as a Deployment Model
- Revocable Resource-and-Effect Capabilities for Coding Agents (PORTICO)
- Root Causes of Vibe-Coded Application Vulnerabilities
- Route Agent Peers by Enrolled Identity, Not Card Name
- Route-Parity Auditing of Agent Safeguards
- Runtime Guard as an Installed Skill (Defense-as-Skill)
- SUDP: Secret-Use Delegation Protocol for Agentic Systems
- Safe Command Allowlisting: Reducing Approval Fatigue
- Safe Outputs Pattern for Trustworthy Agent Responses
- Sandbox + Approvals + Auto-Review Governance Triad
- Sandbox Credential Masking: Authenticate Without Seeing the Secret
- Sandbox-Enforced PII Tokenization in Agent Workflows
- Sandboxed Coding Environments: Containers vs MicroVMs vs OS-Level Isolators
- Scanner-as-MCP-Server: Secret and Dependency Scans as Typed Agent Tools
- Scope Sandbox Rules to Harness-Owned Tools, Not Third-Party
- Scoped Browser DevTools Access for Runtime Diagnosis
- Scoped Credentials via Proxy Outside the Agent Sandbox
- Scoped-Looking Permission Grants
- Scoring Constraint Loss and Tool Reach as Separate Risks
- Secrets Management for AI Agents: Credential Injection
- Security Budget as Token Economics
- Security Constitution for AI Code Generation
- Security Drift in Iterative LLM Code Refinement
- Security Knowledge Priming for Code Generation (SPARK)
- Security for AI Agent Development
- Security-Aware Tool Descriptions for MCP Servers (SpellSmith)
- Selective Network Access in Agent Sandboxes: The allowNetwork Pattern
- Semantic Intent Validation for Agent Skills
- Sensitive Terminal Prompt Interception
- Setup Documentation as an Install-Time Attack Vector
- Severity-Stratified Evaluation of Security Prompts
- Single-Decision Approval of Vendor Skill Suites
- Single-Layer Prompt Injection Defense Anti-Pattern
- Skill Composition Risk in Agent Ecosystems
- Skill Misevolution in Self-Updating Skill Libraries
- Skill Review Without a Token Cost Baseline
- Skill Shell Execution Gate: Disabling Inline Shell from Skills
- Skill Specification Violation Fuzzing
- Skill Supply-Chain Poisoning
- Skill disallowed-tools Frontmatter: Skill-Layer Tool Denial
- Slopsquatting: Hallucinated Package Names as a Supply-Chain Vector
- Static-First Shell Command Gating with Selective Escalation
- Structural Monitoring for Covert Safeguard-Weakening
- Subprocess PID Namespace Sandboxing in Claude Code
- Sufficiency-Tightness Decomposition for Agent-Authored Permissions
- Supply-Chain Security Debt in Agent Pull Requests
- System Prompt Delivery Channels on Shared Runners
- System Prompt as Secret Store (OWASP LLM07)
- Task Alignment: The Selective-Compliance Gap Benchmarks Miss
- Task Category as the Security Review Routing Key
- Task-Based Access Control with Hybrid Inspection
- Team-Scoped Agent Policy Delegation
- The Agent Stack Bet: Architectural Decisions for Production Agents
- The Decoding Harness Is Part of the Agent Attack Surface
- The Post-Authorization Execution Trust Gap in Remote MCP
- The Security Review Gap in AI-Authored PRs
- The Skill Closure Declaration Gap
- Three-Depth In-Session Security Review
- Three-Vector Evasion Taxonomy for Agent Security Tests
- Tool Cloning and Provenance Assessment in Agent Ecosystems
- Tool Signing and Signature Verification for Agents
- Tool-Invocation Attack Surface in Coding Agents
- Trajectory Poisoning of Promoted Agent Skills (PoisonedEvolution)
- Transcript-Driven Permission Allowlist
- Treat Task Scope as a Security Boundary
- Treating Agent Delegation as Routing, Not Authorization
- Treating Agent Safety as Uniform Across a Session (Cold-Start Safety Gap)
- Treating File Secrecy as Skill Confidentiality
- Treating Memory-Injection Rate as Security Evidence
- Treating Read Denial as Confidentiality for a Build Input
- Treating a Clean Static Scan as Security Evidence
- Treating a Local Agent Session Trace as Audit Evidence
- Treating a Worktree as a Safety Boundary
- Trigger-Level Gating for Autonomous Agent Intake
- Trusting Claimed Prior Approval in Agent Review Gates
- Trusting Human Review to Catch Deliberate Agent Sabotage
- Trusting Model-Level Privilege Restraint at Tool Selection
- Trusting Tool Error Messages as Implicit Authority (Error-Path Injection)
- Trusting a Skill Scanner's Verdict as a Security Judgment (Green-Check Fallacy)
- Unbounded Consumption: Bounding Agent Resource Use Against DoS and Denial-of-Wallet
- Usability Pressure as a Silent Security-Regression Vector
- Use a Public-Web Index to Gate Automatic URL Fetching
- Verifying LLM-Generated Cryptographic Code
- Vetting Tool Definitions for Exfiltration Signatures
- Which Task You Delegate Changes Poisoned-Repo Exposure
- Windows Sandboxing for Coding Agents
- Working Inside an Enterprise-Managed Agent Sandbox Policy
- Workload Identity Federation for Agent Runtimes
- Workload-Keyed Sandbox Selection for Agent-Generated Code
- Workspace Topology as an Indirect Injection Attack Vector
- bypassPermissions Silently Overrides allowedTools (The Restricted-Bypass Trap)
skills¶
- Agent Skills: A Cross-Tool Task Knowledge Standard
- Artifact-Only Verification Hides Skipped Skill Steps
- Assuming Loaded Skills Stay Enforced in Long Contexts
- Backlog Triage as a Named Agent Skill
- Benign Skill Wording That Steers Package Hallucination (Neutral Prompting Attack)
- CLI-First Skill Design
- Choosing a Skill Loading Method for Agents
- Compositional Skill Routing for Large Skill Libraries
- Contractual Skill Files: Inspectable SKILL.md for Enterprise Agents
- Cost-Aware Skill Rewriting: Preserve Operational Anchors, Not Skill Tokens
- Coverage-Aware Skill Selection Under a Token Budget
- Credential Hygiene for Agent Skill Authorship
- Daily-Use Skill Library: Encoding Your Process as Agent Skills
- Emulated APIs for Agent Skill Evals
- Enterprise Skill Marketplace: Distribution and Quality
- Frozen-Base Task Mining for Repository Instruction Files
- Generated Procedure Drivers: Skills That Emit a Program
- GitHub Copilot Custom Agents and Skills Extensibility Guide
- Google ADK Skills: Portable SKILL.md Across ADK Agents
- Handoff Skill: Structured Context Transfer Between Agent Sessions
- Incident Log Investigation Skill: Parallel Queries
- Introspective Skill Generation: Mining Agent Patterns
- Judging a Skill's Honesty by the Validity of Its Output
- Listener-State Naming for User-Invoked Agent Skills
- Managing Agent Skills from the GitHub CLI with gh skill
- On-Demand Skill Hooks: Session-Scoped Guardrails via Skill Invocation
- Per-Step Preconditions and Postconditions in Skill Files
- Personalized vs Generic Agent Skills: Where Effort Pays
- Plugin Component Co-Change: Scripts and Their Instructions Move Together
- Project Writing Skill: House Style as Model-Invocable Skill
- Reloading Skills Mid-Session in Claude Code
- Repository Skill Release Drift
- Runtime Guard as an Installed Skill (Defense-as-Skill)
- SDLC-Phase Skill Taxonomy: Full-Lifecycle Skill Libraries
- SKILL.md Frontmatter Reference: All Fields Explained
- Semantic Intent Validation for Agent Skills
- Single-Decision Approval of Vendor Skill Suites
- Skill Authoring Patterns: Description to Deployment
- Skill Authoring as Software Engineering: What Transfers
- Skill Composition Risk in Agent Ecosystems
- Skill Context Isolation: Forking the Skill into a Subagent Window
- Skill Eval Loop
- Skill File Linting: Which Three Checks to Run First
- Skill Library Evolution: Lifecycle Governance for Agents
- Skill Library Refinement Loops: Organizational Feedback for Shared Skills
- Skill Library Technical Debt: Library-Time Maintenance for Agent Skills
- Skill Lift: Measuring What a Skill Adds at Runtime
- Skill Loadout Curation for Coding Agents
- Skill Misevolution in Self-Updating Skill Libraries
- Skill Over-Trust: Treating Topical Relevance as Evidence a Skill Helps
- Skill Reuse as Vendored Forking
- Skill Shell Execution Gate: Disabling Inline Shell from Skills
- Skill Supply-Chain Poisoning
- Skill Test Coverage as a Release Gate
- Skill Tool as Enforcement: Loading Command Prompts at Runtime
- Skill as Instruction Surface and Callable API (Interpreter Skills)
- Skill as Knowledge Pattern for AI Agent Development
- Skill disallowed-tools Frontmatter: Skill-Layer Tool Denial
- Skill or MCP Server: Choosing a Capability's Delivery Mechanism
- Subagent vs In-Context Skill Execution
- Swarm Skills: Multi-Agent Extension of the Agent Skills Standard
- The Merge-Conflict Resolution Skill: What to Encode
- The Skill Closure Declaration Gap
- Throwaway-Prototype Skill: Build to Discard, Keep Only the Answer
- Trajectory Poisoning of Promoted Agent Skills (PoisonedEvolution)
- Treating File Secrecy as Skill Confidentiality
- Trusting a Skill Scanner's Verdict as a Security Judgment (Green-Check Fallacy)
- Typed Pseudocode for Skill Libraries (Skill-as-Pseudocode)
- Video Transcript Skill: Meeting Recording to Markdown
- When a Skill Graph Cannot Beat the Ranker (Pre-Filter Topology Bound)
- Workspace-Hosted Skills: Authorship Outside the Repo
source:opendev-paper¶
- Agent Harness: Initializer and Coding Agent Pattern
- Agent Memory Patterns: Learning Across Conversations
- Context Compression Strategies: Offloading and Summarization
- Context-Injected Error Recovery for AI Agent Development
- Cost-Aware Agent Design: Route by Complexity, Not Habit
- Defense-in-Depth Agent Safety for AI Agent Development
- Designing Agent Tools Like APIs
- Dynamic System Prompt Composition
- Event-Driven System Reminders for AI Agent Development
- Filesystem-Based Tool Discovery for AI Agent Development
- Loop Detection for AI Agents: Stopping Micro-Loops
- Model a Single Agent Turn as Many Inference and Tool-Call
- Objective Drift: When Agents Lose Sight of the Goal
- Reasoning Budget Allocation: The Reasoning Sandwich
- Subagent Schema-Level Tool Filtering for AI Agents
standards¶
- A2UI: Framework-Agnostic Generative UI Standard for Agents
- ACDL: A Language for Describing Agentic LLM Contexts
- AGENTS.md: Project-Level README for AI Coding Agents
- Agent Cards: Capability Discovery Standard for AI Agents
- Agent Definition Formats: How Tools Define Agent Behavior
- Agent Plugins: Portable Packaging With Client-Defined Trust
- Agent Skills: A Cross-Tool Task Knowledge Standard
- Agent-to-Agent (A2A) Protocol for AI Agent Development
- Agentic Resource Discovery: Federated Pre-Invocation Search
- Cross-IDE Plugin Discovery: One Install Surface, Many Consuming Agents
- Directory-Aware Plugin Suggestions via `pluginSuggestionMarketplaces`
- Governance Layer for Agent Interoperability Protocols
- MCP: The Open Protocol Connecting Agents to External Tools
- OAuth Client ID Metadata Documents (CIMD) for MCP Servers
- Open Standards and Protocols for AI Agent Development
- OpenAPI as the Source of Truth for Agent Tool Definitions
- OpenTelemetry for AI Agent Observability and Tracing
- Plugin Dependency Declaration and Disable-Chain Hints
- Plugin and Extension Packaging: Distributing Agent Capabilities
- Portable Agent Definitions: Full-Stack Identity as Code
- Pre-Install Context-Cost Projection in Plugin Marketplaces
- Pre-Install Plugin Transparency: Capability Inventory and Cost Projection
- Reference: Standards, Human Factors, Emerging, and Fallacies
- SUDP: Secret-Use Delegation Protocol for Agentic Systems
- Stateless MCP: One Request per Tool Call
- Swarm Skills: Multi-Agent Extension of the Agent Skills Standard
- Symphony: Open Spec for Issue-Tracker-Driven Coding Agent Orchestration
- Tool Calling Schema Standards for AI Agent Development
- WebMCP: Browser-Hosted Tool Contracts for In-Page AI Agents
- llms.txt: Making Your Project Discoverable to AI Agents
supply-chain¶
- Agent-Emitted Dependency Version Ranges Widen the Supply-Chain Attack Surface
- Benign Skill Wording That Steers Package Hallucination (Neutral Prompting Attack)
- Containment Playbook: npm-to-Signing-Channel Compromise
- Enterprise-Managed Plugin Governance for Agent CLIs
- LLM-Pinned Library Versions Carry Systemic CVE Exposure
- Setup Documentation as an Install-Time Attack Vector
- Single-Decision Approval of Vendor Skill Suites
- Skill Supply-Chain Poisoning
- Slopsquatting: Hallucinated Package Names as a Supply-Chain Vector
- The Skill Closure Declaration Gap
- Tool Signing and Signature Verification for Agents
technique¶
- @import Composition Pattern for Agent Instruction Files
- AGENTS.md Design Patterns for Effective Agent Files
- AI Crawler Policy: robots.txt for the Three-Tier Crawler Landscape
- Agent-Readiness Discovery Surfaces for Docs Sites
- Answer-First Writing: Structure Content for AI Retrieval
- Assertion Density — Stats and Quotes Over Vague Claims
- Atomic Pages and Chunking — One Concept Per Page for RAG
- Calibrated Early Termination and Warm Restart for Agent Runs (FailFast-RestartSmart)
- Chunking Strategy for RAG-Based Code Completion
- Comment Content as Code-Generation Context
- Convergence Detection in Iterative Agent Refinement
- Criticality and Containment: Scoping Which Agent Code You Read
- Cross-Tool Translation: Learning from Multiple AI Assistants
- Evidence-Based Allowlist Auto-Discovery for Agents
- Exhaustive Retrieval for Listing Questions
- GEO for Technical Docs: Developer Documentation Checklist
- Give the Model the Target's Contract, Not Similar Solutions
- Handoff Skill: Structured Context Transfer Between Agent Sessions
- How AI Engines Cite — ChatGPT, Perplexity, Claude, Gemini
- Incident Log Investigation Skill: Parallel Queries
- Instruction-Guided Code Completion: Controlling What Models Generate
- Inversion Analysis: Surface Capabilities Competitors Cannot Replicate
- Issue Requirements Preprocessing: Structured Input Before Code Generation
- Manual Compaction Strategy for Dumb Zone Mitigation
- Measuring GEO Performance for AI Search Visibility
- Post-Compaction Re-read Protocol for Agent Continuity
- Prompt Governance via PRs: Reviewable AI Behavior
- Prompt Transpilation: Instructions as Build Artifacts
- Schema and Structured Data for GEO — AI Citation Guide
- Separating Exposure From Selection in GEO Measurement
- Stuck-Loop Recovery: Detecting and Escaping Non-Converging Agent Loops
- The AX Stack: A Layered Model of an AI Coding Agent's Prompt-to-Compile Path
- Three Reasoning Spaces: Plan-Bead-Code Phase Gates
- Topical Authority — Entity Coverage for AI Citation
- llms.txt: Full Specification, Adoption, and Limitations
testing-verification¶
- AI-Powered Vulnerability Triage for AI Agent Development
- AIRA: Inspection Framework for AI-Generated Code
- AX Evals: Measure the Agent-Facing Surface, Not the Model
- Action-Class Decomposition for Tool-Calling Evals
- Action-Graded Severity for Agent Red-Team Outcomes
- Adaptive Evaluation of Out-of-Band Prompt-Injection Defenses
- Adaptive Generate-Rank-Verify Under Costly Verification
- Adaptive Validation Task Selection
- Adversarial Multi-Model Development Pipeline (VSDD)
- Against-Prior Accuracy: Score the Rules That Fight Defaults
- Agent Approval Authority in Code Review
- Agent Determinism Moves to the Last Unconstrained Axis
- Agent Self-Review Loop for Iterative Self-Improvement
- Agent-Assisted Code Review: Agents as PR First Pass
- Agent-Authored Eval Suites From Repo Context and Traces
- Agent-Driven Deployment: What to Delegate and What to Gate
- Agent-Driven Eval Flywheel: Prove a Fix Generalizes
- Agent-Driven Fuzzing with Human-Gated Crash Triage
- Agent-Generated Verification Reports: A Structured Round-Trip for Human Review
- Agent-Recorded Video Demos as a Verification Artifact
- Agent-Resolved Review Threads
- Agentic Code Review Architecture With Tool-Calling
- Agentic Review Comment Acceptance
- Ambiguity Stability as a Model-Selection Criterion
- Answer-Reachable Eval Environments
- Anti-Reward-Hacking: Rubrics That Resist Gaming
- Artifact-Only Verification Hides Skipped Skill Steps
- Assertion-Free Test Theater in Agent-Authored Patches
- Assuming Loaded Skills Stay Enforced in Long Contexts
- Assumption Propagation: Compounding Agent Misunderstandings
- Audit Your Test Suite With an Agent, Then Certify Each Flag
- Audit the Noise Floor Before Trusting a Benchmark Gap
- Audit-Budget Allocation for Agent Fleets
- Auditing Agent Tool Chains for Silent Partial Success
- Baseline-Aware Test Evaluation for Multi-Agent Issue Resolution (Phoenix)
- Behavior Specs: Grading the Trajectory, Not the Result
- Behavior-Partitioned Security Tests as Executable Specs
- Behavioral Specification Elicitation Before Synthesis (SpecFirst)
- Behavioral Testing for Non-Deterministic AI Agents
- Benchmark Contamination as Eval Risk
- Benchmark Poisoning of Self-Modifying Coding Agents
- Benchmark-Driven Tool Selection for Code Generation
- Black-Box Agent Risk Scoring by Domain
- Blind Tool Deference: Agents Parroting Callable Tools
- Bounded Repair-Loop Iterations
- Bounded Tool Surfaces for Code Review Agents
- Browser as Agent Action Space
- Bug-Discriminating Validation Evidence for Repair Agents (BSG-VA)
- Building Agent Eval Environments With a World Spec
- Canary Tools for Diagnosing Tool-Selection Reasoning
- CausalFlow: Counterfactual Repair for Failed Agent Trajectories
- Chain-of-Verification for Coding Agents
- Cheaper-Per-Token Model Upgrades That Cost More Per Task
- Choosing the Judge Model That Grades Your Agent Evals
- Claim-to-Evidence Trace Graphs for Auditing Agent Runs
- Classification Before Repair in an Analyzer Backlog
- Claude Code Review
- Close the Attack-to-Fix Loop: Adversarially Train Agent
- CoT Robustness in Code Generation
- Code Health as a Signal for Agent-Generated Test Quality
- Committee Review Pattern for Multi-Agent Code Review
- Comparative Judging for Agent Configuration Ranking
- Completion Failure Taxonomy: Why Code Suggestions Miss
- Completion Summary as the Oversight Surface
- ComplexMCP: Three Bottlenecks in Large Interdependent Tool Sandboxes
- Component-Isolated Memory Stress Testing for LLM Agents
- Constraint Decay in Backend Code Generation
- Context Quality as a Leading Indicator of Agent Reliability
- Contract-Domain Tracing for Rubric Credit
- Control Lexical Leakage in Agent-Memory Retrieval Evals (Entity-Collision)
- Controlled Benchmark Rewriting for Agent Safety Judgment
- Corpus-Level Trace Diagnostics for LLM Agents
- Coverage-Guided Agents for Fuzz Harness Generation
- Coverage-Guided Fuzzing for Multi-Agent LLM Systems (FLARE)
- Covert Success Rate for Indirect Prompt Injection
- Cross-Framework Signal Semantics: Re-Measure Borrowed Trajectory Rules
- Cumulative-Best Reporting Hides Repair-Loop Security Regressions
- Cut-Point Replay: Test a Fix Against a Recorded Agent Run
- Data Fidelity Guardrails: Preventing Agent Data Mutation
- Decomposed Red-Teaming for Agent Monitors
- Decomposing Agent Output Variability by Layer (Sampling vs Orchestration State)
- Defense-in-Depth Against Coding Agent Fabrication (Honesty Harness)
- Deletion Avoidance: Agents That Guard Code Instead of Removing It
- Deletion and Cost Rules for an Evolving Agent Harness
- Demand-Driven Repository Auditing
- Demo-to-Production Gap: When Demos Hide Real Costs
- Density-Normalized Quality Metrics Mask AI-Driven Code Growth
- Dependency Gap Validation for AI-Generated Code
- Dependency Inlining Erodes SBOM and License Provenance
- Deriving a Specification From Buggy Code Before Generating Tests
- Destructive-Failure Mechanism Attribution by Mitigation Owner (ClayBuddy Three)
- Detecting Self-Preference in a Single LLM Judge
- Deterministic Guardrails Around Probabilistic Agents
- Deterministic Precondition Gates for Tool-Using Agents
- Diff-Based Review Over Output Review
- Diff-Coverage Gating for Agent-Authored Pull Requests
- Discovery-Only Refactor Pass: Surface Candidates Before Touching Code
- Distillation-Induced Similarity Metrics for Tool-Use Agents
- Dominator-Graph Trajectory Invariants for Non-Deterministic Agents
- Downstream Disclosure Coordination for Agent-Found Defects
- Dual Executable Specifications for Long-Horizon Features
- Emulate Agent-Experience Changes Before Shipping
- Emulated APIs for Agent Skill Evals
- Engineering: Tools, Review, Verification, Security, and Observability
- Enumerate Every Visible Exit Before Scoring Containment
- Equivalence Testing for Agent Configuration Changes
- Escalation Channels: A Reporting Tool Instead of a Reward Hack
- Eval Awareness: Designing Evals Agents Cannot Recognize
- Eval Blind Spots: Structural Gaps in Measurement Methodology
- Eval Difficulty as a Product Smell
- Eval Engineering (Training Module)
- Eval Environment Containment for Cyber-Capable Agents
- Eval Strategy by Agent Generation: A Structure-to-Eval Locator
- Eval-Driven Development Training for AI Agent Teams
- Eval-Driven Development: Write Evals Before Building Agent
- Evaluator Templates: Portable Primitives for Agent Eval Suites
- Evaluator-Optimizer Pattern for AI Agent Development
- Event Sourcing for Agents: Separating Cognitive Intention
- Evidence-Bundled Agent PRs: Sizing the Reviewer's Effort
- Evidence-Chain Run Logs: Bracket the Reported Symptom
- Evidence-Conditioned Execution: Gate Edits on Observations
- Evidence-Gated Lifecycle Control for Coding Agents (Proof-or-Stop)
- Evidence-Grounded Disagreement in Agentic Code Review (Adversarial Review)
- Exactly-Once Enforcement Layer: Model, Harness or Contract
- Execution Budgeting in Agentic Program Repair
- Execution Lineage: DAG of Artifacts vs Agent Loops
- Explained Feedback for LLM Vulnerability Repair
- Failure-Driven Iteration for Improving Agent Workflows
- False-Pass Liability Decides the Next Harness Component
- Feature List Files for Reliable AI Agent Development
- Five-Failure-Layers Diagnostic: Attribute Before Swapping the Model
- Five-Pass Blunder Hunt: Repeated Critique Passes for Plans
- Fleet Harness Attribution: Pinning the Model to Compare Whole Harnesses
- Frozen Task Sets for Affordable Agent A/B Testing
- Frozen-Base Task Mining for Repository Instruction Files
- Frozen-Stimulus Panels for Cross-Vendor Behavior Measurement
- Function-Level Debugger Interfaces for Coding Agents
- Gate Best-of-k Selection on Compliance Before Score
- Generating Tests From Agent-Written Code (Code-First Oracle Bias)
- Generative Provenance Records for Tool-Using Agents
- Golden Journeys: Restartability as a First-Class Verification Primitive
- Golden Query Pairs as Continuous Regression Tests for Agents
- Governed Sources of Truth for Analytics Agents (Structure Over Access)
- Grade Agent Outcomes, Not Execution Paths
- Grading Strategies for Eval-Driven Development
- Handoff-Boundary Fault Injection (llmmas-otel)
- Happy Path Bias: How AI Agents Skip Error Handling
- Hardening Agent Evals for Production-Grade Reliability
- Harness Bug Detection Patterns
- Harness Composition for Scaled Security Audits
- Harness Hill-Climbing: Eval-Driven Iterative Improvement of Agent Harnesses
- Hash-Pinned Command Approval for Non-Interactive Plugin Installs
- Head-to-Head Evaluation of Competing MCP Servers
- Held-Out Tasks as a Harness Shortcut Defense
- Homogeneous Debate Panels as a Groundedness Quality Lever
- Human-Review-Driven Curation of Golden Eval Datasets
- In-Place Atomic Replacement for Agent-Driven Ports
- Incident-to-Eval Synthesis: Production Failures as Evals
- Incremental Verification: Check at Each Step, Not at the End
- Independent Test Generation in Multi-Agent Code Systems
- Inference-Time Tool-Call Reviewer: Pre-Execution Feedback for Tool-Calling Agents
- Inferring Agent Failure from Conversation Evidence (Perceived Error)
- Informed Abstention as a Tool-Boundary Runtime Gate
- Interaction-Pattern Evaluation for Agentic PRs
- Isometric Harness Ablation: Rank Subsystem Investment by Removing One at a Time
- Knowledge Cutoff as a Documentation Boundary
- LLM API Fault Injection at the HTTP Layer (AgentChaos)
- LLM Agent Bug Fix Taxonomy: 23 Fix Patterns from 930 Real Bugs
- LLM Code Review Overcorrection for AI Agent Development
- LLM Self-Review Failure in Code Modernization Tasks
- LLM Static Verification Against Natural-Language Requirements
- LLM Support During the First Detection Pass
- LLM-Driven Benchmark Auditing
- LLM-as-Judge Evaluation with Human Spot-Checking
- Layered Accuracy Defense for Reliable Agent Outputs
- Layered Oracle Stack for Agent IaC Security Repair (TerraProbe)
- Learned Prefix Monitors for Agent Traces
- Learning Execution Guardrails from Agent Failure Traces
- Macro Evals for Agentic Systems: Population-Level Behavior Patterns
- Markov-Chain Reliability for LLM Agents: Audit the Abstraction Before You Trust the Metric
- Measure the Judge Before You Freeze a Gate on It
- Measuring Synthetic Eval Data Quality (SynAE)
- Measuring the Verification Tax on Agent Output
- Meta-Evaluate the LLM Judge Before Trusting Rubric Verdicts
- Minimum-Cost Evidence Selection for Agent Changes (Assurance Envelopes)
- Monolith-to-Sub-Agents Refactor: Five Lessons from a Brittle Prototype
- Multi-Agent RAG for Spec-to-Test Automation
- Multi-Agent Shared State Isolation Anomalies
- Multi-Layer Specification Redundancy as a Robustness Budget
- Multi-Run, Shuffled-Order Evaluation for Self-Improving Agents
- Multi-Turn Conversation Evaluation: Per-Turn and Trace-Level Scoring Together
- Mutation Testing as a Quality Gate for AI-Generated Test Suites
- Mutation Testing for LLM Judges: Scoring an Evaluator on Injected Defects
- Name the Check That Passed Before Accepting AI Code
- Narrative Problem Reformulation for Code Generation
- Natural-Language Documentation as a Code-Review Intermediate (Verifiable Literate Programming)
- Non-Compensatory Readiness Gates Before Agent Release
- Nonstandard Errors in AI Agents: Model-Family Variance
- Observability Feedback Loop: A 7-Step Debug Runbook for Agents
- Observation Contract Preservation in Tool-Augmented Agents
- Offline Evaluation as an Integration Test for LLM Features
- Offline Trajectory Replay for Multi-Agent Workflow Debugging
- Oracle-Based Task Decomposition for AI Agent Development
- Oracle-Gated Delegation Beyond Your Domain Expertise
- Outcome Monitors: Recovery Affordances for Tool Failures
- Overeager-Behavior Elicitation: Scope + Trap Fragments as a Diagnostic for Out-of-Scope Tool Calls
- Overtrusting Human Sign-Off on Generated Assertions
- PASS@(k,T): Evaluate RL for Agents Along Sampling and Interaction Depth
- Parallel Polyglot Ports as a Spec-Ambiguity Oracle
- Per-Attempt Sandboxes for Agents That Change the Filesystem
- Per-Change Deploy Monitors: Report the Verdict, Don't Act on It
- Per-Layer Suppression Accounting in Acceptance Gates
- Per-Line Requirement Citations for Hallucination Detection
- Perceived Model Degradation: Why Vibes Are Not Evals
- Phantom Symbol Detection for LLM API Migration
- Planted-Bug Methodology: Deliberate Bugs as Observability Calibration
- Policy-Graded Evaluation of Coding Agents
- Pooled-Evidence Factuality Checks for MCP Agents (Cross-Source Conflation)
- Post-Merge Fix Signals for Agent Merges
- Pre-Change Impact Analysis: Dependency Maps That Prevent Agent Regressions
- Pre-Completion Checklists for AI Agent Development
- Pre-Execution Failure Scoring with a Draft Model (Speculative Uncertainty)
- Pre-Generation Complexity Scoring for Code Reliability
- Prebuilt Agent Monitoring Dashboard
- Precise Debugging: Measure Edit Precision, Not Just Test Pass Rate
- Predicting Reviewable Code: Pre-Flagging Functions Reviewers Will Delete
- Premature Completion: Agents That Declare Success Too Early
- Prescribing TDD Inside the Agent Loop (Process Theater)
- Probing Unstated Constraints in Generated Code (Intent Violation Rate)
- Profile Your Agent Test Suite Against Measured Practice
- Profiler-Guided Optimization Loops for Coding Agents
- Protecting the Test Oracle From the Agent
- Purpose-Built Eval Suites for Model and Harness Swaps
- QA Session to Issues Pipeline for AI Agent Development
- Quality Score Rubric and Simplification Log for Agent Harnesses
- RAG/Agent Reliability Problem Map: 16-Domain Failure Taxonomy
- RL-Trained Automated Red Teamers for Prompt Injection Discovery
- Rank Resolution: Reading a Converged Coding-Agent Leaderboard
- Re-Run the Original Test Suite After Every Refinement Turn
- Reading a Coding-Agent Vendor's Security Certificate
- Recover the Six Measurement Choices Behind an Attack Success Rate
- Red-Green-Refactor with Agents: Tests as the Spec
- Red-Team Your Blocking Monitor Before You Trust It
- Reducing Fixed CI Overhead Before Adding Shards
- Refactoring Runaway: Tangled Refactorings in Agent Patches
- Rented Sandboxes for Coding-Agent Benchmark Runs
- Repository Perturbation as Context-Reasoning Diagnosis (RepoMirage)
- Reproduce-Before-Report Verification Gate
- Retry-Switch-Abstain: A Runtime Tool-Recovery Policy
- Reverse-Engineered Executable Specifications for Agentic Program Repair
- Review Constraint Tests as a Second Acceptance Gate
- Review-Comment-Derived Benchmarks for Code Review Agents
- Review-Feedback-to-Rule Loop: Promoting Recurring PR Comments into Harness Rules
- Review-Then-Implement Loop for AI Agent Development
- Reviewer Habituation in Agent PR Review
- Reviewer Theme Distribution Audit for AI Code Review
- Risk-Based Shipping: Review by Risk Matrix, Not by Default
- Risk-Based Task Sizing for Agent Verification Depth
- Risk-Score Threshold Calibration for Auto-Approval
- Root Causes of Vibe-Coded Application Vulnerabilities
- Route-Parity Auditing of Agent Safeguards
- RubricRefine: Pre-Execution Rubric Refinement for Code-Mode Tool Use
- Runnable Documentation as Agent Verification
- Scoped Browser DevTools Access for Runtime Diagnosis
- Security Drift in Iterative LLM Code Refinement
- Seed-Variance Reporting and Measurable-Range Eval Design
- Semantic Collapse Under Underspecified Prompts
- Semantic Validation for Schema-Valid Agent Output
- Serving-Stack Confounds in Tool-Call Evaluation
- Severity-Stratified Evaluation of Security Prompts
- Shallow Agent Test Coverage from Premature Termination (Lazy Generation)
- Signal Over Volume in AI Review for AI Agent Development
- Silent Adoption of Corrupted Tool Returns by Agents
- Simulation and Replay Testing for Agent Verification
- Size Agent Comparisons by Run-to-Run Variance
- Skill Eval Loop
- Skill Evals: Measuring Skill Quality as a Dataset-Graded Unit
- Skill Lift: Measuring What a Skill Adds at Runtime
- Skill Specification Violation Fuzzing
- Skill Test Coverage as a Release Gate
- Skill-Use Gates: Trigger, Compliance and Boundary
- Slop Detectors Fail as Per-Item Review Gates
- Solver-Externalized Constraint Reasoning (MaxSAT/SMT Encoding)
- Source-Grounded Test Plan with Pre-Action Assertion Annotation
- Spec-Derived Execution as a Correctness Oracle
- Spec-Driven Test Generation: Contract Coverage Is the Lever
- Specification Authority Boundary: Agents Propose, the Runtime Commits
- Specification-First Convergence Without a Test Oracle
- Specification-Grounded Test Writing
- Specification-Path Testing: Same Contract, Different History
- Staged Evidence Gates for Agentic Program Repair
- Staged Literal Porting with a Per-Stage Numeric Oracle
- State-Bound Evidence and Typed Revision Contracts for Repair Loops
- Stateful Agent Evals via State Snapshots and Transition Assertions
- Static Difficulty Estimation for Agent Issue Triage
- Step-by-Step: Building Your First Eval-Driven Feature
- Stochastic-Deterministic Boundary as First-Class Contract
- Strained Coherence as a Pre-Failure Signal in Agent Trajectories
- Structural Coverage Criteria for Agent Workflows
- Structure-Aware Diff Labeling with Two-Stage LLM Pipelines
- Structured Output Constraints: Reducing Hallucination
- Symptom-First Bug Triage for Agent Code
- Symptom-Reduction-as-Root-Cause: Why Oracle Tests Alone Miss Architectural Drift
- TDD Interaction Models: Throughput Versus Test Quality
- Tail Control for Agent Workflows: Engineering for the Failure Tail, Not the Average
- Task Alignment: The Selective-Compliance Gap Benchmarks Miss
- Task Category as the Security Review Routing Key
- Task Completion as Tool Certification (Silent Tool Rot)
- Task-Uniform Agent Permissions Ignore Where Failures Land
- Test Harness Design for LLM Context Windows
- Test Oracles That Read Their Expectation From the Code
- Test-Driven Agent Development: Tests as Spec and Guardrail
- Test-Driven Intent Clarification: Tests as Intermediate Alignment Artifacts
- The Eval-First Development Loop for AI Agent Features
- The Model Preference Fallacy in Comparison Content
- The Patchwork Problem in LLM-Generated Code
- The Productivity-Experience Paradox in AI-Assisted Development
- The Synthetic Ground Truth Fallacy in Agent Evaluation
- The Test Homogenization Trap: When LLM-Generated Tests Mirror Model Blind Spots
- The Three Loops of Agentic Coding: A Diagnostic Vocabulary
- Tiered Code Review: AI-First with Human Escalation
- Timeout Oracles for Agent-Written Code
- Tool Operability: Interfaces That Survive a Lost Response
- Tool-Call Success as Workflow Effect Evidence
- Tool-Schema Bias: Measure the Interface Before You Ship
- Tool-Use Sim-to-Real Perturbation Taxonomy
- Traces Need Feedback to Power Learning
- Trajectory Decomposition: Diagnose Where Coding Agents Fail
- Trajectory Pre-Filter for Failure Diagnosis (TrajAudit)
- Trajectory Projection: A Flattened View of Agent Traces
- Trajectory-Aware Benchmark Subset Selection for Agents
- Transcript-Measured Review Coverage
- Treating a Clean Final State as Boundary-Compliance Evidence
- Treating a Clean Static Scan as Security Evidence
- Trust Without Verify: Skipping Agent Output Checks
- Trusting Claimed Prior Approval in Agent Review Gates
- Trusting Human Review to Catch Deliberate Agent Sabotage
- Typed Generation Contracts for Grounded Extraction
- Unstated-Contract Bugs: Sort Tickets by Information Gap
- Usability Pressure as a Silent Security-Regression Vector
- Using the Agent to Analyze Its Own Evaluation Transcripts
- Validity-Estimate Stopping for Noisy Verify-Repair Loops (VRR-Stop)
- Variance-Based RL Sample Selection
- Velocity-Quality Asymmetry: Why AI Speed Gains Fade
- Verification Capacity Saturation: Three Levers, One Default
- Verification Capacity as the Agent Quality Ceiling
- Verification Ledger for Tracking Agent Output Quality
- Verification Surface: Match the Tool to the Failure
- Verification-Centric Development for AI-Generated Code
- Verification: Testing, Evals, and Guardrails for Agents
- Verifier-Driven Parallel Coding Agents (Glite ARF)
- Verify Agent Diagnoses and Fix Proposals Before Acting
- Verify Observability in Agent-Generated Code
- Verify-Gated Completion as Admission Control
- Verifying LLM-Generated Cryptographic Code
- Vibe Coding: Outcome-Oriented Agent-Assisted Development
- What Evals Are and Why AI Agents Need Them for Quality
- Writing Your First Agent Evaluation Suite from Scratch
- Zero Violations Is Not Evidence Your Hook Works
- pass@k and pass^k: Capability and Consistency Metrics
token-engineering¶
- Code Cleanliness as an Agent Cost Lever
- Cost-Aware Agent Design: Route by Complexity, Not Habit
- Cost-Quality Pareto Measurement for Agent Configurations
- GitHub's Copilot Cost Levers at Constant Task Quality
- Harness-Controlled Token Economics (The Harness Effect)
- Language Choice as an Agent Token-Cost Lever
- Line-Anchored Feedback: Deliver Change Requests as Inline Comments
- Measuring Refactoring Payback in Tokens
- Pricier-Per-Token Models That Cost Less Per Task
- Probe-Run Calibration for Predicting Agent Token Spend
- Request Shaping to Cut Wasted Agent Turns
- Routing Decision Framework: Which Routing Pattern Fits Which Signal
- Routing Dependency Updates to Repair Agents by Budget
- Splitting an Agent Token Budget at the Scaling Inflection Point
- Static-First Shell Command Gating with Selective Escalation
- Temporal Token Routing: Batch and Flex Tiers for Non-Urgent Work
- The Token Price Index Fallacy in Agent Cost Planning
- Token Engineering: Fewer, Cheaper Tokens Without Losing Quality
- Token-Cost Profiling and Reduction for Always-On Agentic Workflows
- Token-Efficient Code Generation: Structural Beats Prompting
- Token-Efficient Tool Design: Tools That Don't Eat Your Context
- Tokenizer Swap Tax: Budgeting for Model Migrations That Change Token Counts
tool-agnostic¶
- A Governance Framework for Production Agents
- A2UI: Framework-Agnostic Generative UI Standard for Agents
- ACDL: A Language for Describing Agentic LLM Contexts
- ACID for Agent Repository State
- AGENTS.md as a Table of Contents, Not an Encyclopedia
- AGENTS.md: Project-Level README for AI Coding Agents
- AI Abundance Reshapes Software Engineering Identity
- AI Adoption Footprint: The Segmented Shape of Engineering Orgs
- AI Agent Development Anti-Patterns and Failure Modes
- AI Agents in CI/CD with Elevated Permissions and Untrusted Content (GitInject)
- AI Bot CI/CD Workflow Reliability by Agent
- AI Crawler Policy: robots.txt for the Three-Tier Crawler Landscape
- AI Label as Reviewer Attention Redistribution
- AI Slop as a Process Problem: Encoding Quality Standards as Pipeline Gates
- AI-Powered Vulnerability Triage for AI Agent Development
- AIRA: Inspection Framework for AI-Generated Code
- AOCI: Symbolic-Semantic Repository Indexing
- AST-Grounded Critic Loop for Documentation Maintenance
- AX Evals: Measure the Agent-Facing Surface, Not the Model
- AX/UX/DX Triad: Three Experience Layers in Agent Systems
- About
- Absorbing Entry-Level Work into Senior-Agent Workflows
- Abstraction Bloat in AI Agent-Generated Code Output
- Accumulated Behavioral Rules from Review Feedback
- Acknowledged-Debt Ledger with Next-Trigger Conditions
- Action-Audit Divergence: A Four-Mode Taxonomy for Runtime Hardening
- Action-Class Decomposition for Tool-Calling Evals
- Action-Gated Context Trimming for Long-Horizon Agents
- Action-Graded Severity for Agent Red-Team Outcomes
- Action-Selector Pattern: LLM as Intent Decoder with Deterministic Execution
- Adapting AI Assistants to Developer Interaction Style
- Adaptive Evaluation of Out-of-Band Prompt-Injection Defenses
- Adaptive Generate-Rank-Verify Under Costly Verification
- Adaptive Sandbox Fan-Out Controller
- Adaptive Validation Task Selection
- Addressable Recall Compaction: Compact to Citations, Not Summaries
- Adversarial Multi-Model Development Pipeline (VSDD)
- Adversarial-Only Threat Modeling for Agent Data Leakage
- Advisory Prompts Distilled from Reasoning Traces
- Against-Prior Accuracy: Score the Rules That Fight Defaults
- Agent Approval Laundering: Effects Beyond the Named Command
- Agent Backpressure: Automated Feedback for Self-Correction
- Agent Cards: Capability Discovery Standard for AI Agents
- Agent Chat History as a First-Class Artifact
- Agent Circuit Breaker
- Agent Commit Attribution: Signed Commits and Agent Identity
- Agent Composition Patterns for Multi-Agent Workflows
- Agent Config as a Managed Supply Chain: Hashing and Pinning
- Agent Context File Evolution: Treating ACFs as Configuration Code
- Agent Debug Log Panel: Chronological Event Inspection for Session Debugging
- Agent Debugging: Diagnosing Bad Agent Output
- Agent Definition Formats: How Tools Define Agent Behavior
- Agent Determinism Moves to the Last Unconstrained Axis
- Agent Development Lifecycle for Agent Products
- Agent Event Streaming: Consumer Contract Above the Tokens
- Agent Extension Conflicts: When Installed Skills and MCP Servers Fight Each Other
- Agent Failure Trajectories and the Recovery Window
- Agent Governance Plane: Audit Events and Message-Content Surfaces
- Agent Handoff Protocols: Passing Work Between Agents
- Agent Harness: Initializer and Coding Agent Pattern
- Agent Headcount as a Vanity Metric
- Agent JIT Compilation: Compile Tasks Into Executable Plans
- Agent Loop Go/No-Go: When Looping Earns Its Cost
- Agent Loop Middleware — Safety Nets and Message Injection
- Agent Memory Patterns: Learning Across Conversations
- Agent Network Egress Policy: Admin-Controlled Domain Allow/Deny
- Agent PR Volume vs. Value: The Productivity Paradox
- Agent Plugins: Portable Packaging With Client-Defined Trust
- Agent Pushback Protocol for Managing Disagreements
- Agent Retrieval Provenance as an Audit Control
- Agent Rewrites Lose Meaning: The Ownership Rule for AI-Assisted Writing
- Agent Runtime Middleware: Per-Call Interception Pipeline
- Agent Self-Review Loop for Iterative Self-Improvement
- Agent Skills: A Cross-Tool Task Knowledge Standard
- Agent Sprawl: Unmanaged Sub-Agent and Skill Proliferation
- Agent Terminology Disambiguation for AI Coding Systems
- Agent as Tool vs Handoff: Who Keeps the Conversation
- Agent-Assisted Code Review: Agents as PR First Pass
- Agent-Authored Eval Suites From Repo Context and Traces
- Agent-Authored Messages as a Deferred Exfiltration Channel
- Agent-Authored PR Integration: Collaboration Signals That Determine Merge Success
- Agent-Aware CLI Behavior via Environment Variable
- Agent-Client Admission Control for Agentic Traffic
- Agent-Computer Interface (ACI): Tool Design as UX Discipline
- Agent-Discoverable Slash Commands
- Agent-Driven Eval Flywheel: Prove a Fix Generalizes
- Agent-Driven Fuzzing with Human-Gated Crash Triage
- Agent-Driven Greenfield Product Development from Scratch
- Agent-Driven PR Slicing
- Agent-Emitted Dependency Version Ranges Widen the Supply-Chain Attack Surface
- Agent-First Software Design for AI Agent Development
- Agent-Generated Code Maintenance Asymmetry
- Agent-Generated Verification Reports: A Structured Round-Trip for Human Review
- Agent-Initiated Rubric-Gated Self-Compaction (SelfCompact)
- Agent-Laundered Bug Reports
- Agent-Led Dev-Environment Iteration with Validation and Rollback
- Agent-Native Filesystems: Gating Effects, Not Commands
- Agent-Operable Interface Design (Affora)
- Agent-Powered Codebase Q&A and Onboarding Workflow
- Agent-Proposed Merge Resolution
- Agent-Reactive Bugs at the Model-Harness Boundary
- Agent-Readiness Discovery Surfaces for Docs Sites
- Agent-Ready Bug Reports for Software Repair Agents
- Agent-Ready Data Architecture for Analytics Agents
- Agent-Recorded Video Demos as a Verification Artifact
- Agent-Trace Data Layer: Storage for Hours-Long Traces
- Agent-Tuned Code Search: Retrieval Built for the Loop
- Agent-to-Agent (A2A) Protocol for AI Agent Development
- Agentic AI Architecture: From Prompt to Goal-Directed
- Agentic Code Review Patterns and Review Architectures
- Agentic Detection and Response at the MCP Boundary
- Agentic Education: Persona Progression for Teaching AI Coding Tools
- Agentic Flywheel: Self-Improving Agent Systems
- Agentic Framework Landscape: When Each Framework Fits
- Agentic Pattern Vocabulary Crosswalk
- Agentic Resource Discovery: Federated Pre-Invocation Search
- Agentic Review Comment Acceptance
- Agentic Skill Decay: Which Capabilities Erode Under Agent Delegation
- Agentic-Agile: Adapting Agile Rituals for Agent Work
- Agentless vs Autonomous: When Simple Beats Complex
- Agents vs Commands: Separation of Role and Workflow
- Aggregation Bounds for Agent Authorization
- Always-On Agentic PR Security Review
- Ambiguity Stability as a Model-Selection Criterion
- Ambition Scaling: Moving the Target as Model Capability Increases
- An Explicit Update Boundary for Agent Self-State
- Answer-First Writing: Structure Content for AI Retrieval
- Answer-Reachable Eval Environments
- Anthropic's Effective Agents Framework: A Pattern Map
- Anti-Reward-Hacking: Rubrics That Resist Gaming
- App-Window Snapshot as Agent Context
- Applying Coding Agents to Non-Code Tasks
- Approval Gate Granularity in Agent Pipelines
- Architecting a Central Repo for Shared Agent Standards
- Artifact-Driven Workflow Compilation for Agent Execution
- Artifact-Level Accountability Mapping for Agent Workflows
- Artifact-Only Verification Hides Skipped Skill Steps
- Ask the Docs
- Ask-Everything Permission Policies Protect Less than Per-Action Approval
- Assertion Density — Stats and Quotes Over Vague Claims
- Assertion-Free Test Theater in Agent-Authored Patches
- Assuming Agent Interchangeability in Long-Running Teams
- Assuming Loaded Skills Stay Enforced in Long Contexts
- Assumption Propagation: Compounding Agent Misunderstandings
- Async Non-Blocking Subagent Dispatch
- Asynchronous Agent I/O and Speculative Tool Calling
- Atomic Pages and Chunking — One Concept Per Page for RAG
- Attention Latch: When Agents Stay Anchored to Stale Instructions
- Attention Sinks: Why First Tokens Always Win
- Attributed Cache Misses: Reading Why a Prefix Diverged
- Audit Your Test Suite With an Agent, Then Certify Each Flag
- Audit the Noise Floor Before Trusting a Benchmark Gap
- Audit-Budget Allocation for Agent Fleets
- Auditing Agent Tool Chains for Silent Partial Success
- Auth-Isolation as the MCP-vs-CLI Selection Heuristic
- Author-to-Reviewer Role Inversion in AI-Assisted Teams
- Authority Confusion: Untrusted Context Must Not Authorize Side Effects
- Authorization Continuity Across Agent Mutation
- Auto-Merging a Wiki Agent's Documentation Pull Requests
- Auto-Triage Workflow: Bug-Monitoring Agent that Connects Related Reports and Opens Fix PRs
- Autonomous Research Loops: Loops That Know When to Stop
- BYOK Model Token Visibility: Closing the Observability Gap on Self-Hosted Routes
- Background Todo Agent: Offload Plan Maintenance to a Lightweight Model
- Backlog Triage as a Named Agent Skill
- Baseline-Aware Test Evaluation for Multi-Agent Issue Resolution (Phoenix)
- Batch File Operations via Bash Scripts for AI Agents
- Batched Suggestion Application: Bulk-Apply Agent Fixes on PRs
- Behavior Specs: Grading the Trajectory, Not the Result
- Behavior-Partitioned Security Tests as Executable Specs
- Behavioral Drivers of Coding Agent Success and Failure
- Behavioral Firewall for Tool-Call Trajectories
- Behavioral Specification Elicitation Before Synthesis (SpecFirst)
- Behavioral Testing for Non-Deterministic AI Agents
- Belief Inertia After Tool-Map Drift in AI Agents
- Benchmark Contamination as Eval Risk
- Benchmark Poisoning of Self-Modifying Coding Agents
- Benchmark-Driven Tool Selection for Code Generation
- Benign Skill Wording That Steers Package Hallucination (Neutral Prompting Attack)
- Binding an Agent's Effect to the Approval It Claims
- Black-Box Agent Risk Scoring by Domain
- Black-Box Probing in the Agent Build Loop: When It Pays
- Blaming the Model for Scaffolding-Driven Quality Regressions
- Blast Radius Containment: Least Privilege for AI Agents
- Blind Resampling Over Self-Repair in Small Code Models
- Blind Tool Deference: Agents Parroting Callable Tools
- Bootstrapping Coding Agents: The Specification Is the Program
- Boring Technology Bias: When Agents Recommend by Popularity
- Bounded Agent Steps Inside a Deterministic Workflow
- Bounded Batch Dispatch for Parallel Agent Execution
- Bounded Repair-Loop Iterations
- Bounded Tool Surfaces for Code Review Agents
- Bounding a Headless Codex Run Without a Turn Cap
- Brownfield to Agent-First: Repo Maturity Framework
- Browser Automation as a Research Tool: Bypassing Bot Detection
- Browser Sandbox for Agent-Generated HTML (Sandboxed Iframe + Immutable CSP)
- Browser as Agent Action Space
- Budgeted Verification of Inherited Agent Constraints
- Bug-Class Hints as Exploit Input for Coding Agents
- Bug-Discriminating Validation Evidence for Repair Agents (BSG-VA)
- Building Agent Eval Environments With a World Spec
- Building Custom Agents from Substrate to Production (Agents All the Way Down)
- Burn the Boats — Commitment-Forcing Deprecation
- CARE: Three-Party Stage-Gated Engineering of LLM Agents
- CLI Scripts as Agent Tools: Return Only What Matters
- CLI-First Skill Design
- CLI-IDE-GitHub Context Ladder for AI Agent Development
- CRA-Only Review and the Merge Rate Gap
- Cache-Prefix Staggering for Sibling Agent Fan-Out
- Cache-Safe Routing Boundaries: Where a Router May Act
- Calibrated Deciders for In-Loop Agent Decisions
- Calibrated Early Termination and Warm Restart for Agent Runs (FailFast-RestartSmart)
- Caller-Actionable Error Steps in Tool Responses
- Canary Rollout for Agent Policy Changes
- Canary Tools for Diagnosing Tool-Selection Reasoning
- Capability Declarations for Agents That Act on Data
- Capability-Additive Code Interpreters for Untrusted Agent Code
- Capability-Pegged Security Re-Scans: Reviewing Unchanged Code When the Scanner Improves
- Cargo Cult Agent Setup: Copying Without Understanding
- Catastrophic Remembering: Instruction Files That Only Grow
- CausalFlow: Counterfactual Repair for Failed Agent Trajectories
- Centralized LLM Gateway for Per-Dimension Agent Budgets
- Chain-of-Thought Reasoning Fallacy: Traces Are Not Truth
- Chain-of-Verification for Coding Agents
- Chance-Corrected Shortlist Depth Sizing for Tool Retrieval (Bits-over-Random)
- Chat-Platform Agent Delegation: Invoking Cloud Coding Agents from Team Channels
- Cheaper-Per-Token Model Upgrades That Cost More Per Task
- Choosing a Compression Budget for Agent Control Context
- Choosing a Skill Loading Method for Agents
- Choosing an Agent Tool Interface: Shell or Typed Catalog
- Choosing an Integration Layer for an Embedded Agent Harness
- Choosing the Judge Model That Grades Your Agent Evals
- Choosing the Right Surface for a Coding Agent Task
- Chunking Strategy for RAG-Based Code Completion
- Circuit Breakers for Agent Loops
- Claim-Scoped Invalidation for Agent Memory
- Claim-to-Evidence Trace Graphs for Auditing Agent Runs
- Clarification Mode Amplifies Prompt Injection
- Classical SE Patterns as Agent Design Analogues
- Classifier-Gated Auto-Permission for Cloud-IDE Coding Agents
- Classifier-Subagent Run Mode for Per-Call Permission Routing
- Classifying and Auto-Correcting Coding Agent Misbehaviors (Wink)
- Clock-In / Clock-Out Protocol: Bracketed Session Continuity
- Close the Attack-to-Fix Loop: Adversarially Train Agent
- Closed-Loop Agent Training from Tool Schemas
- Closed-Loop CI Failure Remediation with Cloud Coding Agents
- Closed-Loop Role-Based Refinement for Agent Systems
- Closed-World Tool Call Resolution Before the Permission Gate
- Cloud-Agent Session Bootstrap: Cached Install plus Per-Session Start
- Cloud-Agent Three-Layer State Decoupling
- CoALA Decision-Making Loop as an Orchestration Lens
- CoALA Memory Taxonomy as a Classifier for Harness Artifacts
- CoALA Structured Action Space: Internal vs External Actions
- CoT Robustness in Code Generation
- Code Cleanliness as an Agent Cost Lever
- Code Health as a Signal for Agent-Generated Test Quality
- Code Injection Defense in Multi-Agent Pipelines
- Code Interpreter as a Primary Agent Tool
- Code-Native Memory Substrates for Coding Agents
- CodeSlop: Search-Trajectory Residue in Agent Patches
- Codebase Readiness for Agents: Agent-Friendly Code
- Codebase-Derived Pattern Libraries as Agent Context
- Codified Effort and Escalation Policy in the Instruction File
- Coding Agent Scope Expansion: When to Extend Beyond the Codebase
- Coding-Agent Misalignment Forms (Seven-Symptom Taxonomy)
- Coding-Agent Reversibility: Platform Choice as a Two-Way Door
- Coding-Agent Working-Set Coverage (Coherence Debt)
- Cognitive Architectures for Language Agents (CoALA): A Classifier for Agent Harnesses
- Cognitive Poisoning: Untrusted Tool Feedback as a Trajectory Attack
- Cognitive Reasoning vs Execution: A Two-Layer Agent
- Cohesion-Aware Task Partitioning for Multi-Agent Coding
- Comment Content as Code-Generation Context
- Committee Review Pattern for Multi-Agent Code Review
- Comparative Judging for Agent Configuration Ranking
- Comparison-Only Advisor: Steering a Large Actor With a Tiny Comparator
- Compiled Specialist Agents: Muscle Memory for Recurring Intent
- Completion Failure Taxonomy: Why Code Suggestions Miss
- Completion Summary as the Oversight Surface
- ComplexMCP: Three Bottlenecks in Large Interdependent Tool Sandboxes
- Component-Isolated Memory Stress Testing for LLM Agents
- Component-Wise RAG Prioritization for Software Engineering Tasks
- Compositional Skill Routing for Large Skill Libraries
- Compositional Vulnerability Induction in Coding Agents
- Compound Engineering: Learning Loops That Make Each Feature Easier
- Compound Prompt Constraints Degrade More Than Their Parts Predict
- Computer-Systems Lens for Always-On Agent Security
- Concept Map
- Conceptual Integrity Erosion in Agent-Built Codebases
- Concurrent Agent Pull Requests and Merge-Conflict Cost
- Configuration File Structure Does Not Drive Compliance
- Configuration Smells in AGENTS.md Files (Six-Smell Catalog)
- Confirmation Gates for Consequential Agent Actions
- Consistent-format customer capture
- Consolidate Agent Tools to Reduce Cognitive Overhead
- Constraint Decay in Backend Code Generation
- Constraint Degradation in AI Code Generation
- Constraint Drift: Why Safety Must Be Maintained, Not Asserted
- Constraint Encoding Does Not Fix Constraint Compliance
- Constraint Preambles and the Gain Your Scanner Misses
- Constraint Tax: Tool Suppression Under JSON Schema Decoding
- Constraint-Evasive Fabrication in Instruction Sets
- Constraints as a Substrate for Scalable Agent Oversight
- Containment Playbook: npm-to-Signing-Channel Compromise
- Content-Addressed Agent Configurations (Deterministic Control Plane)
- Context Budget Allocation: Spending Every Token Wisely
- Context Compiler: Deterministic Assembly Over Bigger Windows
- Context Compression Strategies: Offloading and Summarization
- Context Engineering (Training Module)
- Context Engineering: The Practice of Shaping Agent Context
- Context Hub: On-Demand Versioned API Docs for Coding Agents
- Context Lifecycle Management: Beyond Store and Retrieve
- Context Poisoning: When Hallucinations Become Premises
- Context Priming: Pre-Loading Files for AI Agent Tasks
- Context Quality as a Leading Indicator of Agent Reliability
- Context Window Anxiety: Countering Premature Task Closure
- Context Window Management: Understanding the Dumb Zone
- Context-Fractured Decomposition Attacks on Tool-Using Agents
- Context-Graph Shared Memory for Multi-Agent Systems
- Context-Injected Error Recovery for AI Agent Development
- Context-Usage Attribution: Per-Source Breakdown of Agent Context
- Contextual Capability Calibration for Multi-Agent Delegation
- Continual Learning for AI Agents: Three Layers of Knowledge Accumulation
- Continuation Authority in Agent Migration
- Continuation Dispatcher: Who Owns the Iterate-or-Stop Call
- Continuous AI (Agentic CI/CD) for AI Agent Development
- Continuous AI: A Navigation Map of Always-On Agent Workflows
- Continuous Agent Improvement: Iterating on Agent Quality
- Continuous Autonomous Task Loop
- Continuous Documentation as an Agent-Driven Practice
- Continuous Triage: Automating Issue Classification with AI Workflows
- Continuously Built Agent Environments
- Contract-Domain Tracing for Rubric Credit
- Contractual Skill Files: Inspectable SKILL.md for Enterprise Agents
- Control Lexical Leakage in Agent-Memory Retrieval Evals (Entity-Collision)
- Control/Data-Flow Separation for Prompt Injection Defense (CaMeL)
- Controlled Benchmark Rewriting for Agent Safety Judgment
- Controlling Agent Output: Concise Answers, Not Essays
- Convenience Loops and AI-Friendly Code in Your Stack
- Convention Over Configuration for Agent Workflows
- Convergence Detection in Iterative Agent Refinement
- Conversation Registers for AI Coding Sessions
- Coordination Channel Policy for Multi-Agent Coding
- Copilot vs Claude Billing Semantics for Enterprise Teams
- Corpus Shape as a Retrieval Design Constraint
- Corpus-Level Trace Diagnostics for LLM Agents
- Cost-Aware Agent Design: Route by Complexity, Not Habit
- Cost-Aware Skill Rewriting: Preserve Operational Anchors, Not Skill Tokens
- Cost-Aware Tracing for Skill Distillation
- Cost-Driven Model Routing Without Quality Monitoring
- Cost-Inefficient Behaviors in Coding Agents
- Cost-Quality Pareto Measurement for Agent Configurations
- Coverage-Aware Skill Selection Under a Token Budget
- Coverage-Guided Agents for Fuzz Harness Generation
- Coverage-Guided Fuzzing for Multi-Agent LLM Systems (FLARE)
- Covert Success Rate for Indirect Prompt Injection
- Credential Hygiene for Agent Skill Authorship
- Critic Agent Pattern: Dual-Model Plan Review
- Critical Instruction Repetition via Primacy and Recency
- Criticality and Containment: Scoping Which Agent Code You Read
- Cross-Component Interference in Agent Scaffolds
- Cross-Cycle Consensus Relay
- Cross-Framework Signal Semantics: Re-Measure Borrowed Trajectory Rules
- Cross-Functional Knowledge Artifacts
- Cross-Iteration Safety State for Agent Loops (LoopHarness)
- Cross-Layer Evidence for Agent Attack Detection
- Cross-Lingual Prompt Preprocessing (Local-LLM Token Arbitrage)
- Cross-OS Library MicroVM Sandbox for Agent Code (smolvm)
- Cross-Reference Dereference Hop in Retrieval Loops
- Cross-Repo Agent Search: GitHub-API-Backed Text Search Beyond the Workspace
- Cross-Repository Security Posture for Agent-Introduced Vulnerabilities
- Cross-Session Re-Implementation of Existing Agent Code
- Cross-Tool Subagent Comparison
- Cross-Tool Translation: Learning from Multiple AI Assistants
- Cross-Vendor Competitive Routing for LLM Selection
- Cryptographic Governance Audit Trail for AI Agents
- Cue-Anchored Working Memory (Delivery, Not Storage)
- Cumulative-Best Reporting Hides Repair-Loop Security Regressions
- Customer-Hosted MCP Tunnel: Outbound-Only Connectivity to Private MCP Servers
- Cut-Point Replay: Test a Fix Against a Recorded Agent Run
- DSLs as a Constraining Harness for LLM Code Generation
- DSPy: Programmatic Prompt Optimization for Compound Agent Systems
- Daily-Use Skill Library: Encoding Your Process as Agent Skills
- Data Fidelity Guardrails: Preventing Agent Data Mutation
- Debugging the Tool-Call Loop Before Reaching for a Framework
- Decentralized Memory for Self-Evolving Multi-Agent Systems
- Decision Unbundling: What Moves Into the Orchestrator
- Decision-Fork Replay: Grading an Agent's Mid-Run Choices
- Declarative Multi-Agent Composition
- Declarative Multi-Agent Topology: Topology-as-Code
- Declared Peer Consensus as Context for a Reviewing Agent
- Decomposed Red-Teaming for Agent Monitors
- Decomposing Agent Output Variability by Layer (Sampling vs Orchestration State)
- Decoupled Search Grounding: A Vendor-Agnostic Grounding Boundary
- Deep Agent Runtime: The Layer Beneath the Harness
- Defense-in-Depth Against Coding Agent Fabrication (Honesty Harness)
- Defense-in-Depth Agent Safety for AI Agent Development
- Deferred Standards Enforcement via Review Agents
- Delegated-Autonomy Boundary Artifacts (AJR and ADP)
- Delegating Change Descriptions to the Agent
- Delegating Multi-Hunk Bug Repair to Coding Agents
- Delegation Threshold Calibration for Orchestrator Agents
- Deletion Avoidance: Agents That Guard Code Instead of Removing It
- Deletion and Cost Rules for an Evolving Agent Harness
- Deliberate AI-Assisted Learning: Accelerating Skill Acquisition
- Deliberation-Inducing Cues That Multiply Reasoning Cost
- Delivery-Bound Tool Authorization: When Progressive Discovery Becomes Access Control
- Delta Channels: Bounded Checkpoint Storage for Append-Only Agent State
- Demand-Driven Repository Auditing
- Demo-to-Production Gap: When Demos Hide Real Costs
- Density-Normalized Quality Metrics Mask AI-Driven Code Growth
- Dependency Gap Validation for AI-Generated Code
- Dependency Inlining Erodes SBOM and License Provenance
- Deriving a Specification From Buggy Code Before Generating Tests
- Design Docs as the Durable Artifact
- Designing Agent Plugins to Survive Co-Installation
- Designing Agent Tools Like APIs
- Designing Agents to Resist Prompt Injection
- Designing for Agent Consumers (Agent Experience)
- Destructive-Failure Mechanism Attribution by Mitigation Owner (ClayBuddy Three)
- Destyling Untrusted Input as a Prompt Injection Defense
- Detecting Memory-Poisoning Exfiltration by Tool-Call Order (Recall-Before-Send Signature)
- Detecting Self-Preference in a Single LLM Judge
- Deterministic Anchoring: Static Facts as Stable Context
- Deterministic Fast Paths: Answer Without a Model Call
- Deterministic Guardrails Around Probabilistic Agents
- Deterministic Orchestration for Structured Modernization
- Deterministic Precondition Gates for Tool-Using Agents
- Dev Containers for AI Coding Agents: Claude Code vs Copilot CLI
- Developer Control Strategies for AI Coding Agents
- Developer as CPU Scheduler: Attention Management with Parallel Agents
- Diagram as the Shared Spec: One Artifact for the Picture and the Prompt
- Diff-Based Review Over Output Review
- Diff-Coverage Gating for Agent-Authored Pull Requests
- Difficulty-Aware Topology Selection for Coding Agents
- Direct Prompt Injection via Collaboration (User as Attack Vector)
- Discoverable vs Non-Discoverable Context for Agents
- Discovering Indirect Injection Vulnerabilities in Your Agent
- Discovery-Only Refactor Pass: Surface Candidates Before Touching Code
- Discrete Phase Separation
- Distillation-Induced Similarity Metrics for Tool-Use Agents
- Distilled Bootstrap Contract: Agent-Authored Repo Setup
- Distractor Interference: Why Relevance Is Not Enough
- Distributed Computing Parallels in Agent Architecture
- Distributed Cross-PR Attacks in Persistent-State AI Control
- Distributing Security Controls Through the Agent Harness
- Do Not Price the Rules in Your Agent Instruction File
- Docker sbx Adoption for Coding Agents
- Document-Borne Prompt Injection Through Agent Read Tools
- Documentation Read Counts Measure Retrievability, Not Value
- Documentation-Grounding MCP Servers for Vendor SDKs
- Documentation-Guided Legacy Migration: Architecture Docs as a C-to-Rust Blueprint
- Documenting Code the Agent Can Already Read
- Domain-Scoped Parallel Exploration for Multi-File Change Localization
- Domain-Specific Agent Challenges
- Domain-Specific System Prompts with Concrete Examples
- Dominator-Graph Trajectory Invariants for Non-Deterministic Agents
- Dormant Memory Payloads Triggered by Sensitive Topics (Trojan Hippo)
- Downstream Disclosure Coordination for Agent-Found Defects
- Dual Executable Specifications for Long-Horizon Features
- Dual-Boundary Sandboxing: Filesystem and Network Isolation
- Dual-Budget Control for Search Agents: VOI Scoring Per Action
- Dual-Graph Alignment for Indirect Prompt Injection Defense (AuthGraph)
- Dual-Trace Memory Encoding: Pair Facts with the Scene They Were Learned In
- Dual-Write Append-Mirror for Agent Transcript Externalization
- Durable Interactive Artifacts: Agent Output Outside the Transcript
- Dynamic System Prompt Composition
- Dynamic Tool Fetching Destroys KV Cache Performance
- Earned-Complexity Agent Maturity Ladder
- Economic Value Signaling in Multi-Agent Networks
- Ecosystem-Level Integration Friction Governance
- Edit Format Selection: Diff vs. Search-Replace vs. Full Rewrite
- Editor and Manager Surface Separation in Agent IDEs
- Effective Feedback Compute (EFC) for Harness Comparison
- Elastic Context Orchestration: A Per-Turn Vocabulary for Long-Horizon Search Agents
- Embedding Inversion: Vector Stores as a Source-Text Disclosure Surface
- Emergent Architecture in AI-Driven Codebases
- Emergent Behavior Sensitivity for AI Agent Development
- Empirical Baseline: Agentic AI Coding Tool Configuration
- Empowerment Over Automation for AI Agent Development
- Emulate Agent-Experience Changes Before Shipping
- Emulated APIs for Agent Skill Evals
- Encode Project Conventions in Distributed AGENTS.md Files
- Encoding AI Writing Tells as a Prose Style Contract
- Encoding Product-Design Taste into Agent Context
- Encoding Tacit Knowledge into Agent Improvement Loops
- Encoding Values in AGENTS.md: Why Prose Without Verification Fails
- Enforced Versus Advisory Controls in LLM-Native IDEs
- Enforcement Modes in Spec-First Agent Frameworks
- Enforcing Who and What Can Trigger an Agent's CI Run
- Engineering: Tools, Review, Verification, Security, and Observability
- Enterprise Agent Hardening: Three Production Gates
- Enterprise Skill Marketplace: Distribution and Quality
- Entity Binding Failures in Tool-Augmented Agents
- Entropy Reduction Agents: Automated Codebase Hygiene
- Enumerate Every Visible Exit Before Scoring Containment
- Environment Specification as Context: Closing the Version Gap
- Episodic Memory Retrieval for AI Coding Agent Loops
- Epistemic Working Memory for Multi-Hop Reasoning (SLEUTH)
- Equivalence Testing for Agent Configuration Changes
- Error Preservation in Context for AI Agent Development
- Escalation Channels: A Reporting Tool Instead of a Reward Hack
- Escape Hatches: Unsticking Stuck Agents
- Eval Awareness: Designing Evals Agents Cannot Recognize
- Eval Blind Spots: Structural Gaps in Measurement Methodology
- Eval Difficulty as a Product Smell
- Eval Engineering (Training Module)
- Eval Environment Containment for Cyber-Capable Agents
- Eval Strategy by Agent Generation: A Structure-to-Eval Locator
- Eval-Driven Development Training for AI Agent Teams
- Eval-Driven Development: Write Evals Before Building Agent
- Evaluating AGENTS.md: When Context Files Hurt More Than Help
- Evaluating Agent Patterns Catalog as a Source
- Evaluator Templates: Portable Primitives for Agent Eval Suites
- Evaluator-Optimizer Pattern for AI Agent Development
- Event Sourcing for Agents: Separating Cognitive Intention
- Event-Driven Agent Routing for Multi-Team AI Pipelines
- Event-Driven System Reminders for AI Agent Development
- Event-Loop Contention in Async Agent Fan-Out
- Evidence-Bundled Agent PRs: Sizing the Reviewer's Effort
- Evidence-Chain Run Logs: Bracket the Reported Symptom
- Evidence-Conditioned Execution: Gate Edits on Observations
- Evidence-First Reports From Failure-Diagnosis Agents
- Evidence-Gated Lifecycle Control for Coding Agents (Proof-or-Stop)
- Evidence-Grounded Disagreement in Agentic Code Review (Adversarial Review)
- Evolving Playbooks: Incremental Context That Preserves Knowledge
- Exactly-Once Enforcement Layer: Model, Harness or Contract
- Example-Driven vs Rule-Driven Instructions
- Exception Handling and Recovery Patterns for AI Coding Agents
- Executable Memory: User State as Code for Personalized Agents
- Execution Budgeting in Agentic Program Repair
- Execution Lineage: DAG of Artifacts vs Agent Loops
- Execution-First Delegation: The AI-as-Executor Pattern
- Execution-Layer Security Invariants for MCP Runtimes
- Execution-State Ledger for Long-Horizon Coding Agents
- Exhaustive Retrieval for Listing Questions
- Experience Graphs as Structured Memory for Self-Evolving Agents
- Experiential-Learning Setup Agents with Snapshot Rollback (SetupX)
- Explained Feedback for LLM Vulnerability Repair
- Explanation-Bound Tool Execution for Agent Gateways
- External Artifacts Treated as Data, Not Adversarial Input
- Externalization in LLM Agents
- Fact Supersession Memory for Code Assistants
- Factory Over Assistant: Orchestrating Parallel Agent Fleets
- Failure-Aware Observability for Multi-Agent LLM Systems
- Failure-Driven Iteration for Improving Agent Workflows
- False-Pass Liability Decides the Next Harness Component
- Fan-Out Synthesis Pattern for AI Agent Development
- Feature List Files for Reliable AI Agent Development
- Feedback as Capability Equalizer: Iterative Feedback Outweighs Model Scale
- Field-Change Intent Instead of Model-Written Diffs
- Field-Level Source Ownership for Agent Capabilities
- File-Based Agent Coordination for AI Agent Development
- Filter and Aggregate Data in the Execution Environment
- First-Party Agent Composition: Agent-Built Features
- First-Proposal Execution in Agent Loops
- Five Design Decisions for MCP Servers and Clients
- Five-Failure-Layers Diagnostic: Attribute Before Swapping the Model
- Five-Pass Blunder Hunt: Repeated Critique Passes for Plans
- Five-Stage Policy Layer Typology for Generalist Agents
- Flattened Tool Specs for Agent Safety Judgment (SafeKeep)
- Fleet Harness Attribution: Pinning the Model to Compare Whole Harnesses
- Fleet-Level Irreversibility Budgets for Agent Effects
- Foresight-Guided Defense Against Infectious Jailbreaks in Multi-Agent Systems
- Forged Reasoning Trace Attacks on Agent Memory (FARMA)
- Forked vs Fresh Subagents: When to Inherit the Parent Conversation
- Formal Process Models as Prompting Scaffolds (Petri Net of Thoughts)
- Foundational Disciplines for AI-Assisted Development
- Foundations: Context Engineering and Instructions
- Four Reporting Levels for Agent Working Memory Evaluation
- Four-Layer Taxonomy of Agent Security Risks
- Four-Phase Agent Delegation with Curated Artifacts
- Framework-First Agent Development: An AI Anti-Pattern
- Frameworks
- Framing Subagent Returns So They Cannot Act as Instructions
- From Preventive to Reactive: Front-Loading Security in AI Coding Prompts
- Frontmatter and Body Rule Drift in Agentic Workflows
- Frozen Playbook Reuse Without Target-Side Validation
- Frozen Spec File: Preserving Intent in AI Agent Sessions
- Frozen Task Sets for Affordable Agent A/B Testing
- Frozen-Base Task Mining for Repository Instruction Files
- Frozen-Stimulus Panels for Cross-Vendor Behavior Measurement
- Function-Level Debugger Interfaces for Coding Agents
- Functional folder taxonomy
- Future-Based Asynchronous Function Calling
- GEO for Technical Docs: Developer Documentation Checklist
- GROUNDING.md: Field-Scoped Hard Constraints and Convention Parameters
- Gate Agent Writes to Executable Config Files as Privileged Actions
- Gate Best-of-k Selection on Compliance Before Score
- Gate Generation on Retrieval Sufficiency, Not Model Confidence
- Generated Procedure Drivers: Skills That Emit a Program
- Generated Programs as Web Agent Action Space
- Generating Tests From Agent-Written Code (Code-First Oracle Bias)
- Generative Agents Memory Stream: Three-Layer Architecture for Long-Running Agent Sessions
- Generative Engine Optimization for Developer Sites
- Generative Provenance Records for Tool-Using Agents
- Getting Started: Setting Up Your Instruction File
- Git-Bound Memory for the Agentic Development Lifecycle
- Give the Model the Target's Contract, Not Similar Solutions
- Goal Contract: Separating the Doer from the Done-Checker
- Goal Monitoring and Progress Tracking for Long-Running Agents
- Goal Recitation: Countering Drift in Long Sessions
- Goal Reframing: The Primary Exploitation Trigger for LLM Agents
- Goal-Driven Autonomous Loop with Budget Cap
- Golden Journeys: Restartability as a First-Class Verification Primitive
- Golden Query Pairs as Continuous Regression Tests for Agents
- Google ADK Skills: Portable SKILL.md Across ADK Agents
- Google Search Console Monitoring Workflow
- Governance Layer for Agent Interoperability Protocols
- Governed Sources of Truth for Analytics Agents (Structure Over Access)
- Graceful Tool-Output Truncation: The PARTIAL Signal
- Grade Agent Outcomes, Not Execution Paths
- Grading Strategies for Eval-Driven Development
- Graph of Thoughts: Directed Graph Reasoning for Multi-Path Problems
- Grill Me: Developer-Initiated Plan Interrogation
- Grounding Agents in Code the Model Has Never Seen
- Guarding Against URL-Based Data Exfiltration in Agentic Workflows
- Guardrails Beat Guidance: Rule Design for Coding Agents
- HTML as Agent Output Format: When to Ask for HTML Instead of Markdown
- Handoff-Boundary Fault Injection (llmmas-otel)
- Happy Path Bias: How AI Agents Skip Error Handling
- Hardening Agent Evals for Production-Grade Reliability
- Harness Bug Detection Patterns
- Harness Composition for Scaled Security Audits
- Harness Design Dimensions and Archetypes
- Harness Engineering (Training Module)
- Harness Engineering for Building Reliable AI Agents
- Harness Hill-Climbing: Eval-Driven Iterative Improvement of Agent Harnesses
- Harness Impermanence: Build Scaffolding To Be Deleted
- Harness Preflight Doctor Command for Agent Diagnostics
- Harness-Controlled Token Economics (The Harness Effect)
- Harness-Memory Coupling as a Design Axis
- Head-to-Head Evaluation of Competing MCP Servers
- Headless-First Services: APIs for Agent Consumers
- Heartbeat-Bound Hierarchical Credentials for Agent Swarms
- Held-Out Tasks as a Harness Shortcut Defense
- Heuristic-Based Effort Scaling in Agent System Prompts
- Hint-Driven Concurrency for Read-Only MCP Tools
- Hints Over Code Samples in Agent Prompts
- History Anchors: Consistency-Cued Continuation of Unsafe Prior Actions
- Homogeneous Debate Panels as a Groundedness Quality Lever
- Hooks and Lifecycle Events: Intercepting Agent Behavior
- Hooks for Enforcement vs Prompts for Guidance: When to Use Each
- Hostname-Allowlist Proxy: The TLS-Inspection Blind Spot
- How AI Engines Cite — ChatGPT, Perplexity, Claude, Gemini
- How Teams Build SE Agents: A Seven-Stage Build Loop
- How the Four Agent Engineering Disciplines Compound
- Human-AI Review Synergy in Agentic Code Review
- Human-Equivalent Hours for Autonomous Coding Agent Productivity
- Human-Facing Docs in the Agent Era: Mental Models Over Reference
- Human-Review-Driven Curation of Golden Eval Datasets
- Human-in-the-Loop Checkpoints as Loop Control
- Human-in-the-Loop Placement: Where and How to Supervise
- Humans and Agents in Software Engineering Loops
- Hybrid Deterministic + Semantic Authorization for Agent Tool Calls
- Hyper-Personalized Software: The Return of RAD
- Hypothesis-Driven Debugging: Instrument Before You Patch
- Hypothetical Classification for Large Label Vocabularies
- Idempotent Agent Operations: Safe to Retry
- Idle-Time Speculative Planning for ReAct Agents
- Improper Output Handling: Validate Agent Output Before Downstream Use
- In-Agent Task Prioritization: Ranking the Next Action
- In-Loop Interception: Custom Logic Between the Model Call and the Tool Call
- In-Place Atomic Replacement for Agent-Driven Ports
- In-Process WebAssembly Sandboxes for Agent-Generated Code
- In-Thread Side-Channel: Bounded Side Questions Without Losing the Main Task
- Incident Log Investigation Skill: Parallel Queries
- Incident-to-Eval Synthesis: Production Failures as Evals
- Incremental Verification: Check at Each Step, Not at the End
- Independent Test Generation in Multi-Agent Code Systems
- Indexed Regex Search for Agent Tools
- Indiscriminate Structured Reasoning on Every Agent Task
- Inference-Time Tool-Call Reviewer: Pre-Execution Feedback for Tool-Calling Agents
- Inferring Agent Failure from Conversation Evidence (Perceived Error)
- Informed Abstention as a Tool-Boundary Runtime Gate
- Initiatives and Community: Tracking the Agentic Engineering Landscape
- Injected-Context Cost Attribution in Agent Workflows
- Inline Safety Harness with Cascade Verification (FinHarness)
- Inline Suggestion Attachment in Agent Code Review
- Install-Once Plugin Trust: Vetting That Never Re-Runs on Update
- Instruction Polarity: Positive Rules Over Negative
- Instruction-Aware Automated Code Review
- Instruction-Guided Code Completion: Controlling What Models Generate
- Intent-Centric Engineering: Oversight Over Authorship
- Intent-Governed Tool Authorization for AI Agents (IGAC)
- Interaction-Pattern Evaluation for Agentic PRs
- Interactive Canvases: Agent-Generated Visual Artifacts as Outputs
- Interactive Clarification for Underspecified Tasks
- Interactive Effort Sliders: Per-Turn Reasoning-Budget Controls
- Internal Hostname Disclosure in Agent-Readable Context
- Intervention Rate as a Diagnostic North Star, Not a Target
- Introspective Skill Generation: Mining Agent Patterns
- Inversion Analysis: Surface Capabilities Competitors Cannot Replicate
- Isometric Harness Ablation: Rank Subsystem Investment by Removing One at a Time
- Issue Requirements Preprocessing: Structured Input Before Code Generation
- Issue-Tracker as Agent Dispatch Surface
- Issue-to-PR Delegation Pipeline for AI Agent Development
- Iterative Binary Feedback for Pattern Adherence
- Judging Agent Safety by Task Completion (Action-Boundary Violations)
- Judging MCP Capability Readiness by Client Adoption
- Judging a Skill's Honesty by the Validity of Its Output
- Judgment Relocation: Where Human Decisions Land in an Agent Factory
- Kaizen-Style Continuous Code Quality Loop (Pomona)
- Knowledge Cutoff as a Documentation Boundary
- Knowledge Gap or Skill Gap: Triage Before Writing Context
- Knowledge Graphs as Provenance-Carrying Agent Memory
- Knowledge-Based Pull Requests for Cross-Trust-Boundary Contributions
- L0 → L1: Making the Repo Readable
- L1 → L2: Adding Feedback Loops
- L2 → L3: Building Mechanical Enforcement
- L3 → L5: Reaching Agent-First
- LLM API Fault Injection at the HTTP Layer (AgentChaos)
- LLM API Routers as Application-Layer Man-in-the-Middle
- LLM Agent Bug Fix Taxonomy: 23 Fix Patterns from 930 Real Bugs
- LLM Code Review Overcorrection for AI Agent Development
- LLM Comprehension Fallacy: When Models Seem to Understand
- LLM Map-Reduce Pattern for Parallel Input Processing
- LLM Refactoring Adoption Patterns
- LLM Self-Review Failure in Code Modernization Tasks
- LLM Static Verification Against Natural-Language Requirements
- LLM Support During the First Detection Pass
- LLM-Driven Benchmark Auditing
- LLM-Driven Logical Retrieval: Boolean Queries over an Inverted Index
- LLM-Pinned Library Versions Carry Systemic CVE Exposure
- LLM-as-Code Agentic Programming for Agent Harnesses
- LLM-as-Judge Evaluation with Human Spot-Checking
- Labels as Locks: Pipelined Backlog Processing with Stage Gates
- Lane-Based Execution Queueing
- Language Choice as an Agent Token-Cost Lever
- Language Selection Scored on Review Cost
- Large-Codebase Coding-Agent Failure Patterns (Sourcegraph Five)
- Late Requirement Arrival in Agent Sessions
- Law of Triviality in AI PRs for AI Agent Development
- Lay the Architectural Foundation by Hand Before Delegating
- Layer Agent Instructions by Specificity Across Scopes
- Layered Accuracy Defense for Reliable Agent Outputs
- Layered Context Architecture for AI Agent Development
- Layered Domain Architecture: A Prescriptive Default for Agent-Built Code
- Layered Mutability: Governing Persistent Self-Modifying Agents
- Layered Oracle Stack for Agent IaC Security Repair (TerraProbe)
- Lead-to-Teammate Plan-Approval Handshake for Multi-Agent Work
- Learned Prefix Monitors for Agent Traces
- Learning Execution Guardrails from Agent Failure Traces
- Legacy Code Archaeology: Reconstruct Intent Before Migrating
- Lethal Trifecta Threat Model for AI Agent Development
- Lexical-First Retrieval for Agentic Search: When BM25 Is Enough
- Lifecycle-Integrated Security Architecture for Agent Harnesses
- Line-Anchored Feedback: Deliver Change Requests as Inline Comments
- Listener-State Naming for User-Invoked Agent Skills
- Live Browser as Agent Context Channel
- Living-Docs-Grounded Agent Design Conversations
- Local Model Viability Factors for Coding
- Lock-State Safeguards for Desktop-Controlling Agents
- Long Context vs Retrieval: The Break-Even Decision
- Long-Running Agents: Durability and Resumability Across Sessions
- Loop Budgeting: Allocating Iteration and Token Budget Across Turns
- Loop Detection for AI Agents: Stopping Micro-Loops
- Loop Strategy Spectrum: Accumulated vs Fresh Context
- Loop Trigger Selection: Pairing the Start with the Stop
- Lost in the Middle: The U-Shaped Attention Curve
- MCP Allowlist by Label, Not by Identity (serverName Trap)
- MCP Approval-View Fidelity Gap and Unicode Concealment
- MCP Client Design: Building Robust Host-Side Logic
- MCP Runtime Control Plane: Policy Evaluation Between Agent and Tool
- MCP Server Design: Building Agent-Friendly Servers
- MCP alwaysLoad: Classifying Servers as Eager or Just-in-Time
- MCP-vs-CLI Cost Ratios Are a Property of the Scaffolding
- MCP: The Open Protocol Connecting Agents to External Tools
- Machine-Readable Error Responses for AI Agents (RFC 9457)
- Macro Evals for Agentic Systems: Population-Level Behavior Patterns
- Magentic Orchestration: Task-Ledger-Driven Adaptive Multi-Agent Planning
- Making Application Observability Legible to Agents
- Managed vs Self-Hosted Agent Harness: Deployment Trade-offs
- Managing Cognitive Load and AI Fatigue for Sustainable Agent Use
- Manual Compaction Strategy for Dumb Zone Mitigation
- Marking Which Artifacts Are for Humans or Agents (Landmarking)
- Markov-Chain Reliability for LLM Agents: Audit the Abstraction Before You Trust the Metric
- Mask Tools Instead of Removing Them
- Match Architecture Spec Format to Model Capability
- Match Tool Instructions to the Agent Workflow
- Measure the Judge Before You Freeze a Gate on It
- Measuring GEO Performance for AI Search Visibility
- Measuring Reacquisition Cost Under Context Compaction
- Measuring Refactoring Payback in Tokens
- Measuring Synthetic Eval Data Quality (SynAE)
- Measuring the Verification Tax on Agent Output
- Memory Retrieval as a Control Decision
- Memory Synthesis: Extracting Lessons from Execution Logs
- Memory Transfer Learning: Cross-Domain Memory Reuse in Coding Agents
- Memory-Induced Tool-Drift in LLM Agents
- Mermaid as Agent Output Format: When to Ask for a Diagram Instead of Prose
- Meta-Evaluate the LLM Judge Before Trusting Rubric Verdicts
- Method Map: Failure-Mode to Smallest-Artifact Triage
- Mid-Session Config Changes as Invisible Cache Invalidators
- Mid-Trajectory Guardrail Selection for Multi-Step Tool Calls
- Minimality Prompts as a Patch-Size Control
- Minimum-Cost Evidence Selection for Agent Changes (Assurance Envelopes)
- Minimum-Sufficient Control Ladder: Escalate by Failure Mode
- Minimum-Sufficient Execution: Estimate Scope Before Spending Budget
- Mise en Place for Agentic Coding
- Model Confidence as Security Verification (Security Calibration Gap)
- Model Deprecation Lifecycle for Agent Workloads
- Model a Single Agent Turn as Many Inference and Tool-Call
- Model-Directed Subagent Tiering: Lead Model Picks the Tier
- Model-ID-as-Dependency: Migration Protocol for Deprecation Churn
- Model-Neutral Agent Architecture: Model Portability Over Cloud Portability
- Model-Set Parity: Reading Harness Efficiency Claims
- Monitor or Wait: The Supervision Choice During Agent Execution
- Monolith-to-Sub-Agents Refactor: Five Lessons from a Brittle Prototype
- Monotonic Capability Attenuation for Composition-Safe Tool Use
- Most-Restrictive-Wins Fusion for Parallel Agent Control Returns
- Multi-Agent RAG for Spec-to-Test Automation
- Multi-Agent SE Design Patterns: A Taxonomy Across 94 Papers
- Multi-Agent Shared State Isolation Anomalies
- Multi-Agent Topology Taxonomy: Centralized, Decentralized
- Multi-Client Session Attachment: One Session, Many Clients
- Multi-Layer Specification Redundancy as a Robustness Budget
- Multi-Model Plan Synthesis for System Architecture
- Multi-Run, Shuffled-Order Evaluation for Self-Improving Agents
- Multi-Shape BYOK Provider: Declare API Family per Endpoint
- Multi-Tool Threshold Poisoning Against MCP (ShareLock)
- Multi-Turn Conversation Evaluation: Per-Turn and Trace-Level Scoring Together
- Multitenant RAG: Closing the Relevance-Authorization Gap
- Mutation Testing as a Quality Gate for AI-Generated Test Suites
- Mutation Testing for LLM Judges: Scoring an Evaluator on Injected Defects
- Name the Check That Passed Before Accepting AI Code
- Narrative Problem Reformulation for Code Generation
- Natural Language Tool Selection (NLT)
- Natural-Language Customization Bootstrap
- Natural-Language Documentation as a Code-Review Intermediate (Verifiable Literate Programming)
- Natural-language git
- Negative Space Instructions: What NOT to Do in Agent Prompts
- Network-less Container + Unix-Socket Egress Proxy for Agent Sandboxes
- Next Edit Suggestions Carry Context You Never Curated
- Non-Compensatory Readiness Gates Before Agent Release
- Non-Human Event Provenance Markers to Block Fabricated Approvals
- Nonstandard Errors in AI Agents: Model-Family Variance
- Notebook-Documented Automation for Repeat Operational Work
- OAuth Client ID Metadata Documents (CIMD) for MCP Servers
- OWASP 2026 Update for Agent Builders: Top 10 Renumbering and the Agent Control Standard
- OWASP LLM Top 10 (2025): Agent Security Crosswalk
- Objective Drift: When Agents Lose Sight of the Goal
- Observability Feedback Loop: A 7-Step Debug Runbook for Agents
- Observability-Driven Harness Evolution
- Observation Contract Preservation in Tool-Augmented Agents
- Observation Masking: Filter Tool Outputs from Context
- Observation-Driven Coordination: CRDT-Based Parallel Agent
- Offline Evaluation as an Integration Test for LLM Features
- Offline Trajectory Replay for Multi-Agent Workflow Debugging
- One-Shot Record and Deterministic Replay for Periodic Agent Tasks
- Open Agent School Pattern Mapping for Practitioners
- Open Standards and Protocols for AI Agent Development
- OpenAI Agents SDK
- OpenAPI Documentation Smells for Agent-Ready APIs
- OpenAPI as the Source of Truth for Agent Tool Definitions
- OpenTelemetry for AI Agent Observability and Tracing
- Opponent Processor / Multi-Agent Debate Pattern
- Oracle Poisoning: Knowledge Graph Corruption Against Tool-Using Agents
- Oracle-Based Task Decomposition for AI Agent Development
- Oracle-Gated Delegation Beyond Your Domain Expertise
- Orchestrator-Worker Pattern for AI Agent Development
- Organizational Context Layer for Agents (Company Brain)
- Organizing Filesystem Agent Memory for Retrieval Cost
- Outcome Monitors: Recovery Affordances for Tool Failures
- Outcome Pricing as a Scope Signal
- Over-Orchestrated Agent Architecture (Prefer the Simplest That Works)
- Overeager-Behavior Elicitation: Scope + Trap Fragments as a Diagnostic for Out-of-Scope Tool Calls
- Override Pattern: Reusing Interactive Commands in Automated Pipelines
- Overtrusting Human Sign-Off on Generated Assertions
- PASS@(k,T): Evaluate RL for Agents Along Sampling and Interaction Depth
- PEEK: Orientation Cache for Recurring-Context Agents
- PM on the AI Exponential
- PR Description Style as a Lever for Agent PR Merge Rates
- PR Scope Creep as a Human Review Bottleneck
- Parallel Agent Sessions Shift the Bottleneck from Writing
- Parallel Polyglot Ports as a Spec-Ambiguity Oracle
- Parameter-Keyed Caching and Dependency-Aware Parallelism for Plan-Execute Pipelines
- Parser-Versus-Shell Evasion in Command Permission Checks
- Parsimonious Agent Routing for Multi-Agent Dispatch
- Path-Scoped Write Contracts for Shared Agent State
- Pattern Replication Risk in Agentic Code Generation
- Pattern Selection Map
- Peer Refusal as a Coordination Control
- Per-Agent Capability Stores Beat One Task-Wide Allowlist
- Per-Attempt Sandboxes for Agents That Change the Filesystem
- Per-Call Budget Hints on Tool Invocations
- Per-Caller Identity: Who an Agent's Tool Call Acts As
- Per-Layer Suppression Accounting in Acceptance Gates
- Per-Line Requirement Citations for Hallucination Detection
- Per-Model Harness Tuning: Treating the Backing Model as a Harness Variable
- Per-Object Context Allocation (Selective Invariance)
- Per-Reviewer Context Views for Code Review Agents
- Per-Run Budget Reservation for Coding Agent Model Calls
- Per-Server MCP Environment Scoping for Credential Isolation
- Per-Step Preconditions and Postconditions in Skill Files
- Per-Task Agent Routing Across Coding Harnesses
- Per-Task Verification Budget: Size the Task to Fit the Check
- Per-Tool Extended Reasoning Opt-In: Tool-Call-Scoped Budgets
- Per-Type Retention Policy for Agent Compaction (Knowledge Triage)
- Per-User Supervisor Process for Background Agent Sessions
- Perceived Model Degradation: Why Vibes Are Not Evals
- Permission Framework Choice Outweighs Model Choice for Limiting Overeager Actions
- Permission Modes as a Defense Against a Tampered Response Path (Response-Path Control Gap)
- Permission-Gated Custom Commands for AI Agent Development
- Permitted Egress Routes as Agent Sandbox Attack Surface
- Permutation Frameworks for Batch Code Generation
- Persistent Shared Search Sub-Agent for Output-Token Reuse
- Persistent-Connection Agent Transport
- Persona-as-Code: Defining Agent Roles as Structured Docs
- Personalized vs Generic Agent Skills: Where Effort Pays
- Phantom Symbol Detection for LLM API Migration
- Phase-Specific Context Assembly for AI Agent Development
- Plan Compliance in Agents: Measure What They Execute, Not What You Wrote
- Plan files as resumable artifacts
- Plan-Then-Execute as the Default for Web Agents
- Planning Stage Preconditions: Budget Headroom and Task Text
- Planted-Bug Methodology: Deliberate Bugs as Observability Calibration
- Plugin and Extension Packaging: Distributing Agent Capabilities
- Poka-Yoke for Agent Tools: Mistake-Proof Tool Interfaces
- Policy File Validation: Catching Silent Non-Enforcement
- Policy-Graded Evaluation of Coding Agents
- Polya Small-Steps: Using AI to Think Better, Not Think Less
- Pooled-Evidence Factuality Checks for MCP Agents (Cross-Source Conflation)
- Portable Agent Definitions: Full-Stack Identity as Code
- Post-Merge Fix Signals for Agent Merges
- Pre-Change Impact Analysis: Dependency Maps That Prevent Agent Regressions
- Pre-Completion Checklists for AI Agent Development
- Pre-Execution Codebase Exploration for AI Coding Agents
- Pre-Execution Failure Scoring with a Draft Model (Speculative Uncertainty)
- Pre-Generation Complexity Scoring for Code Reliability
- Pre-Trust Execution Surface in Coding Agent Harnesses
- Pre-Write Change Intent Admission (Claim Plane)
- Prebuilt Agent Monitoring Dashboard
- Precise Debugging: Measure Edit Precision, Not Just Test Pass Rate
- Predicting Reviewable Code: Pre-Flagging Functions Reviewers Will Delete
- Preempting Agentic PR Rejection by Failure-Mode Category
- Premature Completion: Agents That Declare Success Too Early
- Prescribing TDD Inside the Agent Loop (Process Theater)
- Pressuring a Coding Agent Degrades the Code It Writes
- Pricier-Per-Token Models That Cost Less Per Task
- Prior Dominance Over Feedback in Agent Optimization Loops
- Privacy-Preserving LLM Requests: Eight Techniques and a Practical Combination
- Proactive Idle-Time Anticipation (ProAct)
- Probe-Run Calibration for Predicting Agent Token Spend
- Probe-and-Refine Tuning of Repository Guidance for Coding Agents
- Probing Unstated Constraints in Generated Code (Intent Violation Rate)
- Process Amplification: Scaling Human Work with Agents
- Product-Operation Tool Surfaces for In-Product Agents
- Product-as-IDE: When the Application Becomes the Development
- Production MCP Agent Stack: Sequencing Six Decisions into One Deployment
- Profile Your Agent Test Suite Against Measured Practice
- Profiler-Guided Optimization Loops for Coding Agents
- Programmatic Cloud-Agent Dispatch via REST API and Webhooks
- Programming Language Choice Still Shapes Agent Artifacts
- Progressive Autonomy: Scaling Trust with Model Evolution
- Progressive Disclosure for Layered Agent Definitions
- Progressive Spend Threshold Alerting for Agent Cost Governance
- Project Instruction File Ecosystem
- Project Writing Skill: House Style as Model-Invocable Skill
- Prompt Cache Keepalive for Agent Pauses
- Prompt Caching: Architectural Discipline for Agents
- Prompt Chaining: Sequential LLM Calls for Agent Workflows
- Prompt Compression: Maximizing Signal Per Token
- Prompt Debt: Hand-Tuning Natural-Language Prompts as Technical Debt
- Prompt Engineering for Agent Instructions and Systems
- Prompt File Libraries for Reusable Agent Instructions
- Prompt Governance via PRs: Reviewable AI Behavior
- Prompt Injection: A First-Class Threat to Agentic Systems
- Prompt Layering: How Instructions Stack and Override
- Prompt Transpilation: Instructions as Build Artifacts
- Prompt as Security Knob
- Prompt-Only Baseline Before a Specialized Agent Subsystem
- Prompt-Only Tool Access Control
- Prompt-Rewrite Discipline on Cross-Generation Model Migration
- Prompted Uncertainty Decomposition for Clarification Routing
- Proof of Presence: Re-Authenticating for Agent Actions
- Proprioceptive Context Dashboard: Agent Self-Managed Context
- Protecting Sensitive Files from Agent Context Access
- Protecting the Test Oracle From the Agent
- Prototype Before Optimizing: Establish Quality Baselines Before Token Constraints
- Provenance-Aware Decision Auditing for LLM Agents
- Provider-Hosted Subagent Delegation: One Model Price for the Whole Tree
- Public Placeholder Credentials with a Fail-Closed Injecting Proxy
- Public-Channel Agent Work as Lehrwerkstatt for Team Learning
- Publishing Agent Instructions Outside the Repo (design.md)
- Purpose-Built Eval Suites for Model and Harness Swaps
- Push-Event MCP Channels: Inverting the Pull-Tool Polarity
- QA Session to Issues Pipeline for AI Agent Development
- Quality Score Rubric and Simplification Log for Agent Harnesses
- Query-Conditioned Reuse of Retrieved Agent Trajectories
- RAG Architecture as a Poisoning Robustness Decision
- RAG over Thinking Traces: Index Reasoning Trajectories Instead of Documents
- RAG/Agent Reliability Problem Map: 16-Domain Failure Taxonomy
- RAMP: Committed AI Configuration and the Quality Cost
- RL-Trained Automated Red Teamers for Prompt Injection Discovery
- Rainbow Deployments for Agents: Gradual Version Migration
- Rank Resolution: Reading a Converged Coding-Agent Leaderboard
- Re-Run an Agent's Speed-Up Claim Before Merging
- Re-Run the Original Test Suite After Every Refinement Turn
- ReAct (Reason + Act): Interleaved Reasoning-Action Loops
- Reader-Scoped Trace Views: When to Build Your Own UI
- Reading Visible Edge-Case Handling as a Security Check
- Reading a Coding-Agent Vendor's Security Certificate
- Reasoning Budget Allocation: The Reasoning Sandwich
- Reasoning Effort Over Tool Scaffolding for First-Try Reliability
- Reasoning Retention and Compaction as Harness Settings
- Recover the Six Measurement Choices Behind an Attack Success Rate
- Recoverability-Gated Self-Evolution for Agent Harnesses
- Recurring Control Belongs in Harness Code, Not Context
- Recursive Agent Harnesses (RAH)
- Recursive Best-of-N Delegation
- Recursive Sub-Agent Delegation: Depth Limits and Trade-offs in Nested Hierarchies
- Red-Green-Refactor with Agents: Tests as the Spec
- Red-Team Your Blocking Monitor Before You Trust It
- Reducing Fixed CI Overhead Before Adding Shards
- Refactoring Runaway: Tangled Refactorings in Agent Patches
- Reference: Standards, Human Factors, Emerging, and Fallacies
- Reflective Prompt Evolution with Pareto Selection (GEPA)
- Reframed Exfiltration Defeats Wording-Based Defenses (Framing Gap)
- Reliability of an Automatically Selected Agent Harness
- Remote Agent Host Sessions over SSH and Dev Tunnels
- Remote Session Control for Local CLI Agents
- Rented Sandboxes for Coding-Agent Benchmark Runs
- Repairing Agent Prompts from Trace Contrast, Not Search
- Replayable Encrypted Reasoning Blocks in Agent Traces
- Repository Bootstrap Checklist: Wiring Agent Support
- Repository Map Pattern: AST + PageRank for Dynamic Code
- Repository Perturbation as Context-Reasoning Diagnosis (RepoMirage)
- Repository Skill Release Drift
- Repository-Level Retrieval for Code Generation
- Reproduce-Before-Report Verification Gate
- Reproducibility Artifacts as Agent Context
- Requirement Smells: No Category Signal, Density Unconfirmed
- Rerunnable Claim Graph as Shared Agent Memory
- Residual Completion for Stateful Agent Handoffs (CFRC)
- Restraint Rules Need External Enforcement
- Restricting a Coding Agent to a Single execute_code Tool
- Retrieval-Augmented Agent Workflows: On-Demand Context
- Retry-Switch-Abstain: A Runtime Tool-Recovery Policy
- Reverse-Engineered Executable Specifications for Agentic Program Repair
- Review Constraint Tests as a Second Acceptance Gate
- Review-Comment-Derived Benchmarks for Code Review Agents
- Review-Feedback-to-Rule Loop: Promoting Recurring PR Comments into Harness Rules
- Review-Then-Implement Loop for AI Agent Development
- Reviewer Habituation in Agent PR Review
- Reviewer Precision as a Pipeline Quality Proxy
- Reviewer Theme Distribution Audit for AI Code Review
- Reviewer's Playbook for Agent-Authored Pull Requests
- Revocable Resource-and-Effect Capabilities for Coding Agents (PORTICO)
- Rewriting a CLI Into a JSON Payload for Agents
- Rigor Relocation: Engineering Discipline with AI Agents
- Risk Architecture for AI-Native Engineering Teams
- Risk-Based Shipping: Review by Risk Matrix, Not by Default
- Risk-Based Task Sizing for Agent Verification Depth
- Risk-Score Threshold Calibration for Auto-Approval
- Role Orchestration on a Single Model
- Role-Declared Context Mode: Where the Inherit Decision Lives
- Rollback-First Design: Every Agent Action Should Be Reversible
- Rolling Out CLI Coding Agents at Organization Scale
- Rolling Out a Team-Embedded Agent Like a Tool
- Root Causes of Vibe-Coded Application Vulnerabilities
- Route Agent Peers by Enrolled Identity, Not Card Name
- Route-Parity Auditing of Agent Safeguards
- Router-Imposed Quality Ceiling: Committing Before Output
- Routing Break-Even: When a Cheaper Model Actually Pays
- Routing Decision Framework: Which Routing Pattern Fits Which Signal
- Routing Dependency Updates to Repair Agents by Budget
- RubricRefine: Pre-Execution Rubric Refinement for Code-Mode Tool Use
- Rule Lifecycle Metadata for Prunable Instruction Surfaces
- Run-Status vs Task-Status Confusion in Autonomous Agent Runs
- Runbooks as Agent Instructions: Agent-Followable Ops
- Runnable Documentation as Agent Verification
- Running Several Coding Agents Behind One Harness Interface
- Runtime Guard as an Installed Skill (Defense-as-Skill)
- Runtime Resource Limits as Prompt Context
- Runtime Scaffold Evolution: Agents That Build Tools
- SDLC-Phase Skill Taxonomy: Full-Lifecycle Skill Libraries
- SEO vs GEO — How Signals and Metrics Differ
- SKILL.md Frontmatter Reference: All Fields Explained
- SUDP: Secret-Use Delegation Protocol for Agentic Systems
- Safe Outputs Pattern for Trustworthy Agent Responses
- Sandbox + Approvals + Auto-Review Governance Triad
- Sandbox Forking: Branch Agent Runs From a Warm Snapshot
- Sandbox-Enforced PII Tokenization in Agent Workflows
- Sandboxed Coding Environments: Containers vs MicroVMs vs OS-Level Isolators
- Scanner-as-MCP-Server: Secret and Dependency Scans as Typed Agent Tools
- Scheduled Instruction File Fact-Checker for Accuracy
- Schema and Structured Data for GEO — AI Citation Guide
- Schema-Guided Graph Retrieval
- Scope Sandbox Rules to Harness-Owned Tools, Not Third-Party
- Scope-Matched Retrieval for Persisted Agent Skills
- Scoped Browser DevTools Access for Runtime Diagnosis
- Scoped Credentials via Proxy Outside the Agent Sandbox
- Scoped MCP Server Discovery: Most-Specific-Wins Resolution
- Scoped-Looking Permission Grants
- Scoring Constraint Loss and Tool Reach as Separate Risks
- Scoring a Compaction Policy on Latency and Billed Cost
- Scout-Then-Route: Verify the Handoff Before Routing
- Seamless Background-to-Foreground Handoff
- Secrets Management for AI Agents: Credential Injection
- Security Budget as Token Economics
- Security Constitution for AI Code Generation
- Security Drift in Iterative LLM Code Refinement
- Security Knowledge Priming for Code Generation (SPARK)
- Security-Aware Tool Descriptions for MCP Servers (SpellSmith)
- Seed-Variance Reporting and Measurable-Range Eval Design
- Seeding Agent Context: Breadcrumbs in Code
- Selective Autonomy from Copilot Feedback
- Selective Revalidation for Pending Agent Actions
- Selective Rewind Summarization: Compress Earlier Turns, Keep Recent Ones Intact
- Self-Correcting Memory: Evidence-Backed Claim Repair
- Self-Discover Reasoning: LLM-Composed Reasoning Structures
- Self-Explanation Loop
- Self-Healing Production Agent: Automated Regression Detection and Autofix PR
- Self-Healing Tool Routing
- Self-Reporting Loops: Autonomous Routines That File Their Own Backlog
- Self-Rewriting Meta-Prompt Loop
- Semantic Caching for Multi-Agent Code Systems
- Semantic Collapse Under Underspecified Prompts
- Semantic Context Loading: Language Server Plugins for Agents
- Semantic Density Optimization for Agent Codebases
- Semantic Intent Validation for Agent Skills
- Semantic Tool Output: Designing for Agent Readability
- Semantic Validation for Schema-Valid Agent Output
- Sensitive Terminal Prompt Interception
- Separating Exposure From Selection in GEO Measurement
- Separation of Knowledge and Execution in Agent Systems
- Serving-Stack Confounds in Tool-Call Evaluation
- Session Harness Sandbox Separation for Long-Running Agents
- Session Initialization Ritual: How Agents Orient Themselves
- Session Recap: Goal-Shaped Handoff at Context Boundaries
- Setup Documentation as an Install-Time Attack Vector
- Severity-Stratified Evaluation of Security Prompts
- Shadow Tech Debt Created by Autonomous AI Agent Commits
- Shallow Agent Test Coverage from Premature Termination (Lazy Generation)
- Shared Context Bundle Registry for Agent Teams
- Shortening Old Tool Results Under Context Pressure (Half-Life Truncation)
- Signal Over Volume in AI Review for AI Agent Development
- Silent Adoption of Corrupted Tool Returns by Agents
- Silent Handoff Failure in Delegated Code Search
- Silent-Failure Mechanism Taxonomy in Production Agent Runtimes
- Simulation and Replay Testing for Agent Verification
- Single-Branch Git for Agent Swarms: A Trade-Off Pattern
- Single-CLI Agent Platform: Create to Production in One CLI
- Single-Decision Approval of Vendor Skill Suites
- Single-Layer Prompt Injection Defense Anti-Pattern
- Situated Harness Layers: Fix at the Layer That Owns It
- Size Agent Comparisons by Run-to-Run Variance
- Sizing an Agent Migration Fan-Out by Diff Uniformity
- Skeleton Projects as Agent Scaffolding
- Skill Atrophy: When AI Reliance Erodes Developer Capability
- Skill Authoring Patterns: Description to Deployment
- Skill Authoring as Software Engineering: What Transfers
- Skill Composition Risk in Agent Ecosystems
- Skill Context Isolation: Forking the Skill into a Subagent Window
- Skill Evals: Measuring Skill Quality as a Dataset-Graded Unit
- Skill File Linting: Which Three Checks to Run First
- Skill Library Evolution: Lifecycle Governance for Agents
- Skill Library Refinement Loops: Organizational Feedback for Shared Skills
- Skill Library Technical Debt: Library-Time Maintenance for Agent Skills
- Skill Lift: Measuring What a Skill Adds at Runtime
- Skill Loadout Curation for Coding Agents
- Skill Misevolution in Self-Updating Skill Libraries
- Skill Over-Trust: Treating Topical Relevance as Evidence a Skill Helps
- Skill Packs: Registry Distribution Needs Pinning Discipline
- Skill Program Functions: Executable Guardrails Compiled From Past Failures
- Skill Reuse as Vendored Forking
- Skill Review Without a Token Cost Baseline
- Skill Specification Violation Fuzzing
- Skill Supply-Chain Poisoning
- Skill Test Coverage as a Release Gate
- Skill Tool as Enforcement: Loading Command Prompts at Runtime
- Skill as Instruction Surface and Callable API (Interpreter Skills)
- Skill as Knowledge Pattern for AI Agent Development
- Skill or MCP Server: Choosing a Capability's Delivery Mechanism
- Skill-Use Gates: Trigger, Compliance and Boundary
- Slop Detectors Fail as Per-Item Review Gates
- Slopsquatting: Hallucinated Package Names as a Supply-Chain Vector
- Solver-Externalized Constraint Reasoning (MaxSAT/SMT Encoding)
- Source Code Minification for State-in-Context Agents
- Source-Grounded Test Plan with Pre-Action Assertion Annotation
- Spec Complexity Displacement: When Specs Become Code
- Spec-Anchored Drift-Gated Architecture (Spec Growth Engine)
- Spec-Derived Execution as a Correctness Oracle
- Spec-Driven Development with Spec Kit
- Spec-Driven Test Generation: Contract Coverage Is the Lever
- Specialist Orchestrated Queuing for Multi-Agent SE (SPOQ)
- Specialized Agent Roles for Effective AI Pipelines
- Specialized Small Language Models as Agent Sub-Tools
- Specification Authority Boundary: Agents Propose, the Runtime Commits
- Specification Memory: What a Shared Agent Workspace Keeps
- Specification Portability Across Coding Agents
- Specification-First Convergence Without a Test Oracle
- Specification-Grounded Test Writing
- Specification-Path Testing: Same Contract, Different History
- Splitting an Agent Token Budget at the Scaling Inflection Point
- Splitting the Drift Judge from the Advisor (LivePlan)
- Sprint Contracts: Pre-Coding Success Agreements for Multi-Agent Tasks
- Stacked Agent Sessions on Unmerged Feature Branches
- Stacking Outer Loops Around the Agent
- Stage Elision Before Summarization
- Stage-Targeted Prompt Structure for Pull Request Outcomes
- Staged Evidence Gates for Agentic Program Repair
- Staged Literal Porting with a Per-Stage Numeric Oracle
- Staggered Agent Launch: Preventing Thundering-Herd in Swarms
- Stakeholder Trust Through Evals and Observability
- Stale AI Configuration Artifacts (Context Rot)
- Standard-Grounded NFR Specs: Quality Up, Correctness Flat
- Standards as Agent Instructions for AI Agent Development
- State-Bound Evidence and Typed Revision Contracts for Repair Loops
- State-Conditioned Evidence Selection for Mid-Task Retrieval
- Stateful Agent Evals via State Snapshots and Transition Assertions
- Stateful Iteration State-Carry: Typed Persistent State for Long Agent Loops
- Stateless MCP: One Request per Tool Call
- Static Difficulty Estimation for Agent Issue Triage
- Static Roster vs Runtime Subagent Definition
- Static-First Shell Command Gating with Selective Escalation
- Steering Running Agents: Mid-Run Redirection and Follow-Ups
- Step Budgets and Trust in Agent-Generated Code Tours
- Step-by-Step: Building Your First Eval-Driven Feature
- Stochastic-Deterministic Boundary as First-Class Contract
- Strained Coherence as a Pre-Failure Signal in Agent Trajectories
- Strategy Over Code Generation: Why AI Speed Doesn't Fix Wrong Goals
- Structural Coverage Criteria for Agent Workflows
- Structural Monitoring for Covert Safeguard-Weakening
- Structure Prompts with Static Content First to Maximize Cache Hits
- Structure-Aware Diff Labeling with Two-Stage LLM Pipelines
- Structured Agentic Software Engineering (SASE)
- Structured Domain Retrieval: Knowledge Graphs and Case-Based Reasoning
- Structured Output Constraints: Reducing Hallucination
- Structured Task-State Ledger for Tool-Calling Agents (LedgerAgent)
- Stuck-Loop Recovery: Detecting and Escaping Non-Converging Agent Loops
- Sub-Agents for Fan-Out Research and Context Isolation
- Subagent Schema-Level Tool Filtering for AI Agents
- Subagent vs In-Context Skill Execution
- Subtask-Level Memory for Software Engineering Agents
- Sufficiency-Tightness Decomposition for Agent-Authored Permissions
- Suggestion Gating: Fewer Completions, Better DX
- Supply-Chain Security Debt in Agent Pull Requests
- Suspect the Harness Before the Model on a Regression
- Swarm Migration Pattern
- Swarm Skills: Multi-Agent Extension of the Agent Skills Standard
- Symphony: Open Spec for Issue-Tracker-Driven Coding Agent Orchestration
- Symptom-First Bug Triage for Agent Code
- Symptom-Reduction-as-Root-Cause: Why Oracle Tests Alone Miss Architectural Drift
- System Prompt Altitude: Specific Without Being Brittle
- System Prompt Replacement for Domain-Specific Agent Personas
- System Prompt as Secret Store (OWASP LLM07)
- System-Level Optimization Pipeline
- TDD Interaction Models: Throughput Versus Test Quality
- Tab-Accept Rate as a Proxy for Critical Engagement
- Tail Control for Agent Workflows: Engineering for the Failure Tail, Not the Average
- Task Alignment: The Selective-Compliance Gap Benchmarks Miss
- Task Category as the Security Review Routing Key
- Task Completion as Tool Certification (Silent Tool Rot)
- Task Feasibility Awareness: Stop Before You Start
- Task List Divergence as Instruction Quality Diagnostic
- Task Shape Decides What a Heavier Agent Harness Buys
- Task-Based Access Control with Hybrid Inspection
- Task-Specific Agents vs Role-Based Agents
- Task-Uniform Agent Permissions Ignore Where Failures Land
- Team OS: Coding-Agent Repo as Cross-Functional Team Brain
- Team Onboarding for AI Agent Workflows and Adoption
- Team Shared-Language Desync from Removed Review Friction
- Temporal Token Routing: Batch and Flex Tiers for Non-Urgent Work
- Temporary Compensatory Mechanisms in Agent Harnesses
- Terminal Tool Output Compression: Filtering Predictable Noise at the Harness
- Terminal Tools for Agents: send_to_terminal and Background Interaction
- Terminal-First Agent Interfaces with Browser Escalation
- Test Harness Design for LLM Context Windows
- Test Oracles That Read Their Expectation From the Code
- Test-Driven Agent Development: Tests as Spec and Guardrail
- Test-Driven Intent Clarification: Tests as Intermediate Alignment Artifacts
- The 7 Phases of AI-Assisted Feature Development
- The AI Development Maturity Model: From Skeptic to Agentic
- The AI Knowledge Generation Fallacy: LLMs Recombine, Not Invent
- The AX Stack: A Layered Model of an AI Coding Agent's Prompt-to-Compile Path
- The Addictive Flow State of Agent-Assisted Development
- The Advisor Strategy: Frontier Model as Strategic Advisor
- The Agent Stack Bet: Architectural Decisions for Production Agents
- The Agent-on-Agent Maintenance Penalty
- The Anthropomorphized Agent for AI Agent Development
- The Bottleneck Migration When Humans Supervise Agents
- The Citizen-Agent-Expert Operating Model for AI Coding
- The Compliance Trap: Consuming Conflicting Agent Memory
- The Consistent Capability Fallacy in LLM Agent Design
- The Context Ceiling -- Where AI Fails Expert Architects
- The Copy-Paste Agent Anti-Pattern in AI Development
- The Decoding Harness Is Part of the Agent Attack Surface
- The Delegation Decision: When to Use an Agent vs Do It Yourself
- The Effortless AI Fallacy for AI Agent Development
- The Error-Class Governance Loop for Instruction Libraries
- The Eval-First Development Loop for AI Agent Features
- The First Edit Predicts Whether an AI Completion Survives
- The Handoff Tax: What a Receiving Model Should Inherit
- The Harness as Product: What Listed-Rate Pricing Buys
- The Implicit Knowledge Problem for AI Coding Agents
- The Infinite Context Anti-Pattern in Agent Systems
- The Instruction Compliance Ceiling: How Rule Count Limits AI
- The Kitchen Sink Session Anti-Pattern in AI Agents
- The LLM Laziness Deficit Fallacy: Restraint Comes From Harness, Not Instruction
- The Meat Proxy: Relaying Agent Output Without Reading It
- The Merge-Conflict Resolution Skill: What to Encode
- The Model Economics of Agent Swarms: Cost and Width
- The Model Preference Fallacy in Comparison Content
- The No-Op Test: Prune Agent Docs by Behavior, Not Length
- The Orchestrator's Attention Budget: Delegating to Protect Context
- The Patchwork Problem in LLM-Generated Code
- The Plan-First Loop: Always Design Before Writing Code
- The Post-Authorization Execution Trust Gap in Remote MCP
- The Productivity-Experience Paradox in AI-Assisted Development
- The Prompt Tinkerer Anti-Pattern in Agent Workflows
- The Ralph Wiggum Loop: Fresh-Context Iteration Pattern
- The Reasoning-Complexity Trade-off
- The Recall Trap: Tuning a Code Retriever on Recall@k at a Fixed Slot Budget
- The Research-Plan-Implement Pattern
- The Security Review Gap in AI-Authored PRs
- The Skill Closure Declaration Gap
- The Software Factory Model: Industrializing Agent Loops
- The Specification as Prompt: Existing Artifacts as Agent
- The Subagent Inheritance Contract: What Crosses Down
- The Synthetic Ground Truth Fallacy in Agent Evaluation
- The Task Framing Irrelevance Fallacy in Agent Prompting
- The Test Homogenization Trap: When LLM-Generated Tests Mirror Model Blind Spots
- The Think Tool: Mid-Stream Reasoning for AI Agents
- The Three Loops of Agentic Coding: A Diagnostic Vocabulary
- The Token Price Index Fallacy in Agent Cost Planning
- The Yes-Man Agent: Compliance Without Verification
- Three Knowledge Tiers: Sourced, Unverified, Hallucinated
- Three Reasoning Spaces: Plan-Bead-Code Phase Gates
- Three-Depth In-Session Security Review
- Three-Vector Evasion Taxonomy for Agent Security Tests
- Throwaway-Prototype Skill: Build to Discard, Keep Only the Answer
- Tiered Code Review: AI-First with Human Escalation
- Tiered Memory Architecture: Episodic-to-Semantic Consolidation for Long-Running Agents
- Timeout Oracles for Agent-Written Code
- Token Preservation Backfire for AI Agent Development
- Token Reduction Mistaken for Cost Reduction
- Token-Cost Profiling and Reduction for Always-On Agentic Workflows
- Token-Efficient Code Generation: Structural Beats Prompting
- Token-Efficient Tool Design: Tools That Don't Eat Your Context
- Tokenizer Swap Tax: Budgeting for Model Migrations That Change Token Counts
- Tool Architecture Moves Consistency, Not Resolve Rate
- Tool Calling Schema Standards for AI Agent Development
- Tool Cloning and Provenance Assessment in Agent Ecosystems
- Tool Confirmation Carousel: Batched UI for Per-Call Approvals
- Tool Description Quality for Effective Agent Guidance
- Tool Engineering (Training Module)
- Tool Minimalism and High-Level Prompting
- Tool Necessity Probing: Reading Tool-Call Decisions From Hidden States
- Tool Operability: Interfaces That Survive a Lost Response
- Tool Preamble: User-Visible Status Updates Before Tool Calls
- Tool Signing and Signature Verification for Agents
- Tool-Call Success as Workflow Effect Evidence
- Tool-Invocation Attack Surface in Coding Agents
- Tool-Schema Bias: Measure the Interface Before You Ship
- Tool-Use Sim-to-Real Perturbation Taxonomy
- Tools as Typed Code Stubs (Programmatic Tool Calling)
- Toolset Agentization: Wrapping Co-Used Tools as Sub-Agents
- Topical Authority — Entity Coverage for AI Citation
- Traces Need Feedback to Power Learning
- Training-Data Gravity: Agents Default to Deprecated APIs
- Trajectory Attribution for Context Repair (TRACE)
- Trajectory Decomposition: Diagnose Where Coding Agents Fail
- Trajectory Logging via Progress Files and Git History
- Trajectory Poisoning of Promoted Agent Skills (PoisonedEvolution)
- Trajectory Pre-Filter for Failure Diagnosis (TrajAudit)
- Trajectory Projection: A Flattened View of Agent Traces
- Trajectory as the Monitoring Unit for Production Agents
- Trajectory-Aware Benchmark Subset Selection for Agents
- Trajectory-Conditioned Model Escalation (SWE-Router)
- Transcript-Driven Permission Allowlist
- Transcript-Measured Review Coverage
- Treat Task Scope as a Security Boundary
- Treating Agent Delegation as Routing, Not Authorization
- Treating Agent Safety as Uniform Across a Session (Cold-Start Safety Gap)
- Treating File Secrecy as Skill Confidentiality
- Treating Memory-Injection Rate as Security Evidence
- Treating Read Denial as Confidentiality for a Build Input
- Treating a Clean Final State as Boundary-Compliance Evidence
- Treating a Clean Merge as Compatibility Evidence
- Treating a Clean Static Scan as Security Evidence
- Treating a Local Agent Session Trace as Audit Evidence
- Treating a Worktree as a Safety Boundary
- Trigger-Level Gating for Autonomous Agent Intake
- Trigger-to-Function Architecture for Unprompted Codebase Maintenance
- Trust Without Verify: Skipping Agent Output Checks
- Trusting Claimed Prior Approval in Agent Review Gates
- Trusting Human Review to Catch Deliberate Agent Sabotage
- Trusting Model-Level Privilege Restraint at Tool Selection
- Trusting Tool Error Messages as Implicit Authority (Error-Path Injection)
- Trusting a Skill Scanner's Verdict as a Security Judgment (Green-Check Fallacy)
- Typed Context Buys Addressability, Not Token Savings
- Typed Generation Contracts for Grounded Extraction
- Typed Memory Provenance and Assertion Release Gating
- Typed Pseudocode for Skill Libraries (Skill-as-Pseudocode)
- Typed Schemas at Agent Boundaries for Multi-Agent Systems
- Typed Tool Surfaces and Out-of-Loop Correctness Gates
- Ubiquitous Language for AI Plans
- Unbounded Agent Feedback Paths (Infinite Agentic Loops)
- Unbounded Consumption: Bounding Agent Resource Use Against DoS and Denial-of-Wallet
- Unix CLI as the Native Tool Interface for AI Agents
- Unsignalled Tool Failure: Returning Success With an Unusable Payload
- Unstated-Contract Bugs: Sort Tickets by Information Gap
- Unversioned Scaffolding Commands Pull Stale Templates
- Usability Pressure as a Silent Security-Regression Vector
- Usage-Reinforced Memory Decay for Long-Running Agents
- Use a Public-Web Index to Gate Automatic URL Fetching
- Using the Agent to Analyze Its Own Evaluation Transcripts
- Utility-Model Split: Background Tasks on a Cheaper Model
- Validating Token-Optimized Formats Inside Agentic Loops
- Validity-Estimate Stopping for Noisy Verify-Repair Loops (VRR-Stop)
- Variance-Based RL Sample Selection
- Velocity-Quality Asymmetry: Why AI Speed Gains Fade
- Verbatim Failure Records in Small-Model Agent Transcripts
- Verification Capacity Saturation: Three Levers, One Default
- Verification Capacity as the Agent Quality Ceiling
- Verification Ledger for Tracking Agent Output Quality
- Verification Surface: Match the Tool to the Failure
- Verification-Centric Development for AI-Generated Code
- Verification-Gated Agent Autonomy via Automated Review
- Verifier-Driven Parallel Coding Agents (Glite ARF)
- Verify Observability in Agent-Generated Code
- Verify-Gated Completion as Admission Control
- Verifying LLM-Generated Cryptographic Code
- Version-Controlled Agent Context (Git Context Controller)
- Vetting Tool Definitions for Exfiltration Signatures
- Vibe Coding: Outcome-Oriented Agent-Assisted Development
- Visible Thinking in AI-Assisted Development
- Voting / Ensemble Pattern for AI Agent Development
- WIP=1 and Little's Law: Kanban Throughput Theory for Agent Task Design
- WRAP Framework for Writing Agent-Ready Issue Descriptions
- Weakest Consistent Learning: What Agent Loops Should Persist
- Web Search Agent Loop: Iterative Research Patterns
- WebMCP: Browser-Hosted Tool Contracts for In-Page AI Agents
- What Evals Are and Why AI Agents Need Them for Quality
- What is GEO — Generative Engine Optimization Defined
- When Developers Understand Less of Their Own Codebase
- When a Skill Graph Cannot Beat the Ranker (Pre-Filter Topology Bound)
- Which Task You Delegate Changes Poisoned-Repo Exposure
- Whole-Codebase Visibility as a Migration Prerequisite
- Why an Encoded Rule Still Fails After a Passing Eval
- Wiki Memory: Agent-Maintained Compressed Knowledge Base
- Windows Sandboxing for Coding Agents
- Within-Task Model Cascade: Designing the Escalation Gate
- Workflows for AI Agent Development
- Workload-Keyed Sandbox Selection for Agent-Generated Code
- Workspace Topology as an Indirect Injection Attack Vector
- Workspace-Hosted Skills: Authorship Outside the Repo
- Worktree Isolation: Parallel Agent Sessions in Safe Sandboxes
- Write Agent Rules You Can Grade From the Transcript
- Write Tool Descriptions as Agent Onboarding Documents
- Writing Your First Agent Evaluation Suite from Scratch
- Zero Violations Is Not Evidence Your Hook Works
- llms.txt: Full Specification, Adoption, and Limitations
- llms.txt: Making Your Project Discoverable to AI Agents
- pass@k and pass^k: Capability and Consistency Metrics
tool-engineering¶
- Advanced Tool Use: Scaling Agent Tool Libraries
- Agent Plugins: Portable Packaging With Client-Defined Trust
- Agent-Aware CLI Behavior via Environment Variable
- Agent-Computer Interface (ACI): Tool Design as UX Discipline
- Agent-Discoverable Slash Commands
- Agent-Tuned Code Search: Retrieval Built for the Loop
- Auditing Agent Tool Chains for Silent Partial Success
- Auth-Isolation as the MCP-vs-CLI Selection Heuristic
- Batch File Operations via Bash Scripts for AI Agents
- Browser Automation as a Research Tool: Bypassing Bot Detection
- CLI Scripts as Agent Tools: Return Only What Matters
- CLI-First Skill Design
- Chance-Corrected Shortlist Depth Sizing for Tool Retrieval (Bits-over-Random)
- Choosing an Agent Tool Interface: Shell or Typed Catalog
- Closed-Loop Agent Training from Tool Schemas
- Code Interpreter as a Primary Agent Tool
- Codebase-Derived Pattern Libraries as Agent Context
- Conditional Hook Execution: Filter Hooks by Tool Pattern
- Consolidate Agent Tools to Reduce Cognitive Overhead
- Cross-IDE Plugin Discovery: One Install Surface, Many Consuming Agents
- Cross-Repo Agent Search: GitHub-API-Backed Text Search Beyond the Workspace
- Cursor Customize Page: Unified Surface for Agent Primitives
- Designing Agent Tools Like APIs
- Designing for Agent Consumers (Agent Experience)
- Documentation-Grounding MCP Servers for Vendor SDKs
- Edit Format Selection: Diff vs. Search-Replace vs. Full Rewrite
- Effort-Aware Hooks: Reading the Reasoning Tier from PreToolUse and PostToolUse
- Engineering: Tools, Review, Verification, Security, and Observability
- Filesystem-Based Tool Discovery for AI Agent Development
- Five Design Decisions for MCP Servers and Clients
- Function-Level Debugger Interfaces for Coding Agents
- Future-Based Asynchronous Function Calling
- Generated Procedure Drivers: Skills That Emit a Program
- Google ADK Skills: Portable SKILL.md Across ADK Agents
- Graceful Tool-Output Truncation: The PARTIAL Signal
- Headless-First Services: APIs for Agent Consumers
- Hint-Driven Concurrency for Read-Only MCP Tools
- Hook Catalog for Claude Code Enforcement
- Hook Exec Form vs Shell Form: Shell-Injection-Safe Hook Commands
- Hooks Invoking MCP Tools: Closing the Loop Between Policy and Tool Execution
- Hooks and Lifecycle Events: Intercepting Agent Behavior
- Indexed Regex Search for Agent Tools
- Judging MCP Capability Readiness by Client Adoption
- Lexical-First Retrieval for Agentic Search: When BM25 Is Enough
- Listener-State Naming for User-Invoked Agent Skills
- Local Plugin Scaffolding via `claude plugin init` and Auto-Loaded `.claude/skills`
- MCP Client Design: Building Robust Host-Side Logic
- MCP Elicitation: Servers Requesting Structured Input Mid-Task
- MCP LLM Sampling: Servers Requesting AI Inference Mid-Tool
- MCP Server Design: Building Agent-Friendly Servers
- MCP Tool Result Persistence via _meta Annotation
- MCP alwaysLoad: Classifying Servers as Eager or Just-in-Time
- MCP-vs-CLI Cost Ratios Are a Property of the Scaffolding
- Machine-Readable Error Responses for AI Agents (RFC 9457)
- Managing Agent Skills from the GitHub CLI with gh skill
- MessageDisplay Hook: Transforming Assistant Text at the Display Boundary
- Model-Switch Lifecycle Hooks: Gating a Mid-Session Model Change
- Natural-Language Customization Bootstrap
- On-Demand Skill Hooks: Session-Scoped Guardrails via Skill Invocation
- One-Shot Record and Deterministic Replay for Periodic Agent Tasks
- OpenAPI Documentation Smells for Agent-Ready APIs
- OpenAPI as the Source of Truth for Agent Tool Definitions
- Oracle Poisoning: Knowledge Graph Corruption Against Tool-Using Agents
- Out-of-Band Hook Notifications via terminalSequence
- Override Pattern: Reusing Interactive Commands in Automated Pipelines
- Per-Surface Verification of Agent Plugin Packages
- Plugin Background Monitors: Declarative Supervision Auto-Armed at Session Start
- Plugin-Activated Main-Agent Override and Bin/ PATH Injection
- Poka-Yoke for Agent Tools: Mistake-Proof Tool Interfaces
- PostToolBatch Hook: Once-Per-Decision-Cycle Injection at the Batch Boundary
- PostToolUse Hook for BSD/GNU CLI Incompatibilities
- PostToolUse Hooks: Automatic Formatting and Linting
- PostToolUse Output Replacement: Hooks That Rewrite Tool Results
- PostToolUse continueOnBlock: Refusal With a Load-Bearing Reason
- PowerShell Tool: Native Windows Shell for Claude Code
- Pre-Execution Risk Classification for Terminal Commands
- PreCompact Hook: Vetoing Compaction at Lifecycle Boundaries
- Product-Operation Tool Surfaces for In-Product Agents
- Production MCP Agent Stack: Sequencing Six Decisions into One Deployment
- Project Writing Skill: House Style as Model-Invocable Skill
- Proprietary-to-Open-Standard Tool Migration (Copilot Extensions to MCP)
- Push-Event MCP Channels: Inverting the Pull-Tool Polarity
- Reactive Environment Hooks: CwdChanged and FileChanged
- Reloading Skills Mid-Session in Claude Code
- Restricting a Coding Agent to a Single execute_code Tool
- Rewriting a CLI Into a JSON Payload for Agents
- Runtime Scaffold Evolution: Agents That Build Tools
- SKILL.md Frontmatter Reference: All Fields Explained
- Scanner-as-MCP-Server: Secret and Dependency Scans as Typed Agent Tools
- Scoped MCP Server Discovery: Most-Specific-Wins Resolution
- Self-Healing Tool Routing
- Semantic Tool Output: Designing for Agent Readability
- Skill Authoring Patterns: Description to Deployment
- Skill Authoring as Software Engineering: What Transfers
- Skill Composition Risk in Agent Ecosystems
- Skill Context Isolation: Forking the Skill into a Subagent Window
- Skill Library Evolution: Lifecycle Governance for Agents
- Skill Library Technical Debt: Library-Time Maintenance for Agent Skills
- Skill Reuse as Vendored Forking
- Skill Shell Execution Gate: Disabling Inline Shell from Skills
- Skill Tool as Enforcement: Loading Command Prompts at Runtime
- Skill as Instruction Surface and Callable API (Interpreter Skills)
- Skill as Knowledge Pattern for AI Agent Development
- Skill or MCP Server: Choosing a Capability's Delivery Mechanism
- StopFailure Hook: Observability for API Error Termination
- Symbol Ranking for Agent File Pickers
- Terminal Tool Output Compression: Filtering Predictable Noise at the Harness
- Terminal Tools for Agents: send_to_terminal and Background Interaction
- Terminal-First Agent Interfaces with Browser Escalation
- The Merge-Conflict Resolution Skill: What to Encode
- Token-Efficient Tool Design: Tools That Don't Eat Your Context
- Tool Architecture Moves Consistency, Not Resolve Rate
- Tool Calling Schema Standards for AI Agent Development
- Tool Cloning and Provenance Assessment in Agent Ecosystems
- Tool Description Quality for Effective Agent Guidance
- Tool Engineering (Training Module)
- Tool Engineering: Designing and Managing AI Agent Tooling
- Tool Minimalism and High-Level Prompting
- Tool Necessity Probing: Reading Tool-Call Decisions From Hidden States
- Tool-Schema Bias: Measure the Interface Before You Ship
- Tools: Claude Code, Cursor, and GitHub Copilot
- Toolset Agentization: Wrapping Co-Used Tools as Sub-Agents
- Unix CLI as the Native Tool Interface for AI Agents
- Video Transcript Skill: Meeting Recording to Markdown
- Web Search Agent Loop: Iterative Research Patterns
- Write Tool Descriptions as Agent Onboarding Documents
training¶
- Autonomous Research Loops: Loops That Know When to Stop
- Context Engineering (Training Module)
- Earned-Complexity Agent Maturity Ladder
- Eval Engineering (Training Module)
- Eval-Driven Development Training for AI Agent Teams
- Foundational Disciplines for AI-Assisted Development
- GitHub Copilot Advanced Patterns: Multi-Agent and Automation
- GitHub Copilot Platform Surface Map: All Capabilities
- GitHub Copilot Training Modules for Engineering Teams
- GitHub Copilot: Context Engineering & Agent Workflows
- GitHub Copilot: Customization Primitives and Stack
- GitHub Copilot: Harness Engineering for Agent-Ready Code
- GitHub Copilot: Model Selection, Routing, and Costs
- GitHub Copilot: Team Adoption and Governance Guide
- Grading Strategies for Eval-Driven Development
- Hardening Agent Evals for Production-Grade Reliability
- Harness Engineering (Training Module)
- How the Four Agent Engineering Disciplines Compound
- L0 → L1: Making the Repo Readable
- L1 → L2: Adding Feedback Loops
- L2 → L3: Building Mechanical Enforcement
- L3 → L5: Reaching Agent-First
- Prompt Engineering for Agent Instructions and Systems
- Step-by-Step: Building Your First Eval-Driven Feature
- The Eval-First Development Loop for AI Agent Features
- Tool Engineering (Training Module)
- Training Modules
- What Evals Are and Why AI Agents Need Them for Quality
- Writing Your First Agent Evaluation Suite from Scratch
workflows¶
- A Governance Framework for Production Agents
- AI Adoption Footprint: The Segmented Shape of Engineering Orgs
- AI Bot CI/CD Workflow Reliability by Agent
- AI Crawler Policy: robots.txt for the Three-Tier Crawler Landscape
- AI Slop as a Process Problem: Encoding Quality Standards as Pipeline Gates
- AI-Powered Vulnerability Triage for AI Agent Development
- AST-Grounded Critic Loop for Documentation Maintenance
- Accumulated Behavioral Rules from Review Feedback
- Acknowledged-Debt Ledger with Next-Trigger Conditions
- Adversarial Multi-Model Development Pipeline (VSDD)
- Agent Chat History as a First-Class Artifact
- Agent Commit Attribution: Signed Commits and Agent Identity
- Agent Debug Log Panel: Chronological Event Inspection for Session Debugging
- Agent Debugging: Diagnosing Bad Agent Output
- Agent Development Lifecycle for Agent Products
- Agent Environment Bootstrapping for AI Agent Development
- Agent Governance Policies for AI Agent Development
- Agent Harness: Initializer and Coding Agent Pattern
- Agent Loop Middleware — Safety Nets and Message Injection
- Agent Mission Control for Orchestrating Agent Tasks
- Agent Observability with OpenTelemetry and Trajectory Logging
- Agent PR Volume vs. Value: The Productivity Paradox
- Agent Project State Purge: Clean-Slate Session Reset
- Agent-Authored PR Integration: Collaboration Signals That Determine Merge Success
- Agent-Driven Deployment: What to Delegate and What to Gate
- Agent-Driven Fuzzing with Human-Gated Crash Triage
- Agent-Driven Greenfield Product Development from Scratch
- Agent-Driven PR Slicing
- Agent-Generated Onboarding Guide as a Durable Artifact
- Agent-Laundered Bug Reports
- Agent-Led Dev-Environment Iteration with Validation and Rollback
- Agent-Powered Codebase Q&A and Onboarding Workflow
- Agent-Proposed Merge Resolution
- Agentic Education: Persona Progression for Teaching AI Coding Tools
- Agentic Flywheel: Self-Improving Agent Systems
- Agentic-Agile: Adapting Agile Rituals for Agent Work
- Answer-First Writing: Structure Content for AI Retrieval
- Applying Coding Agents to Non-Code Tasks
- Architecting a Central Repo for Shared Agent Standards
- Artifact-Driven Workflow Compilation for Agent Execution
- Assertion Density — Stats and Quotes Over Vague Claims
- Atomic Pages and Chunking — One Concept Per Page for RAG
- Auto-Triage Workflow: Bug-Monitoring Agent that Connects Related Reports and Opens Fix PRs
- Autonomous Research Loops: Loops That Know When to Stop
- Backlog Triage as a Named Agent Skill
- Batched Suggestion Application: Bulk-Apply Agent Fixes on PRs
- Behavioral Specification Elicitation Before Synthesis (SpecFirst)
- Bounded Agent Steps Inside a Deterministic Workflow
- Bounding a Headless Codex Run Without a Turn Cap
- Brownfield to Agent-First: Repo Maturity Framework
- Browser Automation as a Research Tool: Bypassing Bot Detection
- Building Custom Agents from Substrate to Production (Agents All the Way Down)
- Burn the Boats — Commitment-Forcing Deprecation
- CARE: Three-Party Stage-Gated Engineering of LLM Agents
- CLI-IDE-GitHub Context Ladder for AI Agent Development
- Canary Rollout for Agent Policy Changes
- Chain-of-Verification for Coding Agents
- Channels Permission Relay
- Chat-Platform Agent Delegation: Invoking Cloud Coding Agents from Team Channels
- Choosing the Right Surface for a Coding Agent Task
- Classical SE Patterns as Agent Design Analogues
- Classifier-Subagent Run Mode for Per-Call Permission Routing
- Classifying and Auto-Correcting Coding Agent Misbehaviors (Wink)
- Claude Code --bare Flag
- Claude Code /batch and Worktrees for AI Agent Development
- Claude Code Review
- Clock-In / Clock-Out Protocol: Bracketed Session Continuity
- Closed-Loop Agent Training from Tool Schemas
- Closed-Loop CI Failure Remediation with Cloud Coding Agents
- Closed-Loop Role-Based Refinement for Agent Systems
- Cloud Parallel Review Pattern
- Cloud Planning with Inline-Comment Review and Execute-Anywhere Choice
- Cloud-Agent Session Bootstrap: Cached Install plus Per-Session Start
- Cloud-Agent Three-Layer State Decoupling
- Cloud-Local Agent Handoff for AI Agent Development
- Cloud-Scheduled Routines vs Local Session Scheduling
- Coding Agent Scope Expansion: When to Extend Beyond the Codebase
- Coding-Agent Misalignment Forms (Seven-Symptom Taxonomy)
- Coding-Agent Reversibility: Platform Choice as a Two-Way Door
- Comment-Triggered Agent Dispatch on Issues and PRs
- Committee Review Pattern for Multi-Agent Code Review
- Compound Engineering: Learning Loops That Make Each Feature Easier
- Concurrent Agent Pull Requests and Merge-Conflict Cost
- Consistent-format customer capture
- Containment Playbook: npm-to-Signing-Channel Compromise
- Context-Injected Error Recovery for AI Agent Development
- Continual Learning for AI Agents: Three Layers of Knowledge Accumulation
- Continuous AI (Agentic CI/CD) for AI Agent Development
- Continuous AI: A Navigation Map of Always-On Agent Workflows
- Continuous Agent Improvement: Iterating on Agent Quality
- Continuous Autonomous Task Loop
- Continuous Documentation as an Agent-Driven Practice
- Continuous Triage: Automating Issue Classification with AI Workflows
- Convention Over Configuration for Agent Workflows
- Convergence Detection in Iterative Agent Refinement
- Copilot CLI Agentic Workflows for AI Agent Development
- Copilot CLI BYOK and Local Model Support
- Copilot Cloud Agent Three-Phase Execution Model
- Copilot Unified Sessions View and CLI Agent in JetBrains IDEs
- Cross-Functional Knowledge Artifacts
- Cursor 3 Agents Window: Parallel Agents and Worktree Isolation
- Cursor Automations: Event-Triggered Agents and /automate
- Cursor Multi-Root Workspaces for Cross-Repo Agent Edits
- Daily-Use Skill Library: Encoding Your Process as Agent Skills
- Delegating Change Descriptions to the Agent
- Delegating Delivery Stages to GitHub Agent Apps
- Delegating Dependabot Pull Request Triage to an Agent
- Delegating Multi-Hunk Bug Repair to Coding Agents
- Deliberate AI-Assisted Learning: Accelerating Skill Acquisition
- Dependabot Agent Assignment for AI-Driven Vulnerability Remediation
- Design Docs as the Durable Artifact
- Deterministic Orchestration for Structured Modernization
- Dev Containers for AI Coding Agents: Claude Code vs Copilot CLI
- Developer Control Strategies for AI Coding Agents
- Discovery-Only Refactor Pass: Surface Candidates Before Touching Code
- Distilled Bootstrap Contract: Agent-Authored Repo Setup
- Documentation-Guided Legacy Migration: Architecture Docs as a C-to-Rust Blueprint
- Downstream Disclosure Coordination for Agent-Found Defects
- Earned-Complexity Agent Maturity Ladder
- Ecosystem-Level Integration Friction Governance
- Editor and Manager Surface Separation in Agent IDEs
- Emergent Architecture in AI-Driven Codebases
- Encoding Tacit Knowledge into Agent Improvement Loops
- Enterprise Skill Marketplace: Distribution and Quality
- Enterprise-Managed Plugin Governance for Agent CLIs
- Entropy Reduction Agents: Automated Codebase Hygiene
- Escape Hatches: Unsticking Stuck Agents
- Eval-Driven Development: Write Evals Before Building Agent
- Event-Driven Agent Routing for Multi-Team AI Pipelines
- Evidence-Based Allowlist Auto-Discovery for Agents
- Evidence-Chain Run Logs: Bracket the Reported Symptom
- Execution Lineage: DAG of Artifacts vs Agent Loops
- Execution-First Delegation: The AI-as-Executor Pattern
- Exhaustive Retrieval for Listing Questions
- Experiential-Learning Setup Agents with Snapshot Rollback (SetupX)
- Factory Over Assistant: Orchestrating Parallel Agent Fleets
- Failure-Driven Iteration for Improving Agent Workflows
- Fallacies for AI Agent Development
- Fan-Out Synthesis Pattern for AI Agent Development
- File-Based Agent Coordination for AI Agent Development
- Five-Pass Blunder Hunt: Repeated Critique Passes for Plans
- Four-Phase Agent Delegation with Curated Artifacts
- Frozen Task Sets for Affordable Agent A/B Testing
- GEO for Technical Docs: Developer Documentation Checklist
- Generating Tests From Agent-Written Code (Code-First Oracle Bias)
- Generative Engine Optimization for Developer Sites
- GitHub Agentic Workflows for Automating Dev Processes
- GitHub Copilot Dedicated App as Agent-First Surface
- GitHub Copilot: Context Engineering & Agent Workflows
- GitHub Models in Actions for AI-Driven CI Workflows
- Goal Contract: Separating the Doer from the Done-Checker
- Goal Monitoring and Progress Tracking for Long-Running Agents
- Golden Journeys: Restartability as a First-Class Verification Primitive
- Headless Claude in CI: Using -p and --max-turns for Safe Pipeline Integration
- Hook Catalog for Claude Code Enforcement
- How AI Engines Cite — ChatGPT, Perplexity, Claude, Gemini
- How Teams Build SE Agents: A Seven-Stage Build Loop
- Human-in-the-Loop Placement: Where and How to Supervise
- Humans and Agents in Software Engineering Loops
- Hyper-Personalized Software: The Return of RAD
- Hypothesis-Driven Debugging: Instrument Before You Patch
- In-Place Atomic Replacement for Agent-Driven Ports
- In-Session Transcript Search: Navigating Long Agent Conversations
- In-Thread Side-Channel: Bounded Side Questions Without Losing the Main Task
- Incident Log Investigation Skill: Parallel Queries
- Incremental Verification: Check at Each Step, Not at the End
- Initiatives and Community: Tracking the Agentic Engineering Landscape
- Interaction-Pattern Evaluation for Agentic PRs
- Introspective Skill Generation: Mining Agent Patterns
- Inversion Analysis: Surface Capabilities Competitors Cannot Replicate
- Issue-Tracker as Agent Dispatch Surface
- Issue-to-PR Delegation Pipeline for AI Agent Development
- Kaizen-Style Continuous Code Quality Loop (Pomona)
- Knowledge-Based Pull Requests for Cross-Trust-Boundary Contributions
- L0 → L1: Making the Repo Readable
- L1 → L2: Adding Feedback Loops
- L2 → L3: Building Mechanical Enforcement
- L3 → L5: Reaching Agent-First
- LLM Refactoring Adoption Patterns
- LLM Static Verification Against Natural-Language Requirements
- LLM-as-Judge Evaluation with Human Spot-Checking
- Labels as Locks: Pipelined Backlog Processing with Stage Gates
- Large-Codebase Coding-Agent Failure Patterns (Sourcegraph Five)
- Late Requirement Arrival in Agent Sessions
- Law of Triviality in AI PRs for AI Agent Development
- Lay the Architectural Foundation by Hand Before Delegating
- Lazy Worktree Isolation: Enter the Worktree on First Write, Not on Dispatch
- Legacy Code Archaeology: Reconstruct Intent Before Migrating
- Long-Running Agents: Durability and Resumability Across Sessions
- Loop Detection for AI Agents: Stopping Micro-Loops
- Loop Strategy Spectrum: Accumulated vs Fresh Context
- Managing Agent Skills from the GitHub CLI with gh skill
- Measuring GEO Performance for AI Search Visibility
- Measuring the Verification Tax on Agent Output
- Mise en Place for Agentic Coding
- Model Deprecation Lifecycle for Agent Workloads
- Model a Single Agent Turn as Many Inference and Tool-Call
- Model-ID-as-Dependency: Migration Protocol for Deprecation Churn
- Monitor or Wait: The Supervision Choice During Agent Execution
- Monolith-to-Sub-Agents Refactor: Five Lessons from a Brittle Prototype
- Monorepo Skill and Agent Discovery: Hierarchical Configuration
- Multi-Agent RAG for Spec-to-Test Automation
- Multi-Model Plan Synthesis for System Architecture
- Multi-Repo and No-Repo Coding Agent Automation Templates
- Natural-Language Customization Bootstrap
- Natural-language git
- Notebook-Documented Automation for Repeat Operational Work
- Observability Feedback Loop: A 7-Step Debug Runbook for Agents
- Offline Evaluation as an Integration Test for LLM Features
- One-Click CI Auto-Fix: Human-Triggered Cloud-Agent Remediation for Failing GitHub Actions
- Oracle-Based Task Decomposition for AI Agent Development
- Oracle-Gated Delegation Beyond Your Domain Expertise
- Orchestrator-Worker Pattern for AI Agent Development
- PM on the AI Exponential
- PR Description Style as a Lever for Agent PR Merge Rates
- PR Scope Creep as a Human Review Bottleneck
- PR-Subscribed Agent Ownership: The Agent That Opened the PR Drives It to Green
- Parallel Agent Sessions Shift the Bottleneck from Writing
- Parallel Polyglot Ports as a Spec-Ambiguity Oracle
- Pattern Selection Map
- Per-Change Deploy Monitors: Report the Verdict, Don't Act on It
- Per-Task Agent Routing Across Coding Harnesses
- Per-Task Verification Budget: Size the Task to Fit the Check
- Permutation Frameworks for Batch Code Generation
- Persona-as-Code: Defining Agent Roles as Structured Docs
- Phase-Specific Context Assembly for AI Agent Development
- Plan files as resumable artifacts
- Plan mode for knowledge artifacts
- Polya Small-Steps: Using AI to Think Better, Not Think Less
- Porting a Coding-Agent Harness Beyond Engineering
- PostToolUse Hook for BSD/GNU CLI Incompatibilities
- Pre-Execution Codebase Exploration for AI Coding Agents
- Prebuilt Agent Monitoring Dashboard
- Premature Completion: Agents That Declare Success Too Early
- Prescribing TDD Inside the Agent Loop (Process Theater)
- Prior Dominance Over Feedback in Agent Optimization Loops
- Process Amplification: Scaling Human Work with Agents
- Product-as-IDE: When the Application Becomes the Development
- Programmatic Cloud-Agent Dispatch via REST API and Webhooks
- Prompt Chaining: Sequential LLM Calls for Agent Workflows
- Prompt Governance via PRs: Reviewable AI Behavior
- Prompt-Rewrite Discipline on Cross-Generation Model Migration
- Proprietary-to-Open-Standard Tool Migration (Copilot Extensions to MCP)
- Prototype Before Optimizing: Establish Quality Baselines Before Token Constraints
- Public-Channel Agent Work as Lehrwerkstatt for Team Learning
- QA Session to Issues Pipeline for AI Agent Development
- Re-Run an Agent's Speed-Up Claim Before Merging
- Red-Green-Refactor with Agents: Tests as the Spec
- Reducing Fixed CI Overhead Before Adding Shards
- Refactoring Runaway: Tangled Refactorings in Agent Patches
- Remote Agent Host Sessions over SSH and Dev Tunnels
- Remote Session Control for Local CLI Agents
- Repository Bootstrap Checklist: Wiring Agent Support
- Review-Feedback-to-Rule Loop: Promoting Recurring PR Comments into Harness Rules
- Reviewer's Playbook for Agent-Authored Pull Requests
- Rigor Relocation: Engineering Discipline with AI Agents
- Routing Dependency Updates to Repair Agents by Budget
- Run-Status vs Task-Status Confusion in Autonomous Agent Runs
- Runbooks as Agent Instructions: Agent-Followable Ops
- Runnable Documentation as Agent Verification
- SDLC-Phase Skill Taxonomy: Full-Lifecycle Skill Libraries
- SEO vs GEO — How Signals and Metrics Differ
- Scheduled Instruction File Fact-Checker for Accuracy
- Schema and Structured Data for GEO — AI Citation Guide
- Seamless Background-to-Foreground Handoff
- Self-Healing Production Agent: Automated Regression Detection and Autofix PR
- Self-Reporting Loops: Autonomous Routines That File Their Own Backlog
- Semantic Issue Search from Chat vs Query Syntax
- Session Initialization Ritual: How Agents Orient Themselves
- Session Scheduling with Loop and Cron in Claude Code
- Shadow Tech Debt Created by Autonomous AI Agent Commits
- Simulation and Replay Testing for Agent Verification
- Single-Branch Git for Agent Swarms: A Trade-Off Pattern
- Single-CLI Agent Platform: Create to Production in One CLI
- Sizing an Agent Migration Fan-Out by Diff Uniformity
- Skeleton Projects as Agent Scaffolding
- Skill Library Refinement Loops: Organizational Feedback for Shared Skills
- Spec Complexity Displacement: When Specs Become Code
- Spec-Anchored Drift-Gated Architecture (Spec Growth Engine)
- Spec-Driven Development with Spec Kit
- Specification Portability Across Coding Agents
- Specification-First Convergence Without a Test Oracle
- Stacked Agent Sessions on Unmerged Feature Branches
- Stacking Outer Loops Around the Agent
- Staged Literal Porting with a Per-Stage Numeric Oracle
- Staggered Agent Launch: Preventing Thundering-Herd in Swarms
- Stakeholder Trust Through Evals and Observability
- Steering Running Agents: Mid-Run Redirection and Follow-Ups
- Swarm Migration Pattern
- System-Level Optimization Pipeline
- TDD Interaction Models: Throughput Versus Test Quality
- Team OS: Coding-Agent Repo as Cross-Functional Team Brain
- Team Onboarding for AI Agent Workflows and Adoption
- Terminal Tools for Agents: send_to_terminal and Background Interaction
- The 7 Phases of AI-Assisted Feature Development
- The AI Development Maturity Model: From Skeptic to Agentic
- The Effortless AI Fallacy for AI Agent Development
- The Error-Class Governance Loop for Instruction Libraries
- The Eval-First Development Loop for AI Agent Features
- The First Edit Predicts Whether an AI Completion Survives
- The LLM Laziness Deficit Fallacy: Restraint Comes From Harness, Not Instruction
- The Plan-First Loop: Always Design Before Writing Code
- The Research-Plan-Implement Pattern
- The Software Factory Model: Industrializing Agent Loops
- The Think Tool: Mid-Stream Reasoning for AI Agents
- Three Reasoning Spaces: Plan-Bead-Code Phase Gates
- Throwaway-Prototype Skill: Build to Discard, Keep Only the Answer
- Tiled Agent Layout: Supervising Parallel Agents Through Dedicated Panes
- Topical Authority — Entity Coverage for AI Citation
- Trajectory Logging via Progress Files and Git History
- Trigger-to-Function Architecture for Unprompted Codebase Maintenance
- Turn-Level Context Decisions for AI Coding Sessions
- Using the Agent to Analyze Its Own Evaluation Transcripts
- Velocity-Quality Asymmetry: Why AI Speed Gains Fade
- Verification Capacity Saturation: Three Levers, One Default
- Verification-Centric Development for AI-Generated Code
- Verifying Agent Changes in the Copilot App
- Visible Thinking in AI-Assisted Development
- Voting / Ensemble Pattern for AI Agent Development
- WIP=1 and Little's Law: Kanban Throughput Theory for Agent Task Design
- Web Search Agent Loop: Iterative Research Patterns
- What is GEO — Generative Engine Optimization Defined
- Whole-Codebase Visibility as a Migration Prerequisite
- Workflows for AI Agent Development
- Worktree Isolation: Parallel Agent Sessions in Safe Sandboxes
- llms.txt: Full Specification, Adoption, and Limitations