Skip to content

Verification

How to measure agent output quality, design evaluation suites, and use evals to drive development.

Measuring Quality

Behavioral Testing

Regression Testing

Eval-Driven Development

Review Techniques

  • Five-Pass Blunder Hunt — Run the same critique prompt five times in sequence on a plan or spec; each pass normalizes the issues it finds, forcing later passes deeper into structural and logical problems
  • Pre-Completion Checklists — Block agent completion signals with a mandatory verification sequence
  • Golden Journeys: Restartability as a First-Class Verification Primitive — Name a small set of end-to-end paths with explicit failure signals per step and gate completion on the system restarting cleanly afterward
  • Test-Driven Intent Clarification — Use AI-generated tests to surface specification ambiguity before code review — validate tests instead of code to clarify intent with lower cognitive cost
  • Probing Unstated Constraints in Generated Code (Intent Violation Rate) — A stated suite only checks the constraints your prompt named; two frontier models passed over 92% of stated tests while violating a hidden constraint in more than half of problems, and the violations are deterministic per problem rather than sampled
  • Unstated-Contract Bugs: Sort Tickets by Information Gap — Agents fixed two hard bugs 16/16 and failed an easy one 12/12 because the correct answer lived in user data rather than the repository, so triage tickets by where the information sits and enrich the ticket before the agent starts
  • Source-Grounded Test Plan with Pre-Action Assertion Annotation — Before a UI-driving agent verifies its change, have it write a source-read test plan and annotate each step's expected behavior upfront so it cannot rationalize an unexpected result as a pass
  • Agent-Recorded Video Demos as a Verification Artifact — The agent drives the running app and records a screencast as proof-of-work for human PR reviewers — a visual modality that complements tests where they are weakest, under stated conditions
  • Agent-Generated Verification Reports: A Structured Round-Trip for Human Review — Each parallel agent emits a per-sub-task report with the request quoted verbatim and exact test steps, and the human's verified or not-fixed verdict routes back as the completion signal
  • Spec-Derived Execution as a Correctness Oracle — Judge candidate code against a natural-language spec by deriving inputs from the spec, executing them, and grading the I/O pairs — ground the LLM judge in real execution traces instead of asking it to reason over the code
  • Specification-Grounded Test Writing — Give an agent's test writer the specification as enumerated rules so it catches the missing-validation and boundary bugs its own tests otherwise pass — grounding in intent, not test count, drives a +38pp correct-code gain on spec-completeness defects
  • Deriving a Specification From Buggy Code Before Generating Tests — When no written specification exists and the code under test may be wrong, have the model write a behavioral docstring first and generate tests from that with the implementation removed — recovers part of the lost bug-detection power, not all of it
  • Spec-Driven Test Generation: Contract Coverage Is the Lever — Having the agent document pre-conditions, post-conditions and undefined behaviors before writing tests lifted bug detection from 53.4% to 63.2% across five runs, but only when the extracted contract named the condition the bug violated — 54.9% detection with that coverage against 19.4% without
  • Natural-Language Documentation as a Code-Review Intermediate (Verifiable Literate Programming) — Translate LLM code into a deterministic natural-language documentation layer and review that prose, so intent-code mismatches surface where a human can judge them
  • Audit-Budget Allocation for Agent Fleets — Screen an agent's self-reported confidence for discrimination before ranking a review queue by it; past a mimicry threshold the ranked queue does worse than random, and a larger audit budget flips first
  • Evidence-Chain Run Logs: Bracket the Reported Symptom — Pair every agent tool call with its actual result and bracket the change with one machine-readable measurement of the reported symptom, taken identically before and after — a gate, never a signal the agent iterates against
  • Evidence-Bundled Agent PRs: Sizing the Reviewer's Effort — Attach the reproduction, failing probe, and executed test the agent already ran so a reviewer picks a depth per change; the artifacts carry the value, and the agent's own risk grade carries none
  • Claim-to-Evidence Trace Graphs for Auditing Agent Runs — Rebuild a finished session as a layered graph with typed edges so a reviewer traverses backward from a claim to its artifacts and checks; the record layer is deterministic, the graph above it is inferred
  • Post-Merge Fix Signals for Agent Merges — Agent merges draw a verified follow-up fix at 1.62 times the odds of human merges in the same repositories; watch the first week and rank by the commits it took to land, because review duration carries no within-repository effect
  • Name the Check That Passed Before Accepting AI Code — Across 506 coded accounts of AI use in research code, evaluation confidence tracked trust in the tool (r=.33) rather than the checks actually run (r=.05), so gate acceptance on a recorded check instead
  • Transcript-Measured Review Coverage — Agents skipped at least one in-scope file in 67.9% of 1,140 review runs and 80.4% of those runs did not disclose it; derive coverage from the tool-call trace and flag the mismatch with what the report claims
  • Verify Agent Diagnoses and Fix Proposals Before Acting — In a Claude-built codebase, design-fix proposals scored 79.4% accurate against 94.3% for checkable facts, narrowing to 93.4% once rejected proposals are excluded; run the check before you adopt a diagnosis or proposal, since it predicts rather than reports
  • Task Category as the Security Review Routing Key — Generated-code security findings varied three to six times more by task category than by which tool wrote the code, so route review depth by task and grade model choice separately

Rubric Design

Guardrails

Tooling