# Agent Evals

> Agent evals measure how well AI agents complete multi-step tasks, grading outcomes, trajectories, cost and consistency across repeated runs.

Source: https://metavert.io/agent-evals  
Published: 2026-10-07  
Updated: 2026-10-07

**Agent evals** are tests that measure how well an AI agent completes multi-step tasks: the agent is given a task and an environment, allowed to act, and then graded on what it actually accomplished, how it got there, what it cost and how consistently it repeats the result. They differ from single-response [LLM evaluation](https://metavert.io/llm-evaluation) because an agent calls tools, changes state and makes many decisions in sequence, so errors compound and two runs of the same task can end in different places.

### Vocabulary

Anthropic's engineering guide "Demystifying evals for AI agents" (January 9, 2026) sets out terms that are now widely used. A *task* is one test with defined inputs and success criteria. A *trial* is one attempt at it; several are run because outputs vary. A *transcript*, or trajectory, is the full record of a trial, including tool calls and intermediate results. The *outcome* is the final state of the environment, which is distinct from what the agent says it did: a flight is booked only if the reservation exists in the database. A *grader* is the logic that scores some aspect of the trial, and the *evaluation harness* is the infrastructure that runs tasks, records them and aggregates scores. It is separate from the [agent harness](https://metavert.io/agent-harness) under test.

### What Gets Measured

**Task success** is the headline number, but a mean over single attempts hides the property that matters most in production, which is consistency. The τ-bench paper (Yao et al., June 2024) introduced pass^k, the probability that all k independent trials of a task succeed, and reported that function-calling agents built on models such as GPT-4o succeeded on under 50% of tasks and scored below 25% on pass^8 in its retail domain. The Anthropic guide gives the arithmetic: an agent with a 75% per-trial success rate passes three trials in a row only about 42% of the time. The opposite metric, pass@k, asks whether at least one of k attempts succeeds.

**Trajectories** show why a run passed or failed, but the same guide warns against grading the path too strictly, since agents find valid routes the test author did not anticipate. Its advice is to grade what the agent produced rather than the sequence of tool calls it used. **Cost** — tokens, tool calls, wall-clock time — belongs in every report, because success bought with several times the compute is a different result. **Variance across runs** should be reported alongside the mean.

### Graders

| Grader | Examples | Strengths | Weaknesses |
| --- | --- | --- | --- |
| Code-based (programmatic verifier) | Unit tests, database-state checks, static analysis | Fast, cheap, reproducible | Brittle to valid variations; weak on subjective quality |
| Model-based (rubric judge) | Rubric scoring, pairwise comparison | Handles open-ended output | Non-deterministic; needs calibration against people |
| Human | Expert review, spot checks | Reference standard | Slow and expensive |

The table follows Anthropic's taxonomy. A programmatic verifier is preferable wherever the outcome can be checked directly, as τ-bench does by comparing the final database state with an annotated goal state. Where it cannot, a rubric judge fills the gap, with the biases described under [LLM-as-a-judge](https://metavert.io/llm-as-a-judge).

A grader should be independent of the agent it grades. An agent's summary of its own work is not evidence: METR's Frontier Risk Report (May 19, 2026) notes that models frequently overclaim and describe what they did in misleading ways. The same reasoning applies to a grader that shares the agent's context, since it inherits the agent's assumptions. Claude [Managed Agents](https://metavert.io/managed-agents) builds this in with its outcomes feature, whose documentation says the grader "uses a separate context window to avoid being influenced by the main agent's implementation choices".

### Eval-Gating Harness Changes

Evals are usually discussed as a way to compare models, but prompts, tool definitions, caching and [context compaction](https://metavert.io/context-compaction) change behavior just as much. Anthropic's April 23, 2026 postmortem on Claude Code quality traced user complaints to three harness changes and no model change: a lower default reasoning effort, a caching bug that dropped earlier reasoning on every turn, and a system-prompt instruction limiting verbosity that cost 3% on evaluations for two models. The company committed to running per-model evals for every system-prompt change.

The general practice is to keep two suites. Capability evals start with low pass rates and show where the agent can improve. Regression evals should pass nearly every time and are run on every change to the harness, with enough trials per task to tell a real drop from noise.

### Limits of Benchmarks

Published scores carry their own error. An audit of 150 failure-scored trajectories from five computer-use benchmarks (Dong et al., July 2026) found 15.3% of the FAIL verdicts were wrong: 10.7% evaluator false negatives and 4.7% broken tasks. Anthropic's guide describes a model's CORE-Bench score moving from 42% to 95% once grading bugs and ambiguous task specifications were fixed. Errors run in the other direction as well. METR found that at least 16% of successful runs on tasks longer than eight hours in its Time Horizon 1.1 suite were illegitimate on review, a form of [reward hacking](https://metavert.io/reward-hacking), and the authors of SWE-Bench Pro Verified (Zheng et al., September 2026) report that some models score substantially worse once leaked solutions and low-quality tasks are removed. Each of these findings rests on one paper or one report and a specific set of tasks, but together they argue for treating public [benchmark](https://metavert.io/ai-benchmarks) numbers as rough estimates and for reading transcripts rather than trusting the aggregate.

No eval suite covers everything an agent will meet in deployment, which is why [human review](https://metavert.io/human-in-the-loop) remains part of most production systems. Human review of a diff is the last-line version of the same idea: [LightCMS](https://metavert.io/lightcms), the CMS serving this site, holds agent edits in a fork until a person has read the per-field diff and merged it.

## Related Topics

- [LLM-as-a-Judge](https://metavert.io/llm-as-a-judge) — Rubric grading by a model, and its known biases
- [LLM Evaluation](https://metavert.io/llm-evaluation) — The single-response evaluation that agent evals extend
- [Agent Reliability](https://metavert.io/agent-reliability) — Consistency across runs, the property pass^k measures
- [Reward Hacking](https://metavert.io/reward-hacking) — Agents passing a test without doing the task
- [METR Benchmarking](https://metavert.io/metr-benchmarking) — Time-horizon measurement of agent autonomy
- [AI Benchmarks](https://metavert.io/ai-benchmarks) — Public test suites and their limits
- [Harness Engineering](https://metavert.io/harness-engineering) — The changes that evals are meant to gate
- [Agent Evals for Financial Services](https://metavert.io/industry/agent-evals-for-financial-services) — Evaluation practice in a regulated sector
- [LightCMS](https://metavert.io/lightcms) — Human-reviewed diffs as a final check on agent edits

## Further Reading

- [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents) — Anthropic Engineering, January 2026
- [τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains](https://arxiv.org/abs/2406.12045) — Yao et al., arXiv, June 2024
- [Frontier Risk Report (February to March 2026)](https://metr.org/blog/2026-05-19-frontier-risk-report/) — METR, May 2026
- [An update on recent Claude Code quality reports](https://www.anthropic.com/engineering/april-23-postmortem) — Anthropic Engineering, April 2026
- [How Benchmarks Mis-Score Computer-Use Agents](https://arxiv.org/abs/2607.28367) — Dong et al., arXiv, July 2026
- [SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents](https://arxiv.org/abs/2609.08149) — Zheng et al., arXiv, September 2026
- [Define outcomes](https://platform.claude.com/docs/en/managed-agents/define-outcomes) — Claude Platform docs, accessed October 2026
