Agent Evals for Healthcare

Industry Application
Agent EvalsHealthcare

Agent evals for healthcare are the tests used to determine whether an AI agent, a language model that gathers information, calls tools and takes actions over multiple steps, behaves safely and correctly in clinical or administrative health workflows. They differ from the medical question-answering scores that dominated early coverage in what they inspect: the orders placed, records retrieved and escalations made in an environment, as well as the text produced. The public evidence comes mainly from research benchmarks. As of October 2026, regulator guidance addresses AI-enabled software in general terms and has little that is specific to agents, and no health system has published a full evaluation suite for a deployed clinical agent.

Why exam-style scores do not transfer

AgentClinic (Schmidgall et al., first posted May 2024, latest revision May 2025) made the point empirically. It recasts diagnosis as a sequential task in which the agent must question a simulated patient, request tests and work with incomplete information, across nine specialties and seven languages. The authors report that solving MedQA problems in this sequential format is considerably harder, with diagnostic accuracy in some cases dropping below a tenth of the original. A model's examination score therefore says little about whether it will ask the right questions before reaching a conclusion.

HealthBench (Arora et al., OpenAI, May 2025) addresses a different weakness, the grading of open-ended responses. It contains 5,000 multi-turn conversations with individuals or healthcare professionals, scored against 48,562 rubric criteria written by physicians. Reported scores rose from 16% for GPT-3.5 Turbo to 32% for GPT-4o and 60% for o3, and on a harder variant of the benchmark the top score was 32%. HealthBench measures conversational quality, not actions taken in a system, so it is a necessary component of an agent evaluation and not the whole of one.

Benchmarks that test actions in clinical systems

MedAgentBench, from Stanford researchers and published in NEJM AI in 2025, is the closest public analogue to a real deployment test. It provides 300 clinically derived tasks in 10 categories written by physicians, the records of 100 realistic patients comprising more than 700,000 data elements, and a FHIR-compliant interactive environment of the kind used by modern electronic health record systems. The best model at publication, Claude 3.5 Sonnet v2, completed 69.67% of tasks, with large differences between categories. The design matters as much as the score. Success is judged on the resulting state of the record system, which is the property an agent eval needs and a question-answering test cannot provide.

These are single benchmarks built on synthetic or simulated patients. They do not measure performance on a specific hospital's data, order sets or local protocols, and scores from 2025 describe models that have since been superseded. Their lasting value is as templates for a local suite.

What regulators have published

In the United States, FDA's draft guidance on AI-enabled device software functions (January 2025, docket FDA-2024-D-4488) recommends managing risk across the total product life cycle and sets out what marketing submissions should contain; it remains a draft. The Clinical Decision Support Software guidance (final, January 2026) clarifies which decision-support functions fall outside the device definition under section 520(o)(1)(E) of the FD&C Act, a boundary that matters for agents that act as well as advise. FDA's public list of AI-enabled devices states that the agency will explore methods to identify and tag devices incorporating foundation models, from large language models to multimodal architectures, an acknowledgement that the category is not yet tracked systematically. FDA also lists a discussion paper on regulating generative AI-enabled devices.

The World Health Organization's guidance on the ethics and governance of large multi-modal models covers their use in healthcare, research, public health and drug development. In the UK, the MHRA's AI Airlock is a regulatory sandbox for AI medical devices: its pilot ran until April 2025 with four products, phase 2 concluded in May 2026 with seven innovators, and phase 3 opened on 6 October 2026 with a focus on post-market surveillance, with programme reports published in October 2025 and July 2026. In the EU, Article 14 of the AI Act requires high-risk systems to be designed so that people can understand their limits, remain aware of automation bias, override outputs and stop the system; the Article 14 page, read in October 2026, gives 2 August 2028 as the application date for systems that are high-risk under Annex I, the route that covers regulated products such as medical devices.

Validation and human oversight in practice

Taken together, the benchmarks and guidance imply a layered evaluation. Outcome grading on a sandboxed copy of the record system checks that the agent did the right thing. Physician-written rubrics check that what it said was accurate, complete and appropriately cautious. Repeated trials estimate consistency, since Anthropic's engineering guidance on agent evals (January 2026) notes that the probability of succeeding on every one of k runs falls quickly as k grows, and a clinical process needs that number, not the best-of-k figure. Escalation tests check that the agent hands off when information is missing or risk is high.

Human oversight then has to be evaluated as part of the system and not assumed. A clinician who approves every agent suggestion is not providing oversight, which is why Article 14 names automation bias explicitly. Useful measures include how often reviewers catch seeded errors and how review time changes with volume. Offline evals also cannot replace prospective clinical evaluation. A benchmark success rate is evidence about capability, while safety in use depends on the workflow, the population and monitoring after deployment, which is the direction the MHRA's phase 3 and FDA's life-cycle framing both take.

Applications & Use Cases

Record-system task suites

Tasks executed against a sandboxed FHIR environment and graded on the resulting orders, notes and retrievals, following the MedAgentBench design.

Physician-rubric grading of conversations

Open-ended patient or clinician dialogues scored against clinician-authored criteria, with a model grader checked against physician judgement.

Sequential diagnostic simulation

Simulated patient encounters that test information gathering under uncertainty instead of answer selection.

Escalation and refusal testing

Cases where the safe action is to defer to a clinician, used to measure under-escalation and over-escalation rates.

Reviewer-effectiveness studies

Seeding known errors into agent output to measure how reliably human reviewers detect them.

Post-deployment regression monitoring

Re-running a frozen suite after each model, prompt or tool change, and sampling live traces for clinical review.

Key Players

  • Stanford University — Researchers there built MedAgentBench, the FHIR-based agent benchmark published in NEJM AI in 2025.
  • OpenAI — Released HealthBench, a physician-rubric benchmark of 5,000 health conversations, with reference code under an MIT licence.
  • AgentClinic — Open multimodal benchmark for sequential clinical decision-making by Schmidgall and colleagues.
  • US Food and Drug Administration — Publishes draft life-cycle guidance for AI-enabled device software, the clinical decision support guidance and the AI-enabled device list.
  • MHRA — UK medicines and devices regulator running the AI Airlock sandbox, now in its third phase.
  • World Health Organization — Issued guidance on the ethics and governance of large multi-modal models in health.

Challenges & Considerations

  • Simulation is not clinical validation — Benchmarks use synthetic patients and generic workflows. Local data, order sets and populations can change results, and prospective evaluation is still required.
  • Unsettled regulatory treatment — FDA's AI-specific life-cycle guidance is still a draft and foundation-model devices are not yet systematically tagged; an agent's regulatory status depends on intended use.
  • Grader validity — Model graders scale rubric scoring but must be shown to agree with physicians, and rubrics age as clinical guidance changes.
  • Automation bias — Oversight that consists of routine approval offers little protection. The AI Act names the risk, but measuring reviewer vigilance is rarely part of evaluation plans.
  • Patient data in traces — Complete transcripts are needed for grading and incident review and contain protected health information, so storage, access and retention need the same controls as the record itself.