Agent Evals for Financial Services
Agent evals for financial services are the structured tests, graders and retained records that a bank, broker-dealer or insurer uses to establish that an AI agent, a language model that plans and calls tools across multiple steps, does what it is authorised to do, reaches correct outcomes and leaves an inspectable trail. The topic matters more in finance than in most sectors because of a gap in the rulebook. The revised US model risk guidance explicitly leaves agents out, conduct regulators say existing duties still apply, and firms are left to build the evidence themselves. As of October 2026 no bank has published its agent evaluation suite or results, so public knowledge rests on regulator texts and external benchmarks.
The scope gap in model risk management
On 17 April 2026 the Federal Reserve, OCC and FDIC issued SR 26-2, replacing SR 11-7 as the US supervisory guidance on model risk management. Its definition of a model is a complex quantitative method that applies statistical, economic or financial theories to turn input data into quantitative estimates. A footnote then states: "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." The same footnote adds that a banking organisation's risk management and governance practices should guide the controls for tools not covered.
The practical effect is that the familiar validation template of conceptual soundness, outcomes analysis and ongoing monitoring does not formally apply to an agent, yet nothing replaces it. Most of its ideas transfer by analogy: independent "effective challenge" by people with expertise and objectivity, monitoring after approval, and validation of vendor components even when the internals are proprietary. An agent eval programme is in effect how a firm supplies that analogue, with a fixed task suite standing in for a validation dataset.
What conduct regulators have said
FINRA's 2026 Annual Regulatory Oversight Report added a discussion of AI agents to its generative AI section. It lists risks that read like an evaluation specification: agents acting autonomously without human validation and approval, agents acting beyond the user's actual or intended scope and authority, multi-step reasoning that makes outcomes difficult to trace or explain, mishandling of sensitive data, insufficient domain knowledge, and poorly designed reward functions that lead to decisions which hurt investors. Among the practices it describes are robust testing to understand capabilities and limitations, ongoing monitoring of prompts, responses and outputs, storing prompt and output logs, and deciding where human-in-the-loop oversight sits. It ties these to Rule 3110, under which a firm's supervisory system must be reasonably designed for its business.
The UK Financial Conduct Authority takes a similar line without new rules. Its published approach states that it does not plan to introduce extra regulations for AI and will rely on existing frameworks, naming the Consumer Duty and the Senior Managers and Certification Regime, and it runs an AI Lab to support firms. In the EU, Annex III of the AI Act classifies AI used to assess the creditworthiness of natural persons as high-risk, and Articles 12 and 14 require such systems to record events automatically over their lifetime and to be designed so that people can oversee, override and halt them. The Article 12 page, read in October 2026, gives 2 December 2027 as the application date for Annex III systems.
What an evaluation has to cover
Each FINRA risk maps to a measurable property. Scope and authority are tested with tasks in which the correct behaviour is to refuse or escalate, such as an instruction to move funds above a limit or to act for a client the user does not cover. Outcome correctness is graded on the final state of the system, whether the right order, ledger entry or disclosure exists, and not on the agent's own summary of what it did. Consistency needs repeated trials: Anthropic's engineering guidance on agent evals (January 2026) distinguishes pass@k, the chance of at least one success in k attempts, from pass^k, the chance that all k succeed, and it is the second that matters for a process that runs thousands of times a day. Communications with customers need rubric grading for fairness and balance, typically by a model grader calibrated against compliance reviewers.
Audit trails
Traceability is both an eval input and a regulatory deliverable. The complete transcript of a run, with model version, prompts, tool calls, intermediate results and the human approvals given, is what graders score and what a supervisor will ask for after an incident. FINRA's reference to storing prompt and output logs and the AI Act's logging requirement point the same way. The design consequence is that the trace store has to be built to recordkeeping standards from the start, with retention and tamper-evidence, and that any change to prompt, tool or model version should trigger the regression suite before release.
Public benchmarks and their limits
External benchmarks set expectations for raw capability. FinanceBench (Islam et al., Patronus AI, November 2023) found that GPT-4-Turbo with a retrieval system incorrectly answered or refused 81% of questions on a 150-case sample of open-book questions about public companies. The Finance Agent Benchmark (Bigeard et al., 2025), run by Vals AI, uses 537 expert-authored research questions over SEC filings in nine task categories; in the paper the best model, OpenAI's o3, reached 46.8% accuracy at an average cost of $3.79 per query. The live leaderboard for version 1.1, updated 4 June 2026, shows Claude Opus 4.7 leading at 64.37%, as reported by the benchmark operator. Progress is real, and a one-in-three error rate on analyst-level research tasks remains far from unsupervised use. These benchmarks also measure research accuracy only. None tests authority limits, suitability or recordkeeping, which a firm has to cover with its own tasks.
Applications & Use Cases
Pre-deployment validation packs
A fixed task suite with graded outcomes and transcripts, assembled as the agent-era counterpart of a model validation report for internal risk committees.
Scope-and-authority testing
Adversarial tasks that check the agent refuses or escalates when asked to exceed limits, act outside a mandate or disclose restricted data.
Change-control regression gates
Re-running the suite on every prompt, tool or model update so that behaviour drift is caught before production.
Customer-communication review
Rubric grading of agent-drafted messages against fair-and-balanced communication standards, with human sampling to calibrate the grader.
Research-agent accuracy tracking
Using public finance benchmarks as an external reference point alongside internal document sets.
Supervisory evidence
Retained traces and eval results that support a firm's account of its supervisory system during examinations and incident reviews.
Key Players
- Federal Reserve, OCC and FDIC — Issued SR 26-2 in April 2026, which excludes generative and agentic AI from the model risk guidance while leaving governance to firms.
- FINRA — US broker-dealer regulator; its 2026 oversight report sets out agent-specific risks and testing, monitoring and logging practices.
- Financial Conduct Authority — UK conduct regulator; applies existing rules to AI and operates an AI Lab for firms.
- Vals AI — Operates the Finance Agent Benchmark and its public leaderboard; the task taxonomy was developed in consultation with experts from banks, hedge funds and private equity firms.
- Patronus AI — Maintains FinanceBench and publishes a 150-example open sample.
- Anthropic — Published engineering guidance on agent evals that defines tasks, trials, graders and consistency metrics used in this field.
Challenges & Considerations
- No settled supervisory standard — With agents outside SR 26-2, firms cannot point to an agreed validation template; expectations are being set through examinations and oversight reports instead.
- Grading judgement-heavy outputs — Suitability and fair-communication standards are not reducible to string matching. Model graders need calibration against compliance staff and periodic re-checking.
- Non-determinism — An agent that passes once may fail on the next run. Evidence has to be expressed as rates over repeated trials, which multiplies cost.
- Trace retention and privacy — Full transcripts contain client data. Retaining them for audit conflicts with data-minimisation duties unless access and redaction are designed in.
- Benchmark relevance — Public finance benchmarks test research over filings. They do not measure authority limits, transaction correctness or conduct, so high scores do not transfer.
Further Reading
- SR 26-2: Revised Guidance on Model Risk Management (letter and attached supervisory guidance) — Federal Reserve, OCC and FDIC, April 2026
- 2026 FINRA Annual Regulatory Oversight Report: GenAI section — FINRA, 2026 report
- AI and the FCA: our approach — Financial Conduct Authority, read October 2026
- EU AI Act, Annex III: High-Risk AI Systems — EU Artificial Intelligence Act explorer, read October 2026
- EU AI Act, Article 12: Record-Keeping — EU Artificial Intelligence Act explorer, read October 2026
- EU AI Act, Article 14: Human Oversight — EU Artificial Intelligence Act explorer, read October 2026
- Demystifying evals for AI agents — Anthropic Engineering, January 2026
- Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks (Bigeard et al.) — arXiv, 2025
- Finance Agent benchmark leaderboard (v1.1) — Vals AI, updated June 2026
- FinanceBench: A New Benchmark for Financial Question Answering (Islam et al.) — arXiv, November 2023