2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks
About FinanceBenchmark

The evaluation standard for
AI in financial services

FinanceBenchmark publishes independent rankings of frontier AI models on financial reasoning tasks - from interchange fee verification to forensic audit discovery. We don't run inference on proprietary weights, accept payment for placement, or estimate scores we haven't verified.

9
Models tracked
8
Benchmarks
13,578
Tasks evaluated
2026
Founded
What we do

Independent evaluation for financial AI

The financial services industry is in the early stages of deploying AI on tasks that have real monetary consequences - verifying transaction fees, synthesising audit evidence, running financial models. The quality of these deployments depends in part on having honest, rigorous benchmarks for what frontier models can and cannot do.

FinanceBenchmark exists to be that resource. We aggregate scores from publicly available financial AI benchmarks, conduct original evaluations using our own benchmark suite, and publish everything with full attribution and open methodology. Our goal is not to rank models for marketing purposes - it is to give practitioners and researchers an accurate picture of the state of financial AI reasoning.

01
Ground truth only
Every benchmark we track has deterministic answers. We do not score tasks that rely on human preference or subjective judgement.
02
Source every score
Every score on the leaderboard shows its source - arXiv paper ID or FinanceBenchmark evaluation. No estimates, no aggregator re-scores.
03
Open methodology
The PAVE benchmark dataset, evaluation harness, and ground-truth answers are publicly available on GitHub. All scores are independently verifiable.
04
Disclose conflicts
Our founding organisation participates in PAVE. We disclose this on every page where Avae 2.0 appears, not just on the methodology page.
What we evaluate

Five domains. Eight benchmarks.

We cover five financial reasoning domains: Verification (deterministic fee and compliance checking), Document QA (retrieval and comprehension over financial filings), Forensic Reasoning (cross-document inconsistency detection), Numerical Reasoning (multi-step arithmetic over structured data), and Agent Tasks (end-to-end financial research workflows).

The benchmarks range from PAVE - our own, designed specifically to test deterministic verification at scale - to established academic benchmarks like FinQA and FinanceBench. Tier 1 benchmarks carry the most weight in the FB Composite Score because they test the hardest reasoning. We add new benchmarks quarterly when they meet our ground-truth and open-data requirements.

Why financial services specifically? Because the reliability bar is higher. A model that retrieves the wrong revenue figure from an earnings call is an inconvenience. A model that misclassifies an interchange fee category or misses a discrepancy in a waterfall structure is a financial error. The tasks we select reflect that asymmetry.

Relationship to Avae

Founded by Avae. Independent in operation.

FinanceBenchmark was founded by Avae, a continuous financial verification engine. Avae builds adversarial consensus verification architecture for LLM agents in financial forensics. The PAVE benchmark - our flagship evaluation - was developed as part of Avae's research into interchange fee verification.

Avae
Founding organisation · avae.ai
Avae is a continuous financial verification engine based in Riyadh, Saudi Arabia. It builds adversarial consensus verification architecture for LLM agents in financial forensics - the same domain FinanceBenchmark evaluates.
avae.ai →

FinanceBenchmark does not have commercial agreements with any of the organisations whose models it evaluates. We do not accept payment for placement, evaluation priority, or promotional content. Our operating costs are funded by Avae. We do not display advertising.

Timeline

How we got here

2025 Q3
PAVE specification published
Avae publishes the PAVE Interchange Benchmark specification as part of its arXiv paper on adversarial consensus verification for financial forensics.
2025 Q4
Baseline evaluations begin
First evaluation run across 5 frontier models on PAVE v1.0. Results indicate a significant gap between deterministic verification and LLM performance on Level 3–4 tasks.
2026 Q1
FinanceBenchmark.ai launched
Public launch with 8 benchmarks, 9 models, the FB Composite Score, and full methodology documentation including conflict of interest disclosure.
2026 Q2
Open submission launched
Any organisation can submit scores or request certified evaluations. Certified evaluations free for research institutions and open-source models.
Now
9 models · 8 benchmarks · 13,578 tasks
Leaderboard updated continuously. New benchmarks reviewed quarterly. PAVE v1.2 released with 50,000 transaction records across four complexity levels.
Contact

Get in touch

We respond to all enquiries, typically within two business days. For score submissions and evaluation requests, use the submit page.

Research & methodology
Academic enquiries
Methodology questions, research collaborations, requests to cite our benchmark data, or enquiries about our evaluation protocol.
research@financebenchmark.ai
Scores & evaluation
Submit & evaluate
Submit a score from your own evaluation run, or request a certified evaluation of your system against our full benchmark suite.
submit@financebenchmark.ai