Independent evaluation for financial AI
The financial services industry is in the early stages of deploying AI on tasks that have real monetary consequences - verifying transaction fees, synthesising audit evidence, running financial models. The quality of these deployments depends in part on having honest, rigorous benchmarks for what frontier models can and cannot do.
FinanceBenchmark exists to be that resource. We aggregate scores from publicly available financial AI benchmarks, conduct original evaluations using our own benchmark suite, and publish everything with full attribution and open methodology. Our goal is not to rank models for marketing purposes - it is to give practitioners and researchers an accurate picture of the state of financial AI reasoning.
Five domains. Eight benchmarks.
We cover five financial reasoning domains: Verification (deterministic fee and compliance checking), Document QA (retrieval and comprehension over financial filings), Forensic Reasoning (cross-document inconsistency detection), Numerical Reasoning (multi-step arithmetic over structured data), and Agent Tasks (end-to-end financial research workflows).
The benchmarks range from PAVE - our own, designed specifically to test deterministic verification at scale - to established academic benchmarks like FinQA and FinanceBench. Tier 1 benchmarks carry the most weight in the FB Composite Score because they test the hardest reasoning. We add new benchmarks quarterly when they meet our ground-truth and open-data requirements.
Why financial services specifically? Because the reliability bar is higher. A model that retrieves the wrong revenue figure from an earnings call is an inconvenience. A model that misclassifies an interchange fee category or misses a discrepancy in a waterfall structure is a financial error. The tasks we select reflect that asymmetry.
Founded by Avae. Independent in operation.
FinanceBenchmark was founded by Avae, a continuous financial verification engine. Avae builds adversarial consensus verification architecture for LLM agents in financial forensics. The PAVE benchmark - our flagship evaluation - was developed as part of Avae's research into interchange fee verification.
FinanceBenchmark does not have commercial agreements with any of the organisations whose models it evaluates. We do not accept payment for placement, evaluation priority, or promotional content. Our operating costs are funded by Avae. We do not display advertising.
How we got here
Get in touch
We respond to all enquiries, typically within two business days. For score submissions and evaluation requests, use the submit page.