What FinanceBenchmark is
FinanceBenchmark is an independent evaluation standard for AI systems in financial services. We aggregate scores from publicly available financial AI benchmarks and conduct original evaluations using our own benchmark, PAVE. Every score on the leaderboard is attributed to its original source.
We do not accept payment from AI companies for placement or preferential treatment. We do not run inference on proprietary model weights — all evaluations use publicly accessible API endpoints under standard commercial terms. We do not reproduce scores from other leaderboard websites; we cite only the original research paper.
FinanceBenchmark covers five financial reasoning domains: Verification (deterministic fee and compliance checking), Document QA (question answering over financial filings), Forensic Reasoning (multi-document evidence synthesis), Numerical Reasoning (multi-step arithmetic over structured data), and Agent Tasks (end-to-end financial research workflows).
The FB Composite Score
The FB Composite Score is a weighted average of a model's performance across all benchmarks it has been evaluated on. It is designed to reflect the relative difficulty and financial relevance of different benchmark types, giving more weight to tasks that require deeper financial reasoning.
Tier weights
Every benchmark tracked by FinanceBenchmark is assigned to one of three tiers based on the complexity of reasoning required and the directness of financial relevance. The tier determines the benchmark's weight in the FB Composite Score calculation.
| Tier | Weight | Description | Benchmarks |
|---|---|---|---|
| Tier 1 | 30 per benchmark | Benchmarks testing reasoning over raw structured financial data requiring code execution, multi-file joins, or adversarial verification. Tasks have exact ground-truth answers derived from formal specifications. | PAVE Interchange Fee Benchmark Finance Agent Benchmark |
| Tier 2 | 20 per benchmark | Benchmarks testing numerical reasoning over financial documents. Tasks require multi-step arithmetic and evidence retrieval from earnings reports and financial statements. | FinanceBench FinQA ConvFinQA |
| Tier 3 | 10 per benchmark | Broad financial knowledge and NLP benchmarks covering diverse tasks including sentiment analysis, entity recognition, summarisation, and general financial reasoning. | FinBen MultiHiertt FinAuditing |
Formula
The FB Score is computed as a weighted sum of a model's scores on each benchmark it has been evaluated on, divided by the total weight of those benchmarks. This normalisation ensures that models evaluated on fewer benchmarks are not unfairly penalised or rewarded simply for coverage.
Incomplete evaluations
Models evaluated on fewer than three benchmarks display an asterisk (*) next to their FB Composite Score on the leaderboard. This signals to readers that the composite score is based on limited data and may not be representative of the model's full financial reasoning capability.
A model evaluated on only the PAVE benchmark, for example, will have an FB Score equal to its PAVE score — but this does not mean it performs equally well on document QA or numerical reasoning tasks. Readers should treat single-benchmark scores as indicative, not comprehensive.
Data sources & attribution
Every score on the FinanceBenchmark leaderboard shows its source. We source scores from two types of origin: published arXiv papers by the original benchmark authors, and original evaluations conducted by FinanceBenchmark using our open evaluation harness.
We do not reproduce scores from other leaderboard websites, aggregators, or blog posts. We do not interpolate or estimate scores. If a model has not been evaluated on a benchmark by the original authors or by FinanceBenchmark, the cell shows — and remains — rather than a guess.
| Benchmark | Source type | Attribution | Scores run by |
|---|---|---|---|
| PAVE | FB ORIGINAL | FinanceBenchmark / Avae, 2026 | FinanceBenchmark |
| FinanceBench | ARXIV | arXiv:2311.11944 | FinanceBenchmark |
| FinBen | ARXIV | arXiv:2402.12659 | FinanceBenchmark |
| FinQA | ARXIV | arXiv:2109.00122 | FinanceBenchmark |
| ConvFinQA | ARXIV | arXiv:2210.03849 | FinanceBenchmark |
| Finance Agent Benchmark | ARXIV | arXiv:2508.00828 | FinanceBenchmark |
| MultiHiertt | ARXIV | arXiv:2206.01347 | FinanceBenchmark |
| FinAuditing | ARXIV | arXiv:2510.08886 | FinanceBenchmark |
For benchmarks sourced from arXiv, we run our own evaluations against the published dataset using the original authors' scoring protocol. We do not modify task prompts, scoring rubrics, or evaluation scripts. Where the original evaluation code is available, we use it directly. Where it is not, we document our reimplementation in the benchmark's methodology page.
Conflict of interest
Structural Exclusion Policy
FinanceBenchmark structurally excludes any model or company in which a FinanceBenchmark founder, employee, or officer holds financial equity or other material interest, from all rankings and comparisons. This is an enforced platform rule, not a case-by-case disclosure.
Any such affiliated models will never appear in the leaderboard, composite scoring calculations, or individual model comparison matrices on this platform.
Advertising and commercial relationships
FinanceBenchmark does not display advertising. We do not accept payment from AI companies for evaluation, placement, or promotion. We do not have commercial partnerships with any of the organisations whose models we evaluate. Our operating costs are funded by Avae.
Submitting a benchmark or score
There are two submission paths. Both are free.
Submit a score
If you have run a model against one of our tracked benchmarks and have a reproducible result, you can submit the score for inclusion on the leaderboard. We require a link to the evaluation run or reproducible code — we do not accept self-reported scores without a verifiable source.
Score submissions are reviewed within five business days. We may run the evaluation ourselves to verify the result before publishing. We reserve the right to decline submissions that cannot be independently verified.
Request a certified evaluation
If you want your AI system evaluated against our full benchmark suite with a published certified score and full methodology documentation, you can request a certified evaluation. We run the evaluation against all eight tracked benchmarks using our open harness and publish the results with the same attribution and disclosure standards as all other scores.
Certified evaluations are currently free for research institutions and open-source models. For commercial systems, please contact us at submit@financebenchmark.ai to discuss terms.
Update frequency
The leaderboard is a living document. We update it as new models are released, as existing models are evaluated on additional benchmarks, and as new benchmarks are added to the suite.
Score corrections — where an error in our evaluation is discovered — are published promptly with a changelog entry noting what changed and why. We do not silently update scores. The date shown on each score reflects the evaluation run date, not the publication date on this site.
To be notified of leaderboard updates, subscribe to our mailing list via the footer of any page. We send one email per significant update, never more than once per week.