2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks
Leads on verification
Claude Opus 5
81.0% · Verification overall
Best forensic discovery
Claude Fable 5.1
77.2% · Forensic Reasoning
Best at document QA
GPT-5.5
81.2% · Document QA
FB Index — Overall Score9/15/2026

Top models ranked on composite FB Score across all financial reasoning benchmarks. View full list →

1AnthropicAnthropicClaude Opus 4.774.20%
2OpenAIOpenAIGPT-5.573.10%
3AnthropicAnthropicClaude Fable 570.70%
4DeepSeekDeepSeekDeepSeek V4 ProOPEN70.10%
5GoogleGoogleGemini 3.1 Pro70.00%
6Moonshot AIMoonshot AIKimi K2.6OPEN67.30%
7AnthropicAnthropicClaude Opus 567.10%
8xAIxAIGrok 4.666.70%
9AnthropicAnthropicClaude Fable 5.166.20%
10Zhipu AIZhipu AIGLM-5.1OPEN65.90%
11OpenAIOpenAIGPT-5.6 Sol65.10%
12GoogleGoogleGemini 3.8 Flash64.70%
13xAIxAIGrok 4.364.60%
14OpenAIOpenAIGPT-5.6 Luna64.60%
15MetaMetaMuse Spark 1.264.20%
Full leaderboard
#ModelFB Score ⓘPaveFinanceBenchFinBenFinQAConvFinQABenchmarksSourceDate
🥇
Anthropic Claude Opus 4.7
Anthropic
74.273.088.081.084.086.09 / 41arXiv2026-03
🥈
OpenAI GPT-5.5
OpenAI
73.171.091.079.082.088.010 / 41arXiv2026-02
🥉
Anthropic Claude Fable 5
Anthropic
70.723 / 41arXiv2026-06-09
4
DeepSeek DeepSeek V4 ProOPEN
DeepSeek
70.165.085.078.087.083.09 / 41arXiv2026-01
5
Google Gemini 3.1 Pro
Google
70.069.086.085.080.082.09 / 41arXiv2026-01
6
Moonshot AI Kimi K2.6OPEN
Moonshot AI
67.366.082.075.076.078.08 / 41arXiv2026-02
7
Anthropic Claude Opus 5
Anthropic
67.129 / 41arXiv2026-07-22
8
xAI Grok 4.6
xAI
66.721 / 41arXiv2026-08-12
9
Anthropic Claude Fable 5.1
Anthropic
66.23 / 41arXiv2026-09-01
10
Zhipu AI GLM-5.1OPEN
Zhipu AI
65.963.079.073.074.075.07 / 41arXiv2026-01
11
OpenAI GPT-5.6 Sol
OpenAI
65.129 / 41arXiv2026-07-09
12
Google Gemini 3.8 Flash
Google
64.73 / 41arXiv2026-09-02
13
xAI Grok 4.3
xAI
64.668.084.077.078.079.08 / 41arXiv2026-02
14
OpenAI GPT-5.6 Luna
OpenAI
64.625 / 41arXiv2026-07-09
15
Meta Muse Spark 1.2
Meta
64.221 / 41arXiv2026-08-05
16
Anthropic Claude Sonnet 5
Anthropic
63.924 / 41arXiv2026-06-30
17
Moonshot AI Kimi K3OPEN
Moonshot AI
63.826 / 41arXiv2026-07-16
18
DeepSeek DeepSeek V4 Pro (0813)OPEN
DeepSeek
63.716 / 41arXiv2026-08-13
19
DeepSeek DeepSeek V4 FlashOPEN
DeepSeek
63.219 / 41arXiv2026-07-31
20
OpenAI GPT-5.6 Terra
OpenAI
63.125 / 41arXiv2026-07-09
21
Meta Muse Spark 1.3
Meta
62.33 / 41arXiv2026-09-02
22
Meta Muse Spark 1.1
Meta
62.222 / 41arXiv2026-07-09
23
Alibaba Qwen 3.8 MaxOPEN
Alibaba
62.225 / 41arXiv2026-08-03
24
Meta Muse Spark 1.3 Max
Meta
61.82 / 41arXiv2026-09-02
25
Google Gemini 3.6 Flash
Google
61.622 / 41arXiv2026-07-21
26
Google Gemini 3.7 Flash
Google
61.23 / 41arXiv2026-08
27
Zhipu AI GLM 5.3OPEN
Zhipu AI
59.43 / 41arXiv2026-08
28
OpenAI GPT-6 Astra
OpenAI
59.13 / 41arXiv2026-09-03
29
MiniMax MiniMax M2.7OPEN
MiniMax
58.762.077.071.072.073.08 / 41arXiv2026-01
30
Zhipu AI GLM 5.2OPEN
Zhipu AI
57.520 / 41arXiv2026-06-13
31
Zhipu AI GLM 5.3 FlashOPEN
Zhipu AI
57.42 / 41arXiv2026-08
32
xAI Grok 4.5
xAI
56.224 / 41arXiv2026-07-16
33
DeepSeek DeepSeek V4.1 FlashOPEN
DeepSeek
56.03 / 41arXiv2026-09-10
34
Google Gemini 3.5 Flash Lite
Google
53.523 / 41arXiv2026-07-21
35
Moonshot AI Kimi K2.7 CodeOPEN
Moonshot AI
53.57 / 41arXiv2026-06-12
36
Thinking Machines InklingOPEN
Thinking Machines
51.325 / 41arXiv2026-07-15
37
Thinking Machines Inkling SmallOPEN
Thinking Machines
50.423 / 41arXiv2026-07-30
38
MiniMax MiniMax M3OPEN
MiniMax
48.53 / 41arXiv2026-06
39
NVIDIA Nemotron 3 UltraOPEN
NVIDIA
46.316 / 41arXiv2026-06-04
40
Ant Group Ling 3.0 FlashOPEN
Ant Group
45.018 / 41arXiv2026-07-23
41
Alibaba Qwen 3.7 Plus
Alibaba
39.512 / 41arXiv2026-06-01
42
Zhipu AI GLM-5.5OPEN
Zhipu AI
0 / 41arXiv2026-08
Tracked BenchmarksView all →
Pave Interchange Fee Benchmark
Forensic financial verification across 50,000 synthetic transaction records spanning four complexity levels from deterministic aggregation to strategic impact modelling.
VerificationForensic24 tasksTier 1

avae-2-0100.0
9 modelsFinanceBenchmark 2026
FinanceBench
Financial question answering over real earnings documents and financial statements.
DocumentQA150 tasksTier 2

GPT-5.591.0
8 modelsarXiv:2311.11944
FinBen
Holistic financial benchmark covering 35 tasks across 23 financial datasets.
MultiDomain35 tasksTier 3

Gemini 3.1 Pro85.0
8 modelsarXiv:2402.12659
FinQA
Numerical reasoning over financial documents requiring multi-step computation.
Numerical8,281 tasksTier 2

DeepSeek V4 Pro87.0
8 modelsarXiv:2109.00122
ConvFinQA
Multi-turn conversational financial reasoning over earnings documents.
DocumentQA3,892 tasksTier 2

GPT-5.588.0
8 modelsarXiv:2210.03849
Finance Agent Benchmark
Core financial analyst tasks and equity research evaluated using the Finance Agent (FAB v2) harness.
AgentTasksVariable tasksTier 1

Gemini 3.8 Flash61.4
40 modelsarXiv:2508.00828
MultiHiertt
Numerical reasoning over multi-hierarchical tabular and textual financial data.
Numerical1,200 tasksTier 3

DeepSeek V4 Pro82.0
8 modelsarXiv:2206.01347
FinAuditing
Multi-document financial audit reasoning requiring cross-document evidence synthesis.
ForensicComplianceVariable tasksTier 3

avae-2-088.0
5 modelsarXiv:2510.08886
VALS Index
Holistic evaluation index measuring agent accuracy, reasoning, coding, and specialized domain proficiency.
VerificationNumericalVariable tasksTier 1

Claude Opus 567.2
22 modelsVALS (2026)
MortgageTax Benchmark
Evaluates financial calculations and regulatory tax computations on complex mortgage agreements.
NumericalForensicVariable tasksTier 2

Claude Opus 572.1
17 modelsVALS (2026)
TaxEval v2
Comprehensive benchmark for corporate and individual tax code interpretation and compliance reasoning.
VerificationForensicVariable tasksTier 2

Muse Spark 1.280.4
21 modelsVALS (2026)
Code Migration
Assesses code refactoring, legacy framework migration, and syntax translation accuracy.
AgentTasksVariable tasksTier 3

Claude Opus 557.5
23 modelsVALS (2026)
CyberBench
Evaluates cybersecurity defense, vulnerability discovery, and exploit mitigation.
ForensicVariable tasksTier 3

GPT-5.6 Sol88.1
11 modelsVALS (2026)
EMB
Evaluates multi-step enterprise business operations, accounting workflows, and strategic synthesis.
DocumentQAForensicVariable tasksTier 3

Claude Fable 5.176.7
39 modelsVALS (2026)
Legal Research Bench
Evaluates statutory interpretation, precedent discovery, and legal brief synthesis.
DocumentQAForensicVariable tasksTier 3

Claude Opus 555.3
22 modelsVALS (2026)
MedCode
Evaluates medical ontology mapping, ICD billing codes, and clinical software tasks.
VerificationVariable tasksTier 3

Claude Opus 563.6
20 modelsVALS (2026)
MedScribe
Clinical consultation documentation, transcription accuracy, and medical reporting.
DocumentQAVariable tasksTier 3

Claude Opus 591.0
20 modelsVALS (2026)
ProofBench v1.1
Formal mathematical theorem proving and step-by-step rigorous logical deduction.
NumericalVerificationVariable tasksTier 3

Claude Opus 599.0
13 modelsVALS (2026)
SAGE
Structured Agent Goal Execution across multi-modal tools and interactive environments.
AgentTasksVariable tasksTier 3

Kimi K354.3
17 modelsVALS (2026)
Public Benefits Bench
Evaluates civic policy compliance, public benefits eligibility, and regulatory adjudication.
VerificationForensicVariable tasksTier 3

Claude Opus 576.9
16 modelsVALS (2026)
Vibe Code Bench v1.1
Evaluates end-to-end full-stack frontend/backend generation and interactive software craftsmanship.
AgentTasksVariable tasksTier 3

Claude Fable 590.3
23 modelsVALS (2026)
Harvey's Legal Agent
Enterprise legal agent evaluation covering contract analysis, clause drafting, and risk assessment.
AgentTasksForensicVariable tasksTier 3

Muse Spark 1.225.4
22 modelsVALS (2026)
GPQA Diamond
Graduate-level Google-proof Q&A benchmark across physics, chemistry, biology, and math.
NumericalVerification198 tasksTier 2

GPT-5.6 Sol95.2
20 modelsarXiv:2311.12022
IOI
International Olympiad in Informatics algorithmic competitive programming tasks.
NumericalVariable tasksTier 3

Claude Opus 591.7
8 modelsIOI (2026)
LiveCodeBench
Contamination-free competitive programming and code completion benchmark.
AgentTasksVariable tasksTier 3

Claude Fable 589.8
19 modelsarXiv:2403.07974
LegalBench
Collaboratively built benchmark for measuring legal reasoning in large language models.
DocumentQAForensicVariable tasksTier 3

Claude Fable 588.6
21 modelsarXiv:2308.11462
MMLU Pro
Robust multi-task language understanding reasoning benchmark with reasoning-focused questions.
NumericalVerification12,000 tasksTier 2

Claude Opus 591.6
21 modelsarXiv:2406.01574
MMMU
Massive Multi-discipline Multimodal Understanding and Reasoning benchmark for expert AGI.
DocumentQANumerical11,500 tasksTier 3

Claude Opus 589.9
14 modelsarXiv:2311.16502
ProgramBench
Program synthesis and execution verification across complex algorithmic challenges.
AgentTasksVariable tasksTier 3

Claude Opus 53.0
17 modelsVALS (2026)
SkillsBench
Measures procedural tool orchestration and software skill invocation.
AgentTasksVariable tasksTier 3

Grok 4.566.0
17 modelsVALS (2026)
SWE-bench
Software engineering benchmark evaluating resolution of real-world GitHub issues.
AgentTasks2,294 tasksTier 2

Claude Opus 597.0
22 modelsarXiv:2310.06770
Terminal-Bench 2.1
Interactive Linux terminal command execution, shell debugging, and agent environment handling.
AgentTasksVariable tasksTier 3

GPT-5.6 Sol85.8
23 modelsVALS (2026)
RSI Index
Recursive Self-Improvement and recursive cognitive task orchestration index.
AgentTasksNumericalVariable tasksTier 3

Claude Opus 530.9
3 modelsVALS (2026)
Time Horizon Index
Long-horizon autonomous planning and multi-stage financial forecasting horizon.
ForensicAgentTasksVariable tasksTier 3

Claude Opus 513.7
5 modelsVALS (2026)
ReverseEngBench
Binary analysis, decompilation reasoning, and software reverse engineering.
ForensicVariable tasksTier 3

GPT-5.6 Sol30.5
4 modelsVALS (2026)
Tax Agent Bench
Evaluates agents on professional tax compliance, preparation, and advisory reasoning.
ForensicAgentTasksVariable tasksTier 3

Claude Fable 5.177.6
19 modelsVALS (2026)
Big Finance Bench
Evaluates models on complex financial analysis tasks that require retrieval and multi-step reasoning.
ForensicDocumentQAVariable tasksTier 3

N/A-
0 modelsLLM Stats (2026)
GDPval-MM
Multimodal benchmark evaluating AI performance on economically valuable tasks across documents, spreadsheets, and diagrams.
DocumentQANumericalVariable tasksTier 3

N/A-
0 modelsarXiv:2510.04374
BankerToolBench
Evaluates models on banking and finance tool-use and API interaction workflows.
AgentTasksVariable tasksTier 3

N/A-
0 modelsLLM Stats (2026)
GDPval-Rubrics
Evaluates AI model performance on rubric-graded economically valuable knowledge work tasks.
ForensicVariable tasksTier 3

N/A-
0 modelsLLM Stats (2026)
PRBench-Finance
Professional reasoning evaluations on complex financial tasks.
NumericalForensicVariable tasksTier 3

N/A-
0 modelsLLM Stats (2026)