2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks
Domain
41 benchmarks
Pave Interchange Fee BenchmarkTier 1FB MAINTAINED
Forensic financial verification across 50,000 synthetic transaction records spanning four complexity levels from deterministic aggregation to strategic impact modelling.
24 tasks
9 models evaluated
VerificationForensic
1avae-2-0
100.0%
2Claude Opus 4.7
73.0%
3GPT-5.5
71.0%
4Gemini 3.1 Pro
69.0%
FinanceBenchmark (2026)9 / 42 models · Weight: 30 in FB Score
Finance Agent BenchmarkTier 1
Core financial analyst tasks and equity research evaluated using the Finance Agent (FAB v2) harness.
Variable tasks
40 models evaluated
Agent Tasks
1Gemini 3.8 Flash
61.4%
2Muse Spark 1.2
60.6%
3Muse Spark 1.3 Max
60.0%
4Gemini 3.7 Flash
59.0%
arXiv:2508.0082840 / 42 models · Weight: 30 in FB Score
VALS IndexTier 1
Holistic evaluation index measuring agent accuracy, reasoning, coding, and specialized domain proficiency.
Variable tasks
22 models evaluated
VerificationNumericalAgent Tasks
1Claude Opus 5
67.2%
2Claude Fable 5
66.0%
3GPT-5.6 Sol
63.7%
4GPT-5.6 Luna
59.9%
VALS (2026)22 / 42 models · Weight: 30 in FB Score
FinanceBenchTier 2
Financial question answering over real earnings documents and financial statements.
150 tasks
8 models evaluated
Document QA
1GPT-5.5
91.0%
2Claude Opus 4.7
88.0%
3Gemini 3.1 Pro
86.0%
4DeepSeek V4 Pro
85.0%
arXiv:2311.119448 / 42 models · Weight: 20 in FB Score
FinQATier 2
Numerical reasoning over financial documents requiring multi-step computation.
8,281 tasks
8 models evaluated
Numerical
1DeepSeek V4 Pro
87.0%
2Claude Opus 4.7
84.0%
3GPT-5.5
82.0%
4Gemini 3.1 Pro
80.0%
arXiv:2109.001228 / 42 models · Weight: 20 in FB Score
ConvFinQATier 2
Multi-turn conversational financial reasoning over earnings documents.
3,892 tasks
8 models evaluated
Document QA
1GPT-5.5
88.0%
2Claude Opus 4.7
86.0%
3DeepSeek V4 Pro
83.0%
4Gemini 3.1 Pro
82.0%
arXiv:2210.038498 / 42 models · Weight: 20 in FB Score
MortgageTax BenchmarkTier 2
Evaluates financial calculations and regulatory tax computations on complex mortgage agreements.
Variable tasks
17 models evaluated
NumericalForensic
1Claude Opus 5
72.1%
2Claude Sonnet 5
70.0%
3Claude Fable 5
68.9%
4Gemini 3.5 Flash Lite
68.7%
VALS (2026)17 / 42 models · Weight: 20 in FB Score
TaxEval v2Tier 2
Comprehensive benchmark for corporate and individual tax code interpretation and compliance reasoning.
Variable tasks
21 models evaluated
VerificationForensic
1Muse Spark 1.2
80.4%
2Muse Spark 1.1
79.7%
3Claude Fable 5
76.9%
4GPT-5.6 Luna
76.2%
VALS (2026)21 / 42 models · Weight: 20 in FB Score
GPQA DiamondTier 2
Graduate-level Google-proof Q&A benchmark across physics, chemistry, biology, and math.
198 tasks
20 models evaluated
NumericalVerification
1GPT-5.6 Sol
95.2%
2Grok 4.6
94.7%
3Qwen 3.8 Max
93.7%
4Claude Opus 5
93.4%
arXiv:2311.1202220 / 42 models · Weight: 20 in FB Score
MMLU ProTier 2
Robust multi-task language understanding reasoning benchmark with reasoning-focused questions.
12,000 tasks
21 models evaluated
NumericalVerification
1Claude Opus 5
91.6%
2Claude Fable 5
91.5%
3Grok 4.6
89.4%
4Grok 4.5
89.2%
arXiv:2406.0157421 / 42 models · Weight: 20 in FB Score
SWE-benchTier 2
Software engineering benchmark evaluating resolution of real-world GitHub issues.
2,294 tasks
22 models evaluated
Agent Tasks
1Claude Opus 5
97.0%
2DeepSeek V4 Pro (0813)
96.4%
3GPT-5.6 Sol
96.2%
4Grok 4.6
95.6%
arXiv:2310.0677022 / 42 models · Weight: 20 in FB Score
FinBenTier 3
Holistic financial benchmark covering 35 tasks across 23 financial datasets.
35 tasks
8 models evaluated
Multi-domain
1Gemini 3.1 Pro
85.0%
2Claude Opus 4.7
81.0%
3GPT-5.5
79.0%
4DeepSeek V4 Pro
78.0%
arXiv:2402.126598 / 42 models · Weight: 10 in FB Score
MultiHierttTier 3
Numerical reasoning over multi-hierarchical tabular and textual financial data.
1,200 tasks
8 models evaluated
Numerical
1DeepSeek V4 Pro
82.0%
2Gemini 3.1 Pro
80.0%
3Claude Opus 4.7
78.0%
4GPT-5.5
76.0%
arXiv:2206.013478 / 42 models · Weight: 10 in FB Score
FinAuditingTier 3
Multi-document financial audit reasoning requiring cross-document evidence synthesis.
Variable tasks
5 models evaluated
ForensicCompliance
1avae-2-0
88.0%
2Claude Opus 4.7
75.0%
3DeepSeek V4 Pro
73.0%
4GPT-5.5
72.0%
arXiv:2510.088865 / 42 models · Weight: 10 in FB Score
Code MigrationTier 3
Assesses code refactoring, legacy framework migration, and syntax translation accuracy.
Variable tasks
23 models evaluated
Agent Tasks
1Claude Opus 5
57.5%
2Claude Fable 5
55.1%
3GPT-5.6 Sol
52.9%
4Grok 4.6
44.6%
VALS (2026)23 / 42 models · Weight: 10 in FB Score
CyberBenchTier 3
Evaluates cybersecurity defense, vulnerability discovery, and exploit mitigation.
Variable tasks
11 models evaluated
Forensic
1GPT-5.6 Sol
88.1%
2GPT-5.6 Luna
83.9%
3Kimi K3
79.0%
4DeepSeek V4 Flash
77.5%
VALS (2026)11 / 42 models · Weight: 10 in FB Score
EMBTier 3
Evaluates multi-step enterprise business operations, accounting workflows, and strategic synthesis.
Variable tasks
39 models evaluated
Document QAForensic
1Claude Fable 5.1
76.7%
2Claude Fable 5
73.7%
3Claude Opus 5
73.6%
4GPT-5.6 Sol
72.3%
VALS (2026)39 / 42 models · Weight: 10 in FB Score
Legal Research BenchTier 3
Evaluates statutory interpretation, precedent discovery, and legal brief synthesis.
Variable tasks
22 models evaluated
Document QAForensic
1Claude Opus 5
55.3%
2Claude Fable 5
49.5%
3GPT-5.6 Sol
48.1%
4Grok 4.6
48.1%
VALS (2026)22 / 42 models · Weight: 10 in FB Score
MedCodeTier 3
Evaluates medical ontology mapping, ICD billing codes, and clinical software tasks.
Variable tasks
20 models evaluated
Verification
1Claude Opus 5
63.6%
2Claude Fable 5
56.1%
3Gemini 3.6 Flash
53.1%
4Muse Spark 1.2
49.4%
VALS (2026)20 / 42 models · Weight: 10 in FB Score
MedScribeTier 3
Clinical consultation documentation, transcription accuracy, and medical reporting.
Variable tasks
20 models evaluated
Document QA
1Claude Opus 5
91.0%
2Muse Spark 1.2
90.1%
3Muse Spark 1.1
88.9%
4Claude Fable 5
88.5%
VALS (2026)20 / 42 models · Weight: 10 in FB Score
ProofBench v1.1Tier 3
Formal mathematical theorem proving and step-by-step rigorous logical deduction.
Variable tasks
13 models evaluated
NumericalVerification
1Claude Opus 5
99.0%
2Claude Fable 5
95.0%
3Kimi K3
87.0%
4GPT-5.6 Sol
83.0%
VALS (2026)13 / 42 models · Weight: 10 in FB Score
SAGETier 3
Structured Agent Goal Execution across multi-modal tools and interactive environments.
Variable tasks
17 models evaluated
Agent Tasks
1Kimi K3
54.3%
2GPT-5.6 Sol
52.6%
3Claude Fable 5
51.9%
4Qwen 3.8 Max
51.3%
VALS (2026)17 / 42 models · Weight: 10 in FB Score
Public Benefits BenchTier 3
Evaluates civic policy compliance, public benefits eligibility, and regulatory adjudication.
Variable tasks
16 models evaluated
VerificationForensic
1Claude Opus 5
76.9%
2Claude Fable 5
70.4%
3Muse Spark 1.2
68.5%
4Kimi K3
68.3%
VALS (2026)16 / 42 models · Weight: 10 in FB Score
Vibe Code Bench v1.1Tier 3
Evaluates end-to-end full-stack frontend/backend generation and interactive software craftsmanship.
Variable tasks
23 models evaluated
Agent Tasks
1Claude Fable 5
90.3%
2Claude Opus 5
88.4%
3Kimi K3
85.0%
4DeepSeek V4 Pro (0813)
82.3%
VALS (2026)23 / 42 models · Weight: 10 in FB Score
Harvey's Legal AgentTier 3
Enterprise legal agent evaluation covering contract analysis, clause drafting, and risk assessment.
Variable tasks
22 models evaluated
Agent TasksForensic
1Muse Spark 1.2
25.4%
2Muse Spark 1.1
20.0%
3Grok 4.6
15.8%
4Grok 4.5
12.9%
VALS (2026)22 / 42 models · Weight: 10 in FB Score
IOITier 3
International Olympiad in Informatics algorithmic competitive programming tasks.
Variable tasks
8 models evaluated
Numerical
1Claude Opus 5
91.7%
2GPT-5.6 Sol
86.7%
3Qwen 3.8 Max
73.0%
4GPT-5.6 Luna
72.9%
IOI (2026)8 / 42 models · Weight: 10 in FB Score
LiveCodeBenchTier 3
Contamination-free competitive programming and code completion benchmark.
Variable tasks
19 models evaluated
Agent Tasks
1Claude Fable 5
89.8%
2Claude Opus 5
89.0%
3Grok 4.6
88.2%
4Gemini 3.6 Flash
88.0%
arXiv:2403.0797419 / 42 models · Weight: 10 in FB Score
LegalBenchTier 3
Collaboratively built benchmark for measuring legal reasoning in large language models.
Variable tasks
21 models evaluated
Document QAForensic
1Claude Fable 5
88.6%
2Claude Opus 5
87.0%
3GPT-5.6 Sol
87.0%
4Gemini 3.6 Flash
86.5%
arXiv:2308.1146221 / 42 models · Weight: 10 in FB Score
MMMUTier 3
Massive Multi-discipline Multimodal Understanding and Reasoning benchmark for expert AGI.
11,500 tasks
14 models evaluated
Document QANumerical
1Claude Opus 5
89.9%
2Claude Fable 5
89.3%
3GPT-5.6 Sol
88.8%
4Kimi K3
88.2%
arXiv:2311.1650214 / 42 models · Weight: 10 in FB Score
ProgramBenchTier 3
Program synthesis and execution verification across complex algorithmic challenges.
Variable tasks
17 models evaluated
Agent Tasks
1Claude Opus 5
3.0%
2Claude Fable 5
2.0%
3Kimi K3
2.0%
4GPT-5.6 Sol
1.5%
VALS (2026)17 / 42 models · Weight: 10 in FB Score
SkillsBenchTier 3
Measures procedural tool orchestration and software skill invocation.
Variable tasks
17 models evaluated
Agent Tasks
1Grok 4.5
66.0%
2GPT-5.6 Terra
60.6%
3GPT-5.6 Luna
60.5%
4Claude Opus 5
60.4%
VALS (2026)17 / 42 models · Weight: 10 in FB Score
Terminal-Bench 2.1Tier 3
Interactive Linux terminal command execution, shell debugging, and agent environment handling.
Variable tasks
23 models evaluated
Agent Tasks
1GPT-5.6 Sol
85.8%
2Claude Opus 5
84.6%
3Kimi K3
80.9%
4Claude Fable 5
80.5%
VALS (2026)23 / 42 models · Weight: 10 in FB Score
RSI IndexTier 3
Recursive Self-Improvement and recursive cognitive task orchestration index.
Variable tasks
3 models evaluated
Agent TasksNumerical
1Claude Opus 5
30.9%
2GPT-5.6 Sol
28.4%
3Kimi K3
22.1%
VALS (2026)3 / 42 models · Weight: 10 in FB Score
Time Horizon IndexTier 3
Long-horizon autonomous planning and multi-stage financial forecasting horizon.
Variable tasks
5 models evaluated
ForensicAgent Tasks
1Claude Opus 5
13.7%
2GPT-5.6 Sol
13.0%
3Grok 4.5
6.3%
4Kimi K3
5.2%
VALS (2026)5 / 42 models · Weight: 10 in FB Score
ReverseEngBenchTier 3
Binary analysis, decompilation reasoning, and software reverse engineering.
Variable tasks
4 models evaluated
Forensic
1GPT-5.6 Sol
30.5%
2Claude Opus 5
12.2%
3Grok 4.5
0.8%
4GLM 5.2
0.0%
VALS (2026)4 / 42 models · Weight: 10 in FB Score
Tax Agent BenchTier 3
Evaluates agents on professional tax compliance, preparation, and advisory reasoning.
Variable tasks
19 models evaluated
ForensicAgent Tasks
1Claude Fable 5.1
77.6%
2Claude Opus 5
75.1%
3GLM 5.3
73.1%
4Muse Spark 1.3
71.9%
VALS (2026)19 / 42 models · Weight: 10 in FB Score
Big Finance BenchTier 3
Evaluates models on complex financial analysis tasks that require retrieval and multi-step reasoning.
Variable tasks
0 models evaluated
ForensicDocument QA
LLM Stats (2026)0 / 42 models · Weight: 10 in FB Score
GDPval-MMTier 3
Multimodal benchmark evaluating AI performance on economically valuable tasks across documents, spreadsheets, and diagrams.
Variable tasks
0 models evaluated
Document QANumerical
arXiv:2510.043740 / 42 models · Weight: 10 in FB Score
BankerToolBenchTier 3
Evaluates models on banking and finance tool-use and API interaction workflows.
Variable tasks
0 models evaluated
Agent Tasks
LLM Stats (2026)0 / 42 models · Weight: 10 in FB Score
GDPval-RubricsTier 3
Evaluates AI model performance on rubric-graded economically valuable knowledge work tasks.
Variable tasks
0 models evaluated
Forensic
LLM Stats (2026)0 / 42 models · Weight: 10 in FB Score
PRBench-FinanceTier 3
Professional reasoning evaluations on complex financial tasks.
Variable tasks
0 models evaluated
NumericalForensic
LLM Stats (2026)0 / 42 models · Weight: 10 in FB Score