Pave Interchange Fee Benchmark
Forensic financial verification across 50,000 synthetic transaction records spanning four complexity levels from deterministic aggregation to strategic impact modelling.
VerificationForensic24 tasksTier 1
avae-2-0100.0
9 modelsFinanceBenchmark 2026
FinanceBench
Financial question answering over real earnings documents and financial statements.
DocumentQA150 tasksTier 2
GPT-5.591.0
8 modelsarXiv:2311.11944
FinBen
Holistic financial benchmark covering 35 tasks across 23 financial datasets.
MultiDomain35 tasksTier 3
Gemini 3.1 Pro85.0
8 modelsarXiv:2402.12659
FinQA
Numerical reasoning over financial documents requiring multi-step computation.
Numerical8,281 tasksTier 2
DeepSeek V4 Pro87.0
8 modelsarXiv:2109.00122
ConvFinQA
Multi-turn conversational financial reasoning over earnings documents.
DocumentQA3,892 tasksTier 2
GPT-5.588.0
8 modelsarXiv:2210.03849
Finance Agent Benchmark
Core financial analyst tasks and equity research evaluated using the Finance Agent (FAB v2) harness.
AgentTasksVariable tasksTier 1
Gemini 3.8 Flash61.4
40 modelsarXiv:2508.00828
MultiHiertt
Numerical reasoning over multi-hierarchical tabular and textual financial data.
Numerical1,200 tasksTier 3
DeepSeek V4 Pro82.0
8 modelsarXiv:2206.01347
FinAuditing
Multi-document financial audit reasoning requiring cross-document evidence synthesis.
ForensicComplianceVariable tasksTier 3
avae-2-088.0
5 modelsarXiv:2510.08886
VALS Index
Holistic evaluation index measuring agent accuracy, reasoning, coding, and specialized domain proficiency.
VerificationNumericalVariable tasksTier 1
Claude Opus 567.2
22 modelsVALS (2026)
MortgageTax Benchmark
Evaluates financial calculations and regulatory tax computations on complex mortgage agreements.
NumericalForensicVariable tasksTier 2
Claude Opus 572.1
17 modelsVALS (2026)
TaxEval v2
Comprehensive benchmark for corporate and individual tax code interpretation and compliance reasoning.
VerificationForensicVariable tasksTier 2
Muse Spark 1.280.4
21 modelsVALS (2026)
Code Migration
Assesses code refactoring, legacy framework migration, and syntax translation accuracy.
AgentTasksVariable tasksTier 3
Claude Opus 557.5
23 modelsVALS (2026)
CyberBench
Evaluates cybersecurity defense, vulnerability discovery, and exploit mitigation.
ForensicVariable tasksTier 3
GPT-5.6 Sol88.1
11 modelsVALS (2026)
EMB
Evaluates multi-step enterprise business operations, accounting workflows, and strategic synthesis.
DocumentQAForensicVariable tasksTier 3
Claude Fable 5.176.7
39 modelsVALS (2026)
Legal Research Bench
Evaluates statutory interpretation, precedent discovery, and legal brief synthesis.
DocumentQAForensicVariable tasksTier 3
Claude Opus 555.3
22 modelsVALS (2026)
MedCode
Evaluates medical ontology mapping, ICD billing codes, and clinical software tasks.
VerificationVariable tasksTier 3
Claude Opus 563.6
20 modelsVALS (2026)
MedScribe
Clinical consultation documentation, transcription accuracy, and medical reporting.
DocumentQAVariable tasksTier 3
Claude Opus 591.0
20 modelsVALS (2026)
ProofBench v1.1
Formal mathematical theorem proving and step-by-step rigorous logical deduction.
NumericalVerificationVariable tasksTier 3
Claude Opus 599.0
13 modelsVALS (2026)
SAGE
Structured Agent Goal Execution across multi-modal tools and interactive environments.
AgentTasksVariable tasksTier 3
Kimi K354.3
17 modelsVALS (2026)
Public Benefits Bench
Evaluates civic policy compliance, public benefits eligibility, and regulatory adjudication.
VerificationForensicVariable tasksTier 3
Claude Opus 576.9
16 modelsVALS (2026)
Vibe Code Bench v1.1
Evaluates end-to-end full-stack frontend/backend generation and interactive software craftsmanship.
AgentTasksVariable tasksTier 3
Claude Fable 590.3
23 modelsVALS (2026)
Harvey's Legal Agent
Enterprise legal agent evaluation covering contract analysis, clause drafting, and risk assessment.
AgentTasksForensicVariable tasksTier 3
Muse Spark 1.225.4
22 modelsVALS (2026)
GPQA Diamond
Graduate-level Google-proof Q&A benchmark across physics, chemistry, biology, and math.
NumericalVerification198 tasksTier 2
GPT-5.6 Sol95.2
20 modelsarXiv:2311.12022
IOI
International Olympiad in Informatics algorithmic competitive programming tasks.
NumericalVariable tasksTier 3
Claude Opus 591.7
8 modelsIOI (2026)
LiveCodeBench
Contamination-free competitive programming and code completion benchmark.
AgentTasksVariable tasksTier 3
Claude Fable 589.8
19 modelsarXiv:2403.07974
LegalBench
Collaboratively built benchmark for measuring legal reasoning in large language models.
DocumentQAForensicVariable tasksTier 3
Claude Fable 588.6
21 modelsarXiv:2308.11462
MMLU Pro
Robust multi-task language understanding reasoning benchmark with reasoning-focused questions.
NumericalVerification12,000 tasksTier 2
Claude Opus 591.6
21 modelsarXiv:2406.01574
MMMU
Massive Multi-discipline Multimodal Understanding and Reasoning benchmark for expert AGI.
DocumentQANumerical11,500 tasksTier 3
Claude Opus 589.9
14 modelsarXiv:2311.16502
ProgramBench
Program synthesis and execution verification across complex algorithmic challenges.
AgentTasksVariable tasksTier 3
Claude Opus 53.0
17 modelsVALS (2026)
SkillsBench
Measures procedural tool orchestration and software skill invocation.
AgentTasksVariable tasksTier 3
Grok 4.566.0
17 modelsVALS (2026)
SWE-bench
Software engineering benchmark evaluating resolution of real-world GitHub issues.
AgentTasks2,294 tasksTier 2
Claude Opus 597.0
22 modelsarXiv:2310.06770
Terminal-Bench 2.1
Interactive Linux terminal command execution, shell debugging, and agent environment handling.
AgentTasksVariable tasksTier 3
GPT-5.6 Sol85.8
23 modelsVALS (2026)
RSI Index
Recursive Self-Improvement and recursive cognitive task orchestration index.
AgentTasksNumericalVariable tasksTier 3
Claude Opus 530.9
3 modelsVALS (2026)
Time Horizon Index
Long-horizon autonomous planning and multi-stage financial forecasting horizon.
ForensicAgentTasksVariable tasksTier 3
Claude Opus 513.7
5 modelsVALS (2026)
ReverseEngBench
Binary analysis, decompilation reasoning, and software reverse engineering.
ForensicVariable tasksTier 3
GPT-5.6 Sol30.5
4 modelsVALS (2026)
Tax Agent Bench
Evaluates agents on professional tax compliance, preparation, and advisory reasoning.
ForensicAgentTasksVariable tasksTier 3
Claude Fable 5.177.6
19 modelsVALS (2026)
Big Finance Bench
Evaluates models on complex financial analysis tasks that require retrieval and multi-step reasoning.
ForensicDocumentQAVariable tasksTier 3
N/A-
0 modelsLLM Stats (2026)
GDPval-MM
Multimodal benchmark evaluating AI performance on economically valuable tasks across documents, spreadsheets, and diagrams.
DocumentQANumericalVariable tasksTier 3
N/A-
0 modelsarXiv:2510.04374
BankerToolBench
Evaluates models on banking and finance tool-use and API interaction workflows.
AgentTasksVariable tasksTier 3
N/A-
0 modelsLLM Stats (2026)
GDPval-Rubrics
Evaluates AI model performance on rubric-graded economically valuable knowledge work tasks.
ForensicVariable tasksTier 3
N/A-
0 modelsLLM Stats (2026)
PRBench-Finance
Professional reasoning evaluations on complex financial tasks.
NumericalForensicVariable tasksTier 3
N/A-
0 modelsLLM Stats (2026)