Pave Interchange Fee BenchmarkTier 1FB MAINTAINED
Forensic financial verification across 50,000 synthetic transaction records spanning four complexity levels from deterministic aggregation to strategic impact modelling.
Top modelsFull leaderboard →
FinanceBenchmark (2026)9 / 42 models · Weight: 30 in FB Score
Finance Agent BenchmarkTier 1
Core financial analyst tasks and equity research evaluated using the Finance Agent (FAB v2) harness.
Top modelsFull leaderboard →
arXiv:2508.0082840 / 42 models · Weight: 30 in FB Score
VALS IndexTier 1
Holistic evaluation index measuring agent accuracy, reasoning, coding, and specialized domain proficiency.
Top modelsFull leaderboard →
VALS (2026)22 / 42 models · Weight: 30 in FB Score
FinanceBenchTier 2
Financial question answering over real earnings documents and financial statements.
Top modelsFull leaderboard →
arXiv:2311.119448 / 42 models · Weight: 20 in FB Score
FinQATier 2
Numerical reasoning over financial documents requiring multi-step computation.
Top modelsFull leaderboard →
arXiv:2109.001228 / 42 models · Weight: 20 in FB Score
ConvFinQATier 2
Multi-turn conversational financial reasoning over earnings documents.
Top modelsFull leaderboard →
arXiv:2210.038498 / 42 models · Weight: 20 in FB Score
MortgageTax BenchmarkTier 2
Evaluates financial calculations and regulatory tax computations on complex mortgage agreements.
Top modelsFull leaderboard →
VALS (2026)17 / 42 models · Weight: 20 in FB Score
TaxEval v2Tier 2
Comprehensive benchmark for corporate and individual tax code interpretation and compliance reasoning.
Top modelsFull leaderboard →
VALS (2026)21 / 42 models · Weight: 20 in FB Score
GPQA DiamondTier 2
Graduate-level Google-proof Q&A benchmark across physics, chemistry, biology, and math.
Top modelsFull leaderboard →
arXiv:2311.1202220 / 42 models · Weight: 20 in FB Score
MMLU ProTier 2
Robust multi-task language understanding reasoning benchmark with reasoning-focused questions.
Top modelsFull leaderboard →
arXiv:2406.0157421 / 42 models · Weight: 20 in FB Score
SWE-benchTier 2
Software engineering benchmark evaluating resolution of real-world GitHub issues.
Top modelsFull leaderboard →
arXiv:2310.0677022 / 42 models · Weight: 20 in FB Score
FinBenTier 3
Holistic financial benchmark covering 35 tasks across 23 financial datasets.
Top modelsFull leaderboard →
arXiv:2402.126598 / 42 models · Weight: 10 in FB Score
MultiHierttTier 3
Numerical reasoning over multi-hierarchical tabular and textual financial data.
Top modelsFull leaderboard →
arXiv:2206.013478 / 42 models · Weight: 10 in FB Score
FinAuditingTier 3
Multi-document financial audit reasoning requiring cross-document evidence synthesis.
Top modelsFull leaderboard →
arXiv:2510.088865 / 42 models · Weight: 10 in FB Score
Code MigrationTier 3
Assesses code refactoring, legacy framework migration, and syntax translation accuracy.
Top modelsFull leaderboard →
VALS (2026)23 / 42 models · Weight: 10 in FB Score
CyberBenchTier 3
Evaluates cybersecurity defense, vulnerability discovery, and exploit mitigation.
Top modelsFull leaderboard →
VALS (2026)11 / 42 models · Weight: 10 in FB Score
EMBTier 3
Evaluates multi-step enterprise business operations, accounting workflows, and strategic synthesis.
Top modelsFull leaderboard →
VALS (2026)39 / 42 models · Weight: 10 in FB Score
Legal Research BenchTier 3
Evaluates statutory interpretation, precedent discovery, and legal brief synthesis.
Top modelsFull leaderboard →
VALS (2026)22 / 42 models · Weight: 10 in FB Score
MedCodeTier 3
Evaluates medical ontology mapping, ICD billing codes, and clinical software tasks.
Top modelsFull leaderboard →
VALS (2026)20 / 42 models · Weight: 10 in FB Score
MedScribeTier 3
Clinical consultation documentation, transcription accuracy, and medical reporting.
Top modelsFull leaderboard →
VALS (2026)20 / 42 models · Weight: 10 in FB Score
ProofBench v1.1Tier 3
Formal mathematical theorem proving and step-by-step rigorous logical deduction.
Top modelsFull leaderboard →
VALS (2026)13 / 42 models · Weight: 10 in FB Score
SAGETier 3
Structured Agent Goal Execution across multi-modal tools and interactive environments.
Top modelsFull leaderboard →
VALS (2026)17 / 42 models · Weight: 10 in FB Score
Public Benefits BenchTier 3
Evaluates civic policy compliance, public benefits eligibility, and regulatory adjudication.
Top modelsFull leaderboard →
VALS (2026)16 / 42 models · Weight: 10 in FB Score
Vibe Code Bench v1.1Tier 3
Evaluates end-to-end full-stack frontend/backend generation and interactive software craftsmanship.
Top modelsFull leaderboard →
VALS (2026)23 / 42 models · Weight: 10 in FB Score
Harvey's Legal AgentTier 3
Enterprise legal agent evaluation covering contract analysis, clause drafting, and risk assessment.
Top modelsFull leaderboard →
VALS (2026)22 / 42 models · Weight: 10 in FB Score
IOITier 3
International Olympiad in Informatics algorithmic competitive programming tasks.
Top modelsFull leaderboard →
IOI (2026)8 / 42 models · Weight: 10 in FB Score
LiveCodeBenchTier 3
Contamination-free competitive programming and code completion benchmark.
Top modelsFull leaderboard →
arXiv:2403.0797419 / 42 models · Weight: 10 in FB Score
LegalBenchTier 3
Collaboratively built benchmark for measuring legal reasoning in large language models.
Top modelsFull leaderboard →
arXiv:2308.1146221 / 42 models · Weight: 10 in FB Score
MMMUTier 3
Massive Multi-discipline Multimodal Understanding and Reasoning benchmark for expert AGI.
Top modelsFull leaderboard →
arXiv:2311.1650214 / 42 models · Weight: 10 in FB Score
ProgramBenchTier 3
Program synthesis and execution verification across complex algorithmic challenges.
Top modelsFull leaderboard →
VALS (2026)17 / 42 models · Weight: 10 in FB Score
SkillsBenchTier 3
Measures procedural tool orchestration and software skill invocation.
Top modelsFull leaderboard →
VALS (2026)17 / 42 models · Weight: 10 in FB Score
Terminal-Bench 2.1Tier 3
Interactive Linux terminal command execution, shell debugging, and agent environment handling.
Top modelsFull leaderboard →
VALS (2026)23 / 42 models · Weight: 10 in FB Score
RSI IndexTier 3
Recursive Self-Improvement and recursive cognitive task orchestration index.
Top modelsFull leaderboard →
VALS (2026)3 / 42 models · Weight: 10 in FB Score
Time Horizon IndexTier 3
Long-horizon autonomous planning and multi-stage financial forecasting horizon.
Top modelsFull leaderboard →
VALS (2026)5 / 42 models · Weight: 10 in FB Score
ReverseEngBenchTier 3
Binary analysis, decompilation reasoning, and software reverse engineering.
Top modelsFull leaderboard →
VALS (2026)4 / 42 models · Weight: 10 in FB Score
Tax Agent BenchTier 3
Evaluates agents on professional tax compliance, preparation, and advisory reasoning.
Top modelsFull leaderboard →
VALS (2026)19 / 42 models · Weight: 10 in FB Score
Big Finance BenchTier 3
Evaluates models on complex financial analysis tasks that require retrieval and multi-step reasoning.
Top modelsFull leaderboard →
LLM Stats (2026)0 / 42 models · Weight: 10 in FB Score
GDPval-MMTier 3
Multimodal benchmark evaluating AI performance on economically valuable tasks across documents, spreadsheets, and diagrams.
Top modelsFull leaderboard →
arXiv:2510.043740 / 42 models · Weight: 10 in FB Score
BankerToolBenchTier 3
Evaluates models on banking and finance tool-use and API interaction workflows.
Top modelsFull leaderboard →
LLM Stats (2026)0 / 42 models · Weight: 10 in FB Score
GDPval-RubricsTier 3
Evaluates AI model performance on rubric-graded economically valuable knowledge work tasks.
Top modelsFull leaderboard →
LLM Stats (2026)0 / 42 models · Weight: 10 in FB Score
PRBench-FinanceTier 3
Professional reasoning evaluations on complex financial tasks.
Top modelsFull leaderboard →
LLM Stats (2026)0 / 42 models · Weight: 10 in FB Score