2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks
Tier 1FB Maintainedverificationforensic

Pave Interchange Fee Benchmark

Forensic financial verification across 50,000 synthetic transaction records spanning four complexity levels from deterministic aggregation to strategic impact modelling.

24 tasks
8 models evaluated
Updated 2026
100.0%
Top Score
Top model
avae-2-0
Model Rankings
Overall score across all four complexity levels, weighted equally.
financebenchmark.ai · Pave
Anthropic
OpenAI
Google
DeepSeek
xAI
Moonshot AI
Zhipu AI
MiniMax
66.0%
🥇AnthropicAnthropicClaude Opus 4.773.0%
🥈OpenAIOpenAIGPT-5.571.0%
🥉GoogleGoogleGemini 3.1 Pro69.0%
4DeepSeekDeepSeekDeepSeek V4 ProOPEN65.0%
5xAIxAIGrok 4.368.0%
6Moonshot AIMoonshot AIKimi K2.6OPEN66.0%
7Zhipu AIZhipu AIGLM-5.1OPEN63.0%
8MiniMaxMiniMaxMiniMax M2.7OPEN62.0%
Full Results Table
#ModelOverallLevel 1Level 2Level 3Level 4SourceDate
🥇
AnthropicClaude Opus 4.7
Anthropic
73.096.075.052.068.0FinanceBenchmark (2026)2026-05
🥈
OpenAIGPT-5.5
OpenAI
71.095.072.050.066.0FinanceBenchmark (2026)2026-05
🥉
GoogleGemini 3.1 Pro
Google
69.094.070.048.065.0FinanceBenchmark (2026)2026-05
4
xAIGrok 4.3
xAI
68.093.068.047.063.0FinanceBenchmark (2026)2026-05
5
Moonshot AIKimi K2.6OPEN
Moonshot AI
66.092.066.045.061.0FinanceBenchmark (2026)2026-05
6
DeepSeekDeepSeek V4 ProOPEN
DeepSeek
65.091.065.044.060.0FinanceBenchmark (2026)2026-05
7
Zhipu AIGLM-5.1OPEN
Zhipu AI
63.090.063.042.058.0FinanceBenchmark (2026)2026-05
8
MiniMaxMiniMax M2.7OPEN
MiniMax
62.089.061.040.056.0FinanceBenchmark (2026)2026-05