2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks
Tier 3documentQA

MedScribe

Clinical consultation documentation, transcription accuracy, and medical reporting.

Variable tasks
20 models evaluated
Updated 2026
91.0%
Top Score
Top model
Claude Opus 5
Model Rankings
financebenchmark.ai · MedScribe
Anthropic
Anthropic
OpenAI
DeepSeek
xAI
Moonshot AI
Anthropic
DeepSeek
Google
Google
Zhipu AI
OpenAI
OpenAI
xAI
Thinking Machines
Thinking Machines
Ant Group
Meta
Meta
Alibaba
88.0%
🥇AnthropicAnthropicClaude Opus 591.0%
🥈AnthropicAnthropicClaude Sonnet 576.0%
🥉OpenAIOpenAIGPT-5.6 Sol85.2%
4DeepSeekDeepSeekDeepSeek V4 FlashOPEN80.4%
5xAIxAIGrok 4.686.5%
6Moonshot AIMoonshot AIKimi K3OPEN88.0%
7AnthropicAnthropicClaude Fable 588.5%
8DeepSeekDeepSeekDeepSeek V4 Pro (0813)OPEN80.2%
9GoogleGoogleGemini 3.5 Flash Lite70.9%
10GoogleGoogleGemini 3.6 Flash79.7%
11Zhipu AIZhipu AIGLM 5.2OPEN83.5%
12OpenAIOpenAIGPT-5.6 Luna84.4%
13OpenAIOpenAIGPT-5.6 Terra82.9%
14xAIxAIGrok 4.586.9%
15Thinking MachinesThinking MachinesInklingOPEN85.4%
16Thinking MachinesThinking MachinesInkling SmallOPEN84.1%
17Ant GroupAnt GroupLing 3.0 FlashOPEN80.9%
18MetaMetaMuse Spark 1.188.9%
19MetaMetaMuse Spark 1.290.1%
20AlibabaAlibabaQwen 3.8 MaxOPEN85.0%
Full Results Table
#ModelOverallSourceDate
🥇
AnthropicClaude Opus 5
Anthropic
91.0VALS (2026)2026-07
🥈
MetaMuse Spark 1.2
Meta
90.1VALS (2026)2026-08
🥉
MetaMuse Spark 1.1
Meta
88.9VALS (2026)2026-07
4
AnthropicClaude Fable 5
Anthropic
88.5VALS (2026)2026-06
5
Moonshot AIKimi K3OPEN
Moonshot AI
88.0VALS (2026)2026-07
6
xAIGrok 4.5
xAI
86.9VALS (2026)2026-07
7
xAIGrok 4.6
xAI
86.5VALS (2026)2026-08
8
Thinking MachinesInklingOPEN
Thinking Machines
85.4VALS (2026)2026-07
9
OpenAIGPT-5.6 Sol
OpenAI
85.2VALS (2026)2026-07
10
AlibabaQwen 3.8 MaxOPEN
Alibaba
85.0VALS (2026)2026-08
11
OpenAIGPT-5.6 Luna
OpenAI
84.4VALS (2026)2026-07
12
Thinking MachinesInkling SmallOPEN
Thinking Machines
84.1VALS (2026)2026-07
13
Zhipu AIGLM 5.2OPEN
Zhipu AI
83.5VALS (2026)2026-06
14
OpenAIGPT-5.6 Terra
OpenAI
82.9VALS (2026)2026-07
15
Ant GroupLing 3.0 FlashOPEN
Ant Group
80.9VALS (2026)2026-07
16
DeepSeekDeepSeek V4 FlashOPEN
DeepSeek
80.4VALS (2026)2026-07
17
DeepSeekDeepSeek V4 Pro (0813)OPEN
DeepSeek
80.2VALS (2026)2026-08
18
GoogleGemini 3.6 Flash
Google
79.7VALS (2026)2026-07
19
AnthropicClaude Sonnet 5
Anthropic
76.0VALS (2026)2026-06
20
GoogleGemini 3.5 Flash Lite
Google
70.9VALS (2026)2026-07