2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks
Tier 3agentTasksnumerical

RSI Index

Recursive Self-Improvement and recursive cognitive task orchestration index.

Variable tasks
3 models evaluated
Updated 2026
30.9%
Top Score
Top model
Claude Opus 5
Model Rankings
financebenchmark.ai · RSI Index
Anthropic
OpenAI
Moonshot AI
22.1%
🥇AnthropicAnthropicClaude Opus 530.9%
🥈OpenAIOpenAIGPT-5.6 Sol28.4%
🥉Moonshot AIMoonshot AIKimi K3OPEN22.1%
Full Results Table
#ModelOverallSourceDate
🥇
AnthropicClaude Opus 5
Anthropic
30.9VALS (2026)2026-07
🥈
OpenAIGPT-5.6 Sol
OpenAI
28.4VALS (2026)2026-07
🥉
Moonshot AIKimi K3OPEN
Moonshot AI
22.1VALS (2026)2026-07