2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks2026-09-10MODELDeepSeek-V4.1-Flash released with extreme agent memory efficiency and evaluated across VALS2026-09-03MODELGPT-6 Astra released by OpenAI - 71.7% on EMB, 63.3% on Tax Agent Bench2026-09-02MODELGoogle Gemini 3.8 Flash takes #1 rank on Finance Agent (v2) with 61.4% and 72.2% on EMB2026-09-02MODELMeta releases Muse Spark 1.3 & 1.3 Max - 59.0%-60.0% on Finance Agent (v2)2026-09-01MODELClaude Fable 5.1 leads Tax Agent Bench (77.6%) and EMB (76.7%) with 75% cheaper cache reads2026-09-11BENCHTax Agent Bench and FAB v2 benchmarks updated with latest frontier model rankings2026-09-15UPDATEPlatform updated: 43 models tracked across 41 independent benchmarks
Open evaluation · Free submission

Submit to FinanceBenchmark

Two paths. Submit a score you've already run, or request a certified evaluation from us against the full benchmark suite.

Free
Path 1
Submit a score
You've run a model against one of our tracked benchmarks and have a reproducible result to share.
Free for research
Path 2
Request certified evaluation
We run your AI system against our full benchmark suite and publish a certified score with full methodology documentation.
Score submission
%
We do not accept self-reported scores without a verifiable source. Link to a GitHub repo, Colab notebook, arXiv paper, or the open evaluation harness run.
We'll email you when the score is published or if we need clarification.
Submissions reviewed within 5 business days. We may re-run the evaluation to verify.
What happens after you submit
1
We receive your submission
Your score and source link go to submit@financebenchmark.ai for review.
2
We verify the result
We may re-run the evaluation independently. If we can't reproduce the result, we'll contact you.
3
Published within 5 days
Verified scores are added to the leaderboard with full attribution and a link to your source.
Requirements
Source link must be publicly accessible — GitHub repo, notebook, or arXiv paper with evaluation code.
Evaluation must use the standard protocol for that benchmark. Deviations must be documented in the notes.
Score must be numeric (0–100). Partial task scores should be averaged using the benchmark's official weighting.
We do not accept scores run against modified versions of the benchmark dataset.
Certified evaluations are currently free for research institutions and open-source models. Commercial systems: contact us to discuss terms before submitting the form.