Vals Finance Agent
Vals AI / Professional work · Everyday office tasks for financial analysts, to see how work such as research reports and modelling is done.
What it measures, and how
Only the official Overall accuracy of Finance Agent v2 is used.
How this evidence is used
3% of the industry professional tasks budget; currently tests finance directly and does not claim to cover all industries.
Results
Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.
| Rank there | Model there | Raw score | Representative configuration |
|---|---|---|---|
| 1 | google/gemini-4-argonGoogle | 65.4% | High reasoning |
| 2 | google/gemini-3.8-flashGoogle | 61.4% | High reasoning |
| 3 | meta/muse_spark_1_2Meta | 60.6% | xHigh reasoning |
| 4 | meta/muse_spark_1_3_maxMeta | 60.0% | Max reasoning |
| 5 | google/gemini-3.7-flashGoogle | 59.0% | High reasoning |
| 7 | anthropic/claude-fable-5-1Anthropic | 58.9% | Max reasoning |
| 8 | anthropic/claude-opus-5Anthropic | 58.6% | Max reasoning |
| 9 | anthropic/claude-opus-5-5Anthropic | 58.6% | Max reasoning |
| 10 | anthropic/claude-sonnet-5-5anthropic | 58.1% | Max reasoning |
| 11 | google/gemini-3.5-flashGoogle | 57.9% | High reasoning |
| 12 | zai/glm-5.3-flashZ.ai | 57.9% | Max reasoning |
| 13 | xiaomi/mimo-v2.6-proXiaomi | 57.3% | Source default configuration |
| 14 | meta/muse_spark_1_1Meta | 57.2% | xHigh reasoning |
| 15 | anthropic/claude-fable-5Anthropic | 56.3% | Max reasoning |
| 16 | google/gemini-3.6-flashGoogle | 56.3% | High reasoning |
| 17 | xiaomi/mimo-v2.6-flashXiaomi | 56.3% | Source default configuration |
| 18 | zai/glm-5.3Z.ai | 55.8% | Max reasoning |
| 19 | tencent/hy4-previewTencent | 55.1% | Source default configuration |
| 20 | openai/gpt-5.6-lunaOpenAI | 55.0% | Max reasoning |
| 21 | ant/ling-3.0-flash-af-rc3Ant | 54.9% | Source default configuration |
| 22 | openai/gpt-5.6-terraOpenAI | 54.4% | Max reasoning |
| 23 | kimi/kimi-k3Moonshot AI | 54.4% | Source default configuration |
| 24 | anthropic/claude-opus-4-8Anthropic | 53.9% | Max reasoning |
| 25 | anthropic/claude-sonnet-5Anthropic | 53.9% | Max reasoning |
| 26 | openai/gpt-5.6-solOpenAI | 53.8% | Max reasoning |
| 27 | grok/grok-4.6xAI | 53.7% | High reasoning |
| 28 | openai/gpt-6-astraOpenAI | 53.5% | Max reasoning |
| 29 | deepseek/deepseek-v4.1-flashDeepSeek | 53.5% | High reasoning |
| 30 | grok/grok-4.7xAI | 52.3% | xHigh reasoning |
| 31 | openai/gpt-6.1-sol— | 52.0% | Max reasoning |
Limits and data attribution
Private questions cannot be recalculated item by item; it belongs to office tasks like APEX and must share one budget.
Data licence: Official public results; public re-display retains official attribution, and full redistribution authorisation remains subject to Vals terms
Scores published by Vals AI. Raw scores and the News consensus score use different scales and cannot be added directly.