Skip to content
Evaluation sources

Vals Finance Agent

Vals AI / Professional work · Everyday office tasks for financial analysts, to see how work such as research reports and modelling is done.

Official evaluation
In NewsRanked
Evidence budget3%
Upstream data as of09/29 08:00
Last synced10/01 20:05

What it measures, and how

Only the official Overall accuracy of Finance Agent v2 is used.

How this evidence is used

3% of the industry professional tasks budget; currently tests finance directly and does not claim to cover all industries.

Results

Under fixed rules, each public model uses one representative configuration. Anonymous test variants are not shown, original leaderboard ranks are kept, and each item lists at most 30.

Rank thereModel thereRaw scoreRepresentative configuration
1google/gemini-4-argonGoogle65.4%High reasoning
2google/gemini-3.8-flashGoogle61.4%High reasoning
3meta/muse_spark_1_2Meta60.6%xHigh reasoning
4meta/muse_spark_1_3_maxMeta60.0%Max reasoning
5google/gemini-3.7-flashGoogle59.0%High reasoning
7anthropic/claude-fable-5-1Anthropic58.9%Max reasoning
8anthropic/claude-opus-5Anthropic58.6%Max reasoning
9anthropic/claude-opus-5-5Anthropic58.6%Max reasoning
10anthropic/claude-sonnet-5-5anthropic58.1%Max reasoning
11google/gemini-3.5-flashGoogle57.9%High reasoning
12zai/glm-5.3-flashZ.ai57.9%Max reasoning
13xiaomi/mimo-v2.6-proXiaomi57.3%Source default configuration
14meta/muse_spark_1_1Meta57.2%xHigh reasoning
15anthropic/claude-fable-5Anthropic56.3%Max reasoning
16google/gemini-3.6-flashGoogle56.3%High reasoning
17xiaomi/mimo-v2.6-flashXiaomi56.3%Source default configuration
18zai/glm-5.3Z.ai55.8%Max reasoning
19tencent/hy4-previewTencent55.1%Source default configuration
20openai/gpt-5.6-lunaOpenAI55.0%Max reasoning
21ant/ling-3.0-flash-af-rc3Ant54.9%Source default configuration
22openai/gpt-5.6-terraOpenAI54.4%Max reasoning
23kimi/kimi-k3Moonshot AI54.4%Source default configuration
24anthropic/claude-opus-4-8Anthropic53.9%Max reasoning
25anthropic/claude-sonnet-5Anthropic53.9%Max reasoning
26openai/gpt-5.6-solOpenAI53.8%Max reasoning
27grok/grok-4.6xAI53.7%High reasoning
28openai/gpt-6-astraOpenAI53.5%Max reasoning
29deepseek/deepseek-v4.1-flashDeepSeek53.5%High reasoning
30grok/grok-4.7xAI52.3%xHigh reasoning
31openai/gpt-6.1-sol—52.0%Max reasoning
Limits and data attribution

Private questions cannot be recalculated item by item; it belongs to office tasks like APEX and must share one budget.

Data licence: Official public results; public re-display retains official attribution, and full redistribution authorisation remains subject to Vals terms

Scores published by Vals AI. Raw scores and the News consensus score use different scales and cannot be added directly.