Vending-Bench 2
Andon Labs / Professional work · Run a simulated shop for a year, handling procurement, negotiation, inventory and pricing.
What it measures, and how
Compares end-of-period balances on the same V2 single-agent task, keeping variance across multiple runs; multiplayer competitive Arena, different tools and reasoning configurations are handled separately. Business results in the task are evaluation scores, and API prices do not count towards capability scores.
How this evidence is used
Under observation: the value of adding long-term operations and e-commerce capabilities is clear; a stable raw-score protocol, run configuration and permission for re-display are still missing. Profits may also come from using simulators, which cannot be directly equated with trustworthy operations.
Limits and data attribution
Under observation: the value of adding long-term operations and e-commerce capabilities is clear; a stable raw-score protocol, run configuration and permission for re-display are still missing. Profits may also come from using simulators, which cannot be directly equated with trustworthy operations.
Data licence: Leaderboard data usage boundaries to be confirmed
Scores are published by Andon Labs; the original scores and the News consensus score use different scales and cannot be added directly.