Hyper-τ-bench
Sierra / Coding · Builds and tests a working customer service agent from business materials.
What it measures, and how
Measures agent development and debugging, kept separate from τ³-Banking, which directly performs customer service; development tools, run time and acceptance environment are fixed, and the highest score cannot be cherry-picked across systems.
How this evidence is used
Under observation: can add coding and business system development; raw scores, licence and fixed run conditions must first be reviewed; no automatic scoring just because of a new release.
Limits and data attribution
Under observation: can add coding and business system development; raw scores, licence and fixed run conditions must first be reviewed; no automatic scoring just because of a new release.
Data licence: Leaderboard data usage boundaries to be confirmed
Scores published by Sierra. Raw scores and the News consensus score use different scales and cannot be added directly.