Skip to content
Evaluation sources

Hyper-τ-bench

Sierra / Coding · Builds and tests a working customer service agent from business materials.

Official evaluation
In NewsObserving
Evidence budgetNot scored
Upstream data as ofTo be confirmed
Last syncedNot collected yet

What it measures, and how

Measures agent development and debugging, kept separate from τ³-Banking, which directly performs customer service; development tools, run time and acceptance environment are fixed, and the highest score cannot be cherry-picked across systems.

How this evidence is used

Under observation: can add coding and business system development; raw scores, licence and fixed run conditions must first be reviewed; no automatic scoring just because of a new release.

Limits and data attribution

Under observation: can add coding and business system development; raw scores, licence and fixed run conditions must first be reviewed; no automatic scoring just because of a new release.

Data licence: Leaderboard data usage boundaries to be confirmed

Scores published by Sierra. Raw scores and the News consensus score use different scales and cannot be added directly.