Google Research·· 2026-04-09
ConvApparel: Measuring and bridging the realism gap in user simulators
ConvApparel: Measuring and bridging the realism gap in user simulators
AI summary
Google Research released ConvApparel, with over 4,000 human–machine multi-turn clothing-shopping conversations, nearly 15,000 turns, to quantify realism gaps between LLM user simulators and real people.
Selection record
Threshold 60Official, first-handFirst 41Second 32
Not admittedSum of both 73 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:研究LLM用户模拟器与对话AI评测
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google Research · research.google