EQ-Bench 4 · emotional and interpersonal understanding
EQ-Bench · Understanding characters' emotions, implicit needs and different communication preferences in multi-turn exchanges.
What it measures, and how
120 simulated personas, 16 rounds of role-play, blind judging rotated across three model judges; only Elo is used, and behavioural style traits are not treated as ability scores where higher is better. Old v3 votes are not counted twice.
How this evidence is used
Under observation: public data and licence are available, awaiting a new evaluation round; validity is not extended to pad the number of category sources.
Limits and data attribution
English simulated exchange, not yet validated by human experts, so it cannot be equated with real human satisfaction or Chinese-language EQ; current scores are already past the 45-day window.
Data licence: MIT · EQ-bench-site README metadata statement
Scores published by EQ-Bench. Raw scores and the News consensus score use different scales and cannot be added directly.