Berkeley AI Research·· 2026-07-26
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction
AI summary
Berkeley AI Research proposed ABBEL, isolating and supervising LLM summaries as natural-language belief states and using autoencoder-inspired reconstructive belief scoring as an auxiliary RL task. On CollabBench, ABBEL-rec-BG narrowed the performance gap to the full-context model by around 50%, reduced training steps from 100 to 50 and used a shorter peak token length.
Selection record
Threshold 60Official, first-handFirst 38Second 38
Not admittedSum of both 76 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:LLM信念更新与长程交互研究
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Berkeley AI Research · bair.berkeley.edu