Skip to content
Berkeley AI Research·· 2026-07-26

Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

Teaching LLMs to Update Beliefs for Efficient Long-Horizon Interaction

AI summary

Berkeley AI Research proposed ABBEL, isolating and supervising LLM summaries as natural-language belief states and using autoencoder-inspired reconstructive belief scoring as an auxiliary RL task. On CollabBench, ABBEL-rec-BG narrowed the performance gap to the full-context model by around 50%, reduced training steps from 100 to 50 and used a shorter peak token length.

Selection record

Not admittedSum of both 76 < twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:LLM信念更新与长程交互研究

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Berkeley AI Research · bair.berkeley.edu