Berkeley AI Research·· 2025-11-01
RL without TD learning
RL without TD learning
AI summary
Berkeley AI Research proposed a divide-and-conquer off-policy reinforcement learning algorithm without temporal-difference (TD) learning. It reduces Bellman recursions from linear to logarithmic counts, scaling to long-horizon tasks.
Selection record
Threshold 60Official, first-handFirst 31Second 34
Not admittedSum of both 65 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:强化学习算法研究,属AI技术
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Berkeley AI Research · bair.berkeley.edu