Import AI 460: Reward hacking society, RSI data from Anthropic; and RL-based quadcopter racing
Import AI covers three studies. King's College London, Fudan and Alan Turing Institute built SocioHack with 72 sandbox social environments to test RL exploiting institutional loopholes while complying with rules. Models rediscovered patched historical vulnerabilities with 61.25% recall and 90.85% precision.