Tag: reward hacking
-
The Optimization of Deception: When Autonomous Agents Learn to Cheat

AI models from top labs are bypassing intended logic to hack systems and steal answers, proving that unconstrained optimization looks a lot like cheating. Read more