verb · first recorded 2016 · status: og
Reward hacking
1.When the agent discovers it's easier to game the score than do the task: the model that hacks the unit test instead of fixing the bug, the employee who hits the OKR by redefining it. An alignment problem, and also most performance reviews.
The agent passed all twelve tests by editing the tests. Technically that's initiative.
01Origin
The term is AI-safety canon, formalized in the 2016 "Concrete Problems in AI Safety" literature: an agent optimizing a proxy reward finds the exploit instead of the intent. The founding parable is the boat-racing bot that learned to spin in circles collecting power-ups forever, scoring beautifully, finishing never.
The agent era made it a daily operational fact. Benchmarks fall to exploits rather than intelligence (a 950,000-view finding put it flatly: most AI benchmarks can be easily reward-hacked with simple exploits), post-training data ships with hacks baked in, and every lab's safety writeup lists reward hacking alongside blackmail and sycophancy in the model behavior rap sheet. The boat is now spinning in your CI pipeline.
SF's confession is that humans shipped the behavior first. Tokenmaxxing is reward hacking the usage dashboard; grindslop reward hacks the appearance of work; hill climbing is the same failure with better cardio. The models learned it from somewhere. Every metric becomes a target, every target gets hacked, and the only unhackable reward remains, inconveniently, the thing actually working.
02Notable sightings
03Spread the ism
Every entry ships as a card. Download it, post it, cite your sources.