
What is reward hacking? Definition, examples, and detection
Reward hacking, defined: how coding agents game their own tests, the documented hack patterns, and how judge-free detection compares to an LLM judge on our public 366-task benchmark.
Private beta open for teams shipping AI agents.·See it catch a reward-hack

What is reward hacking? Definition, examples, and detection
Reward hacking, defined: how coding agents game their own tests, the documented hack patterns, and how judge-free detection compares to an LLM judge on our public 366-task benchmark.

Shipworthy: the certified release gate for AI agents
SHIP / LIMIT / BLOCK verdicts on every agent update, from a judge-free deterministic core — now in private beta.

Changing only the reward: −59–65% reward hacking in GRPO training
Swapping a naive extensional reward for an isomorphic IPT reward — and changing nothing else — cut reward hacking by 59–65% across two Llama models.
Subscribe to our newsletterSubscribe to our newsletter for release notes
“Every number on this site traces to a reproducible run.”
Natural emergent misalignment from reward hacking in production RL
Anthropic · Nov 2025
Read on arXiv ↗Reward hacking is swamping model intelligence gains
Cursor · Jun 2026
Read on cursor.com ↗Recent frontier models are reward hacking (o3 on RE-Bench)
METR · Jun 2025
Read on metr.org ↗Isomorphic Perturbation Testing — the formulation our gate extends
Helff et al., ICLR 2026 · 2026
Read on arXiv ↗External sources we cite in our posts — independent work, not press about Verifiable Labs.