Multimodal large language models can process text, images, and other forms of information, but do they perform consistently when the same content is presented in different forms? Researchers at the University of Amsterdam investigated this and found systematic inconsistencies with implications for how reliable these models are in practice. Their paper has been accepted at CVPR 2026.
In the paper “Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs,” Angela van Sprang, Laurens Samson, Ana Lucic, Erman Acar, Sennay Ghebreab, and Yuki M. Asano examine how multimodal language models respond when equivalent material is presented across different modalities. The finding is that these models regularly reach different conclusions, even when the underlying information is the same.
This cross-modal behavior is relevant for applications where consistency is critical, such as medical image processing, financial document analysis, or legal information systems. The researchers place their findings in the broader context of AI system robustness and reliability, contributing to a more realistic understanding of the limits of current multimodal models. The paper was presented at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026).