Vision-language models (VLMs) are considered augmentations of text-only large language models (LLMs) with additional access to the visual world. However, it is known that VLMs typically perform worse than text-only LLMs on reasoning tasks that involve only text (e.g., math word problems), raising the question of whether the benefits of visual augmentation justify the loss […]
Read More