Understanding and improving compositional generalization in vision-language models

Compositionality — the ability to understand new combinations of familiar parts — is a core aspect of human cognition and reasoning, but current vision-language models such as CLIP still struggle with it. These limitations reduce the reliability of systems that must correctly interpret objects, attributes, and relations in real-world scenarios. Although many methods have been proposed to strengthen CLIP’s compositional abilities, it remains unclear which parts of the model they affect or why certain approaches lead to better results. This project will develop a standardized diagnostic framework to identify where compositional behaviour improves within the models and what factors drive those improvements. The resulting insights and tools will help guide the development of more reliable multimodal AI systems. This project will benefit both participating institutions by supporting their shared interests in advancing reliable, scalable AI methods, fostering a new research collaboration between the groups, and creating reusable resources that will support future research and student training.

Faculty Supervisor:

Martin Ester

Student:

Partner:

Korea Advanced Institute of Science and Technology

Discipline:

Computer science

Sector:

Artificial Intelligence; Technology

University:

Simon Fraser University

Program:

Globalink Research Award

Current openings

Find the perfect opportunity to put your academic skills and knowledge into practice!

Find Projects