Mitigating hallucination in vision language models for autonomous systems

Autonomous systems such as intelligent vehicles and robotic platforms must reliably perceive their surroundings and reason about complex environments to make safe, informed decisions. Recent advances in Vision-Language Models (VLMs) and Vision-Language-Action (VLA) frameworks combine visual understanding with language-based reasoning, creating unified architectures for perception, interpretation, and decision-making. This integration enables autonomous agents to describe scenes, reason about spatial and temporal context, and translate perception into executable actions, improving adaptability and interpretability. However, despite strong performance on benchmarks, VLMs and VLAs remain prone to hallucination, producing confident yet incorrect outputs that can compromise safety. In reasoning tasks, hallucinations may cause scene misinterpretation, while in action-oriented systems they can lead to flawed control decisions. This project will develop and evaluate novel multimodal VLM architectures that assess and verify their own reasoning before generating outputs or actions. By embedding internal verification and grounding mechanisms, the research aims to mitigate hallucination at its source, enhancing the reliability and transparency of model reasoning. The collaboration will strengthen institutional leadership in explainable and dependable AI, expanding expertise in reliable multimodal learning and advancing the development of safer autonomous technologies.

Faculty Supervisor:

Krzysztof Czarnecki

Student:

Partner:

Stanford University

Discipline:

Computer science

Sector:

Education

University:

University of Waterloo

Program:

Globalink Research Award

Current openings

Find the perfect opportunity to put your academic skills and knowledge into practice!

Find Projects