L2M-Augmented Reality-Based Training System using Multimodal Language Model for Context-Aware Guidance and Activity Recognition in Complex Machine Operations

Many industrial companies still rely on traditional training methods and struggle to keep up with evolving skill requirements. These conventional approaches, such as manuals, videos, and classroom instruction, are ineffective in delivering the hands-on skills required to operate complex machinery.
Our project introduces an Augmented Reality (AR) based training platform powered by Multi-Large Language Models (MLLMs) that acts as an intelligent instructor. It can understand what the user is doing, read the machine’s feedback, and guide the user directly on the equipment. Unlike existing AR systems that follow a fixed, pre-scripted path, our system continuously adapts to user actions and updates the instructions automatically, allowing trainees to learn safely and independently without constant supervision.
The system integrates structured prompt design, model-target detection, and MLLM-based reasoning to interpret visual and textual cues in real time. It can also be easily customized for different machine types and industry-specific workflows, enabling rapid deployment across diverse applications. These capabilities position the solution as technically innovative and practically scalable, bridging the gap between AI-driven research and real-world machine-operation training.

Faculty Supervisor:

Qingjin Peng

Student:

Partner:

North Forge

Discipline:

Engineering

Sector:

Professional, scientific and technical services

University:

University of Manitoba

Program:

Business Strategy Internship

Current openings

Find the perfect opportunity to put your academic skills and knowledge into practice!

Find Projects