The main focus of this research is to develop representation learning architectures and algorithms that can help perform various multimodal understanding tasks, and at the same time reduce the need for human supervision in the form of costly annotations. To achieve this goal, a learning system must be able to: (1) learn new tasks or […]
Read More