Computer Vision & Multimodal AI
Advancing visual understanding, vision-language models, object detection, and multimodal AI for real-world applications.
Overview
Work in this area covers how machines represent and reason about visual information, and how visual representations combine with other modalities such as text. Malaya AIR members publish on generative and diffusion models, deep hashing and retrieval, object and scene text detection, video understanding and industrial inspection. Several members work on both the vision and the language side of multimodal systems, which is why this area overlaps closely with the centre's natural language processing research and with its medical imaging work.
Topics
Researchers
8 academic members of Malaya AIR work in this area.
Related projects
Lingua Cognito
NLP and text processing, with inclusive multimodal systems for education and heritage.
Selected publications
Collaboration
Interested in research collaboration, technical exchange or translational opportunities related to this area?
Discuss a Collaboration