Multimodal AI Lab
@ Inha University
@ Inha University
Notice
If you are interested in joining our team, please contact us at pilhyeon.lee@inha.ac.kr. We are currently recruiting for the following positions.
[MS/PhD]
We are looking for prospective MS/PhD students with a strong interest in video understanding (e.g., retrieval, grounding, captioning) and multimodal learning (e.g., cross-modal transfer, multimodal robustness). Positions are open for 2027-Spring and Fall admission.
[Intern/Undergrad]
We plan to run a research internship program during the 2026-Winter semester. There are no restrictions on research topics, and you are welcome to explore any area of interest. If you are interested, we encourage you to contact us before applying.
Our research focuses primarily on understanding and how to interactively fuse multimodal information from diverse sources such as images, videos, text, audio, and etc. Specfically, our research topics cover various problems of Computer Vision (CV), Natural Languge Processing (NLP), and Signal Processing (SP), which includes (but not limited to):
Vision & Language
Scene Understanding (e.g., detection, segmentation)
Cross-modal Retrieval (e.g., text-to-image, audio-to-video)
Video Understanding
Human Behavior Analysis
Knowledge Distillation
Large-scale Foundation Models (e.g., LLMs, LVMs)
Generative Models
Model Robustness
Semi- / Weakly-supervised Learning
Model Debiasing & Fairness