Fahad Shahbaz Khan
MBZUAI, UAE ; Linköping University, Sweden
Biography
Fahad Khan is currently a Full Professor at the MBZUAI, Abu Dhabi, United Arab Emirates. He also holds a faculty position (Universitetslektor + Docent) at Computer Vision Laboratory, Linköping University, Sweden. He received the M.Sc. degree in Intelligent Systems Design from Chalmers University of Technology, Sweden and a Ph.D. degree in Computer Vision from Computer Vision Center Barcelona and Autonomous University of Barcelona, Spain. From 2012 to 2014, he was a postdoctoral fellow at Computer Vision Laboratory, Linköping University, Sweden. From 2014 to 2018, he was a research fellow at Linköping University, Sweden. In 2018, he was awarded the Docent title in computer vision from Linköping University, Sweden. He has achieved top ranks on various international challenges (Visual Object Tracking VOT: 1st 2014 and 2018, 2nd 2015, 1st 2016; VOT-TIR: 1st 2015 and 2016; OpenCV Tracking: 1st 2015; 1st PASCAL VOC Segmentation and Action Recognition tasks 2010). He received the best paper award in the computer vision track at IEEE ICPR 2016, best student paper award at VISAAP 2023, best paper finalist at CVPR 2022, best paper finalist at MICCAI 2024, and best student paper honorable mention at ACCV 2024. He has published over 150 reviewed conference papers, journal articles, and book contributions, with over 80,000 citations according to Google Scholar. His research interests include a wide range of topics within computer vision and machine learning. He serves as a regular senior program committee member for leading AI conferences such as, CVPR, ICCV, ECCV, and NeurIPS, will serve as a General Chair for ACCV 2028, and is an Associate Editor of leading AI journals such as, IEEE TPAMI and CVIU.
News
Recent news, preprint and papers (see publications):
+ Mbzuai-Oryx Generative AI Project (GLaMM, MobiLlama, Video-ChatGPT, AiN, LlamaV-O1, GeoChat): Git Repo
+ Four papers accepted at ECCV'26, Three papers accepted at ICLR'26, One paper accepted at ICML'26.
+ I will be an Area Chair (AC) for ICLR'26, ICML'26, AAAI'26, NeurIPS'26!
Selected Research Projects
𝐖𝐨𝐫𝐥𝐝 𝐌𝐨𝐝𝐞𝐥𝐬 powered by Diffusion Transformers are powerful, but painfully slow. In this work, we rethink 𝐡𝐨𝐰 𝐭𝐨 𝐚𝐜𝐜𝐞𝐥𝐞𝐫𝐚𝐭𝐞 𝐖𝐨𝐫𝐥𝐝 𝐌𝐨𝐝𝐞𝐥𝐬 𝐢𝐧𝐟𝐞𝐫𝐞𝐧𝐜𝐞 𝐰𝐢𝐭𝐡𝐨𝐮𝐭 𝐫𝐞𝐭𝐫𝐚𝐢𝐧𝐢𝐧𝐠. We propose WorldCache, a Perception-Constrained Dynamical Caching framework that improves both when and how to reuse features. WorldCache introduces motion-adaptive thresholds, saliency-weighted drift estimation, optimal approximation via blending and warping, and phase-aware threshold scheduling across diffusion steps. Our cohesive approach enables adaptive, motion-consistent feature reuse without retraining. The proposed approach achieves up to 3.0x speed-up across 𝐂𝐨𝐬𝐦𝐨𝐬, 𝐖𝐀𝐍, 𝐚𝐧𝐝 𝐃𝐫𝐞𝐚𝐦𝐃𝐨𝐣𝐨, with ~99% quality retention.
How can we make 𝐋𝐌𝐌𝐬 𝐭𝐨 𝐮𝐧𝐝𝐞𝐫𝐬𝐭𝐚𝐧𝐝 𝐚𝐧𝐝 𝐫𝐞𝐚𝐬𝐨𝐧 𝐚𝐛𝐨𝐮𝐭 𝐩𝐡𝐲𝐬𝐢𝐜𝐚𝐥 𝟑𝐃 𝐰𝐨𝐫𝐥𝐝 𝐞𝐯𝐞𝐧 𝐚𝐭 𝐚 𝐟𝐢𝐧𝐞-𝐠𝐫𝐚𝐢𝐧𝐞𝐝 𝐥𝐞𝐯𝐞𝐥? Natural-language queries about 3D environments become actionable when responses are verifiable and metric. Verifiability requires explicit grounding to the referred 3D region, while metric answers report physical measurements in real-world units (e.g., size, thickness, clearance, and distance). We present Ground3D-LMM, a unified model that takes a point cloud and an optional RGB image as input and supports 3D spatial conversation with (i) point-grounded responses and (ii) metric numeric outputs at both object and part granularity, including multi-object queries. To evaluate this intersection of grounding and measurement, we define the 3D Grounded Measurement task, which requires predicting the referred 3D region and the corresponding metric quantities in real-world units. We introduce a large-scale dataset with dense object and part annotations and roughly 2.5M question-answer pairs spanning eight tasks, with a manually verified test set. Extensive experiments on multiple datasets and tasks show that our proposed Ground3D-LMM model provides a strong baseline for grounded, metric-aware 3D conversational understanding.
Existing 𝗟𝗮𝗿𝗴𝗲 𝗠𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹 𝗠𝗼𝗱𝗲𝗹𝘀 (𝗟𝗠𝗠𝘀) still rely heavily on 𝗹𝗮𝗻𝗴𝘂𝗮𝗴𝗲 𝗽𝗿𝗶𝗼𝗿𝘀 rather than truly attending to the 𝘃𝗶𝘀𝘂𝗮𝗹 𝗶𝗻𝗽𝘂𝘁. We propose VISE (Visual Invariance Self-Evolution), a purely unsupervised self-evolving framework that directly regularizes the model's visual conditioning policy through two complementary invariance-based rewards: a geometric invariance reward that enforces spatial consistency under known transformations, and a semantic invariance reward that penalizes evidence-agnostic generation by requiring the model to recognize the absence of evidence when predicted regions are perturbed. VISE operates within a single model without specialist roles, external reward models, or annotations, and is trained on raw unlabeled images. Experiments on 18 benchmarks demonstrate the efficacy of our approach.
we introduce Agent-X, a large-scale benchmark for evaluating vision-centric agents’ multi-step and deep reasoning capabilities in real-world, multimodal settings. Agent-X features 828 agentic tasks with authentic visual contexts, including images, multi-image comparisons, videos, and instructional text. These tasks span six major agentic environments: general visual reasoning, web browsing, security and surveillance, autonomous driving, sports, and math reasoning. Our benchmark requires agents to integrate tool use with explicit, stepwise decision-making in these diverse settings. In addition, we propose a fine-grained, step-level evaluation framework that assesses the correctness and logical coherence of each reasoning step and the effectiveness of tool usage throughout the task. Our results reveal that even the best-performing models, including GPT, Gemini, and Qwen families, struggle to solve multi-step vision tasks.
Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model's ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively.
Recent advancements in Large Vision-Language Models (VLMs) have shown great promise in natural image domains, allowing users to hold a dialogue about given visual content. However, such general-domain VLMs perform poorly for Remote Sensing (RS) scenarios. We propose GeoChat - the first versatile remote sensing VLM that offers multitask conversational capabilities with high-resolution RS images. Specifically, GeoChat can not only answer image-level queries but also accepts region inputs to hold region-specific dialogue. Furthermore, it can visually ground objects in its responses by referring to their spatial coordinates. GeoChat demonstrates robust zero-shot performance on various RS tasks, e.g., image and region captioning, visual question answering, scene classification, visually grounded conversations and referring detection.