Recent advances in AI-Generated Content (AIGC), including diffusion and flow-based generative models, have enabled impressive progress in video synthesis, dynamic scene generation, and emerging world-modeling systems. Despite remarkable visual realism, most existing models still lack explicit physical grounding and frequently produce temporally inconsistent motions, implausible interactions, and violations of fundamental physical laws such as gravity, momentum, and collision constraints. These limitations significantly restrict the applicability of AIGC in emerging multimedia scenarios that require stable, long-horizon, and physically coherent dynamics, including XR/VR, digital twins, robotics simulation, and interactive virtual environments.
This workshop, Physics-Driven AIGC: Physically-Consistent Video, 4D Scene, and World Generation, aims to establish a dedicated forum for exploring generative models that integrate physical priors, dynamic constraints, and causal structures into multimedia content creation. The workshop focuses on bridging generative modeling with physics-informed reasoning, addressing challenges in dynamics-consistent video generation, 4D scene and world modeling, and physically grounded evaluation. Topics include physics-driven diffusion and flow-matching models, physical constraints and energy consistency, scene evolution and human–object interaction, physics-informed score distillation, efficient physics-aware AIGC, and benchmarks for physical consistency.
By bringing together researchers from multimedia, computer vision, graphics, simulation, and machine learning, this workshop seeks to foster cross-disciplinary collaboration and advance the development of physically-consistent, reliable, and controllable AIGC systems. As the first ICME workshop explicitly centered on physics-driven AIGC, it fills a critical gap in the current landscape and aligns closely with ICME’s mission to advance multimedia technologies, systems, and applications.
Please submit your work at https://cmt3.research.microsoft.com/IEEEICMEW2026/Track/12/Submission
Submission Deadline: Apr 15 2026
Notification Date: Apr 25 2026
Workshop Date: Jul 05 2026, 09:30-12:30
Researchers are welcome to submit their work in this workshop. More details will be announced soon.
Committee members:
Dr. Ping Liu, University of Nevada, Reno, pingl at unr.edu
Dr. Yawei Luo, Zhejiang University, yaweiluo at zju.edu.cn
Dr. Feifei Shao, Zhejiang University, sff at zju.edu.cn
Dr. Jun Xiao, Zhejiang University, junx at cs.zju.edu.cn
Title: The Paradigm Shift of 3D Generation
Speaker: Dr. Wei Yang, Associate Professor, Huazhong University of Science and Technology (HUST)
The automated creation of 3D content is undergoing a profound transformation. As digital and physical realities continue to converge, the ability to efficiently generate high-quality 3D assets has become a key enabling technology for applications including virtual reality, gaming, digital twins, manufacturing, and spatial computing. This keynote reviews the evolution of 3D generation and highlights the recent paradigm shift that is redefining the field.
Early approaches addressed the scarcity of 3D training data by leveraging large-scale 2D image priors. Methods such as Score Distillation Sampling (SDS), introduced by DreamFusion, and multi-view image diffusion models successfully bridged the gap from 2D to 3D generation. As the demand for higher geometric fidelity and structural consistency has increased, the field has rapidly shifted toward native 3D generation. The talk will introduce the fundamental techniques behind this transition, with a particular focus on representative frameworks including 3DShape2VecSet, CLAY, and Trellis, which directly operate on 3D representations and achieve substantial improvements in quality, scalability, and efficiency.
Building upon these native approaches, the keynote will further discuss recent extensions and downstream applications that demonstrate their versatility across diverse scenarios. Finally, it will examine the major open challenges that remain, including production-quality mesh generation, physically accurate material synthesis, and the extension from static 3D content to dynamic 4D animation. The talk aims to provide a comprehensive overview of the past, present, and future directions of 3D generation research.
Dr. Wei Yang is an Associate Professor in the School of Computer Science at Huazhong University of Science and Technology (HUST), where he co-leads the HUST Media Lab. His research interests include computational imaging, 3D vision, computer graphics, and 3D generative modeling. His recent work on native 3D generation has received broad recognition, including CLAY, which received an Honorable Mention Award at SIGGRAPH 2024, and CAST, which won the Best Paper Award at SIGGRAPH 2025.
Before joining HUST, Dr. Yang worked at Google ATAP on advanced sensing and on-device intelligence, and later served as a Principal Scientist at DGene, where he led the development of real-time volumetric human capture systems. He received his Ph.D. in Computer Science from the University of Delaware in 2017 under the supervision of Prof. Jingyi Yu. Dr. Yang is an active member of the computer vision and machine learning communities and regularly serves as an Area Chair for leading conferences including CVPR, ICCV, NeurIPS, and ICML.
Title: From Multimodal Generative Models to Unified World Modeling
Speaker: Dr. Ziwei Liu, Associate Professor (Provost’s Chair in AI), Nanyang Technological University (NTU), Singapore
Recent advances in multimodal generative AI have significantly improved machines’ ability to perceive and generate visual content. However, true intelligence requires models that can understand, predict, and interact with the physical world. This keynote presents a vision of Unified World Modeling, which extends multimodal generative models with physical understanding, dynamic scene modeling, and actionable reasoning. The talk will highlight recent progress toward AI systems that jointly model geometry, materials, motion, and interaction, enabling richer representations of the physical world and laying the foundation for next-generation embodied AI, robotics, and autonomous agents.
Dr. Ziwei Liu is currently an Associate Professor (Provost’s Chair in AI) at Nanyang Technological University (NTU), Singapore. His research focuses on computer vision, machine learning, and generative AI. He has published extensively in leading conferences and journals, including CVPR, ICCV, ECCV, NeurIPS, ICLR, SIGGRAPH, TPAMI, TOG, and Nature Machine Intelligence.
Dr. Liu is the recipient of numerous prestigious honors, including the PAMI Mark Everingham Prize, CVPR Best Paper Award Candidate, Asian Young Scientist Fellowship, International Congress of Basic Science Frontiers of Science Award, MIT Technology Review Innovators Under 35 Asia Pacific, and the Singapore President’s Young Scientist Award. He regularly serves as an Area Chair for CVPR, ICCV, ECCV, NeurIPS, and ICLR, and is an Associate Editor of IEEE TPAMI and IJCV.