NeurIPS 2026 Workshop · Sydney, Australia · December 11 or 12, 2026 · In person
Physical understanding is the missing bridge from foundation and world models to reliable decision-making agents.
Paper Submission Deadline: August 26th, 2026, 08:00 UTC
AI systems are increasingly deployed in physical settings such as robotics, autonomous driving, and laboratory automation. In these settings, actions must respect physical laws. Objects have mass, contacts transmit forces, constraints limit feasibility, and interventions produce irreversible consequences. An agent that lacks physical understanding (the ability to reason about objects, forces, affordances, causal mechanisms, and long-horizon interaction) cannot act safely or reliably, no matter how accurately it predicts the next video frame.
Foundation models for decision-making, including vision-language-action models, robot foundation models, agentic systems, and interactive world models, have sharply improved long-horizon video generation, embodied policy learning, and multimodal action grounding. Yet most of these models are still trained and evaluated primarily for visual realism or short-term prediction accuracy. The field now stands at an inflection point. Foundation and world models are powerful enough to serve as the backbone of physical agents, yet the community lacks shared definitions, evaluation protocols, and benchmarks for physical understanding as a distinct and measurable capability.
This workshop brings together researchers from machine learning, reinforcement learning, robotics, computer vision, simulation, causality, embodied AI, autonomous driving, and safety around one central question: how can AI systems understand the physical world well enough to make reliable decisions within it?
We welcome theoretical, algorithmic, empirical, benchmark, systems, and position contributions across four themes.
Modeling Physical Systems. Object- and scene-centric dynamics, contact, friction, fluids, deformable objects, material properties, 3D and 4D world models, neural simulation, and video-to-physics. Models that capture actionable physical structure rather than appearance alone.
Physical Perception and Active Interaction. Acting to reduce uncertainty, learning affordances, and discovering physical constraints under partial observability. How agents can use interaction to build richer physical representations.
Causal and Counterfactual Physical Reasoning. Interventions, mechanisms, counterfactual outcomes, stability analysis, and distribution shift. Can models reason about what would happen under actions they have not yet taken?
Decision-Making, Evaluation, and Deployment. Planning, model-based RL, control, manipulation, navigation, and long-horizon execution. Evaluation methodology covering controllability, calibration, causal validity, sim-to-real transfer, and robustness, plus deployment concerns such as hallucinated actions, constraint satisfaction, risk estimation, and failure recovery.