Autonomous driving systems traditionally consist of specialized components for perception, mapping, prediction, and planning. However, some modules like perception or planning might struggle in novel and complex scenarios. End-to-end (E2E) autonomous driving has emerged as a potential approach to address challenging scenarios. Current E2E systems combine large language models (LLMs) for scene understanding and visual data from the camera, however, lacking strong 3D spatial reasoning. Hence, fusing LiDAR data with camera images can enhance the ability of the E2E model to understand the spatial relationship and mitigate depth error.
Recent Papers:
Before a vehicle or robot can decide where to go, it first has to reliably understand what is around it. Our work in perception and multi-modal sensor data fusion addresses this by combining complementary sensors — cameras, LiDAR, and radar — into a single, resilient picture of the environment. The guiding premise is that no sensor is sufficient on its own: cameras capture rich texture but fail at night, LiDAR delivers precise 3D geometry but degrades in fog, and radar sees through bad weather yet only sparsely — so the goal is to use each sensor where it is most useful and combine the information as efficiently as possible.
Recent Papers:
Autonomous vehicles heavily rely on LiDAR sensors for perception tasks, where accurate intensity information is essential. However, obtaining real-world LiDAR data with intensity information is challenging and expensive. As a result, simulation has emerged as a promising alternative. Existing physics-based simulation approaches often oversimplify the complex relationship between LiDAR rays and the objects they interact with, leading to a large simulation-to-real gap. Hence, bridging the gap between simulated and real-world LiDAR intensity data is crucial for developing robust LiDAR perception algorithms in autonomous vehicles.
Recent Papers: