Due to the maturity of large language model (LLM) technology, to make robots generate proper actions according to humans’ language commands and robots’ vision, Vision-Language-Action (VLA) is seen as the key technology of extension from LLM to actions. However, the computational and data resource of VLA is huge. This project proposes Vision-Language-Submodular-Action (VLSA). It is to develop VLA, inverse VLA and transfer learning technologies for arm placement, finger motion missions in different environmental and size tasks via submodularity. This project is to boost the robot learning speed and performance via the near-optimality, sparsity and transferability of submodular functions.
The goal of this research is to explore key issues of VLA:
(1) VLSA: After adding submodularity to VLA, what’s the learning speed? What’s the theoretical guarantee?
(2) Inverse VLSA: To learn experts’ reward functions via submodularity, what’s the learning speed? What’s the theoretical guarantee?
(3) Transfer learning: After utilizing transfer learning of submodular functions, what’s the learning speed? What’s the theoretical guarantee?
With the rapid advancement of NVIDIA digital twin technology, robots are becoming capable of learning highly complex skills through simulation before being deployed in the real world. The key idea is to construct a simulation environment that closely resembles the physical world, creating a digital twin of the real environment. Robots are first trained in the simulator with GPU-accelerated learning, and the learned policies are then transferred to real robotic platforms. This simulation-to-real (Sim-to-Real) paradigm not only reduces the cost of training but also minimizes the risks associated with learning directly in real-world environments. In this project, we will leverage digital twin technology to develop the YO-KAI Intelligent Chef, enabling the robot to acquire sophisticated cooking skills in simulation before deploying them to a real robotic system.
This project aims to develop a vision-inertial multi-UAV search system for deployment on a chip in complex three-dimensional environments. Multi-UAV inspection faces three major challenges: maximal coverage, routing, and task allocation. To address the issues of insufficient coverage, high path costs, and unbalanced workloads, this project adopts the Multi-Robot Search with Matroid Constraints (MRSM) algorithm. Given a pre-constructed 3D map, each UAV independently constructs a spanning tree to perform efficient inspection planning with theoretical performance guarantees. The proposed algorithm will be implemented using ROS 2 and deployed on a domestically developed chip for performance validation.
The AI community has been paying more attention to multi-robot informative path planning (MIPP). MIPP is to plan trajectories for robots to maximize information gathering and subject to path costs. Potential applications include cooperative map exploration, spatial search, disinfection robots and distributed mobile charging stations etc. Finding optimal solutions for the aforementioned problems are to solving three NP-hard problems, the multi-robot assignment problem, the maximal coverage problem and travelling salesman problem (TSP). To make a breakthrough of the MIPP research status, this research utilizes the submodularity of the submodular tree and matroid to boost the theoretical guarantees. This research further analyzes MIPP via Fourier methods to find its invariance. The goal of this research is to explore key issues of MIPP:
(1) Optimality: What’s the theoretical guarantees of MIPP?
(2) Invariance: For transfer learning applications, when environments are different for each robot, how to speed up the learning via sharing information?
(3) Adaptability: When environmental parameters are changed, how to adjust the path planning?
(4) Acceleration: When designing AI chips for deep learning and submodular functions, what’s the acceleration rate of AI chips?
[Abstract]
The AI community has been paying more attention to informative path planning (IPP). IPP is to plan trajectories for robots to maximize information gathering and subject to a path cost. Potential applications include cooperative map exploration, spatial search and disinfection robots etc. However, finding optimal solutions for the aforementioned problems are to solving two NP-hard problems. To make a breakthrough of the IPP research status, this research proposes the submodular tree structure for path cost functions instead of focusing on information functions as prior work and further analyzes it via Fourier methods. The goal of this research is to explore some issues of IPP:
(1) When the submodular tree is adopted as path cost functions, what’s the boosting theoretical guarantees?
(2) What’s the sparsity of the submodular tree in the Fourier domain?
(3) For transfer learning applications, when the environments are changed, what’s the invariant property for IPP problems?
(4) When the problem is multiple robots IPP, how to distribute multiple submodular trees?
[Abstract]
The AI community has been paying more attention to multi-robot informative path planning (MIPP). MIPP is to plan trajectories for robots to maximize information gathering. If the robots are cooperative, potential applications include cooperative map exploration, cooperative search and cooperative disinfection etc. If robots are adversarial, potential applications include pursuit evasion games. However, finding optimal solutions for the aforementioned problems are NP-hard. Hence, this research proposes an inverse reinforcement learning approach to improve MIPP performance of robots through imitate how humans solve MIPP problems in daily lives (e.g., cooperative search).
To make a breakthrough of the MIPP research status, this research analyzes MIPP problems through adaptive submodularity, topology, and adversary. To consider the state uncertainty, the adaptive submodularity of MIPP will be explored. To consider the invariant property, the topology of MIPP will be explored. To consider the adversarial status, the adversary of MIPP will be explored.
The goal of this research is to explore some issues of MIPP:
(1) When the targets and environments are probabilistic, could the MIPP has theoretical guarantees?
(2) When the targets and environments are dynamic, what’s the invariant property for MIPP problems?
(3) When the targets are adversarial, could the MIPP has theoretical guarantees?
(4) When the targets execute adversarial attack, could the MIPP has theoretical guarantees?
[Abstract]
The AI community has been paying more attention to the concept of informative path planning (IPP). The difference between path planning and IPP is that IPP is to maximize information gathering instead of avoiding obstacles. There are different applications depending on the definition of information (e.g., detection of infected plants, search for structural failure, mountain rescue and search, illegal logging, monitor of pollutions and 3D mapping). However, finding optimal solutions for these problems is NP-hard, so finding approximate solutions is a feasible way. To make a breakthrough of the IPP research status, this research proposed a deep inverse reinforcement learning approach to improve IPP performance of robots through analyzing how humans solve IPP problems in daily lives. The project will take three years. The focus of the first year is to explore the reward functions of that humans solve IPP problems via deep inverse reinforcement learning. The focus of the second year is to analyze the transfer learning of that humans solve different IPP problems. The focus of the third year is to explore human-robot cooperative IPP problems. The goal of this research is to explore three issues of IPP:
(1) IPP is learnable? If it is learnable, how much data robots need?
(2) How do humans transfer their knowledge for different IPP problems?
(3) What’s the difference and respective strengths of humans and robots?
[Abstract]
Since the popularity of unmanned aerial vehicles (UAVs), robots can help humans solving practical and high-dimensional problems (e.g., detection of infected plants, search for structural failure, mountain rescue and search for illegal logging). These problems involve autonomously guidance in 3D environments and need to real-time process large data. These problems are proved as NP-hard problems. Hence, professional pilots are still necessary for UAVs. This project is to collect the UAV’s flying data via pilots’ control in complex 3D environments. And then, the UAV will learn from the data. It could make the flying ability of UAV be equal to or greater than the flying ability of pilots. The goal of this project is to explore three key issues of 3D search problems.
(1) The ability of UAV is better than human pilots?
(2) The UAV search problems are learnable? If yes, how much data the UAV need?
(3) What’s the optimality of the proposed solutions?
[Abstract]
The goal of AI maker lab is to build a teaching lab at the Mathematics department in NCU. The difference between AI maker lab and other maker labs is that students only build AI programs instead of gears or circuits. The students in AI lab only focus on math and AI programs.
Currently, the AI maker lab has 10 Minibots (Mobile robots) and 8 bebop (UAVs). If the students took AI related courses, they must build their AI programs on these robots for final projects. After taking theses courses, the students can access to AI maker lab anytime to develop their own AI programs. The AI-related courses are as follows:
Perception and estimation in robotics,
Modern artificial intelligence
Introduction to data science
[Abstract]
During the age of AI, it’s important to teach students AI. However, most of AI platforms are simulators, which cannot satisfy the requirement of AI industrial. If the general AI platform was built, students can implement AI algorithms on real robot dog platforms and sensors. These platforms will train AI engineers for AI industrial. The goal of this project is as follows:
(1) Develop 10 robot dogs with RGBD cameras, laser, and IMU.
(2) Develop supervised learning, unsupervised learning, and reinforcement learning lectures based on robot dog platform.
(3) Evaluate students’ study performance.
Key words: Robot dog, general platform, supervised learning, unsupervised learning, and reinforcement learning.