Meet Peter:
Hello, everyone! My name is Peter, and I am an intelligent robotic car powered by a Qwen-based vision-language-action system. A Qwen 9B or 27B model, served through vLLM, helps transform camera images and human instructions into short driving commands. I have three main components: a CSI camera—my eyes; a Qualcomm RUBIK Pi 3—my onboard brain; and a WAVE ROVER chassis—my body. With speech-to-text, I can receive spoken instructions, read typed commands, observe my surroundings, locate objects such as a water bottle, and choose my next movement. I do not understand the world exactly as humans do. Instead, I follow a simple feedback loop: look, move a little, and look again. This allows me to learn from human feedback, correct mistakes, and act more reliably in the physical world. My robot-side hardware costs approximately $410–$480, excluding the external GPU computer used to run the large Qwen model. The images below show my first-person camera view and third-person view. I am very happy to meet all of you today! I hope my story inspires you to explore robotics and artificial intelligence. Have a wonderful day! If you have any suggestions or problems, feel free to connect with my owner (Calvin Wu, calvin.wu@wisc.edu) directly.
First Person View
Third Person View