The RoboChess Challenge:
From GPU-Accelerated Simulation to Scalable and Generalizable Real Robot Policy Learning
(CORL 2026 Workshop Proposal)
(CORL 2026 Workshop Proposal)
Codebase: https://github.com/yizhouzhao-nvidia/RoboChess-Challenge
Hosted on the CoRL 2026 Workshop on GPU-Accelerated Simulation to Scalable and Generalizable Robot Policy Learning, the RoboChess Challenge is a robotic manipulation challenge that benchmarks generalizable, sim-to-real transferable robot policies through the task of physical chess play. Rather than scripted pick-and-place, RoboChess requires policies to perceive a chess board, reason about legal moves, and execute precise single-arm manipulation of physical chess pieces — across both GPU-accelerated simulators and real robot embodiments.
The challenge directly targets the three open problems motivating the workshop: the lack of a standardized, quantitative evaluation framework for robot learning; the gap between visually convincing demonstrations and truly robust, repeatable task completion; and policy robustness to real-world variation in lighting, viewpoint, and object placement. For more details on the workshop, see the proposal and project page.
We host the challenge leaderboard on EvalAI. To participate, follow these steps:
Refer to the Github codebase and Challenge Setup Guide (links to be added) to set up your workspace, download the simulation and teleoperation datasets, and reproduce the baseline results.
Log into EvalAI (register if necessary) and go to the RoboChess challenge page. Read the task details and select a track (Simulation or Real-World) and a phase.
Train and evaluate your own policy in your chosen simulation engine(s) and/or on your real robot embodiment, using the provided assets and datasets.
Make your submission. You will need to create a team in order to submit — register a new team or join an existing one using an invite code. Please make sure your submission is set to public status to be included on the leaderboard and considered for the challenge prizes.
Built on three GPU-accelerated physics engines — Isaac Lab, MuJoCo,, and Genesis World — enabling diverse, scalable data generation and policy evaluation in simulation.
Real-world evaluation across multiple single-arm embodiments, including LeRobot SO-101, Franka, UR, Piper, and Yam/Rebot, allowing systematic measurement of sim-to-real transfer and cross-embodiment robustness.
A complete digital asset suite for reproducible setup, including robot models and chess pieces/board provided in both .stl and .usd formats.
Large-scale, paired sim-and-real datasets: simulated rollouts from IsaacLab, Genesis, and MuJoCo, alongside real teleoperation data (SO-101 teleop, Yam/Franka/UR teleop).
Baseline benchmarking against state-of-the-art VLA foundation models, including PI, GR00T, Cosmos Action, and others, so new methods can be compared against strong existing policies out of the box.
A standardized, quantitative evaluation protocol covering in-domain performance as well as generalization to novel lighting, viewpoints, and piece/board configurations.
Participants are also encouraged to submit a short paper describing their approach, results, and any failure-case analysis, in line with the workshop's contributed papers track.
EvalAI. We host the evaluation servers on EvalAI, an open-source platform for organizing and participating in AI challenges.
Tracks. RoboChess runs two complementary tracks: a Simulation Track, evaluated across MuJoCo, Isaac Lab, and Genesis, and a Real-World Track, evaluated across the LeRobot SO-101, Franka, UR, Piper, and Yam/Rebot platforms.
Data splits. Evaluation data is divided into dev, test-challenge, and test-reserve splits. Dev is used for debugging and sanity checks, with a daily submission cap, and does not include the full generalization suite (novel lighting, viewpoints, or board/piece configurations). Test-challenge is the standard evaluation split for the challenge, with a public leaderboard updated on submission. Test-reserve guards against overfitting: a large gap between a method's test-challenge and test-reserve scores will be flagged for further investigation, and test-reserve scores are not publicly revealed.
Phases. We maintain two phases per track: Dev and Test. The Dev phase uses the dev split for development and sanity checking; the Test phase uses the test-challenge and test-reserve splits. We encourage participants to submit to Dev first. Submission procedures are identical across phases. Only public submissions to the Test phase count toward official challenge participation and the leaderboard; private submissions and Dev-phase submissions do not.
We host the challenge leaderboard on EvalAI. To participate, follow these steps:
Refer to the Github codebase and Challenge Setup Guide (links to be added) to set up your workspace, download the simulation and teleoperation datasets, and reproduce the baseline results.
Log into EvalAI (register if necessary) and go to the RoboChess challenge page. Read the task details and select a track (Simulation or Real-World) and a phase.
Train and evaluate your own policy in your chosen simulation engine(s) and/or on your real robot embodiment, using the provided assets and datasets.
Make your submission. You will need to create a team in order to submit — register a new team or join an existing one using an invite code. Please make sure your submission is set to public status to be included on the leaderboard and considered for the challenge prizes.
Dates below follow the confirmed CoRL 2026 workshop schedule; challenge submission dates are placeholders to be finalized.
TBD (Summer 2026): Challenge launch — codebase, datasets, and digital assets released; submissions open
TBD (Fall 2026): Submission deadline for leaderboard entries and contributed short papers
November 9, 2026: Winners announced at the CoRL 2026 Workshop on GPU-Accelerated Simulation to Scalable and Generalizable Robot Policy Learning, JW Marriott Austin, Austin, TX (the workshop day ahead of the main CoRL 2026 conference, November 10–12, 2026)
The RoboChess Challenge is organized by a team spanning industry and academia, bringing together expertise in GPU-accelerated simulation, large-scale compute infrastructure, and robot learning research. The team includes researchers from NVIDIA, whose work on simulation platforms and accelerated computing underpins much of the challenge's simulation track, which contributes cloud compute infrastructure and supports the sponsored credits offered to participants. Academic members from the University of Wisconsin–Madison round out the team with a research focus on robot policy learning and generalization.
Together, the organizers were motivated by a shared interest in closing the gap between simulation-based training and reliable real-world robot deployment — the central challenge this workshop and its associated chess-playing benchmark are designed to address. The group combines hands-on experience building and scaling simulation tools with a research agenda centered on evaluation, robustness, and sim-to-real transfer, and is committed to making the challenge accessible to participants across career stages and institutions through shared datasets, digital assets, and compute support.
(CORL Workshop Proposal Challenge Site)
Updated 06/17/26