Anonymous Authors
Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills. However, retaining task performance does not ensure the behavior remains grounded in language because policies may rely on scene cues, object associations, or memorized task structure. We introduce a benchmark protocol to study how language-guided behavior changes as robotic policies learn successive tasks. We construct meaning-preserving and meaning-changing instruction variants for the Goal, Spatial, Object, and Long suites of LIBERO. Policy experiments focus on LIBERO-Goal, evaluating Original and Paraphrase instructions after each continual-learning stage. We compare representative continual imitation learning methods under their original assumptions while separating task competence from language sensitivity. The proposed diagnostics complement standard learning and forgetting metrics by measuring semantic robustness, goal adaptation, and language sensitivity. Results show that strong continual-learning performance does not always translate to reliable language grounding, and our diagnostics help determine whether retained skills remain correctly guided by their instructions.
We build our evaluation protocol on LIBERO (https://libero-project.github.io/intro.html), a dataset for continual robot learning containing four suites with 10 tasks each: LIBERO-Goal, LIBERO-Spatial, LIBERO-Object, and LIBERO-Long. We construct language perturbations across all four suites to evaluate benchmark coverage. Our main continual imitation learning experiments use LIBERO-Goal, whose tasks vary in actions, objects, targets, and spatial relations, making it suitable for controlled language-grounding evaluation.
We show an illustrative set of experiments from our benchmark using the RaccoonBot, comparing the original instruction, a minimal contrast, and a collision. The underlined word marks the semantic change: in the minimal contrast, the relation changes from on to near, while in the collision, the attribute changes from red to green, referring to an object that is not present in the scene.
Original instruction – Place the red mug on
the counter
Minimal contrast – Place the red mug near the counter
Semantic slot changed: Relation
On -> Near
Collision– Place the green mug near the counter
Semantic slot changed: Attribute
Red -> Green (Entity not in scene)
Result 1: Every method degrades under paraphrases
Rewording the instruction while keeping the meaning costs every method roughly five to eight AUC points. Final checkpoint vs. stage-wise measurement tells very different stories. ER's final gap is 19.4 pts vs. a 6.7 stage-wise AUC gap.
Result 2: Goal switching is weak, even for the best performing method
Strong continual imitation learners can still fall back to the original task even when a minimal contrast specifies a valid, executable new goal.
The behavior rarely switches even with impossible requests on collision tasks.
Continual imitation learning methods that forget less are not automatically more grounded in language.
Paraphrase, minimal-contrast, and collision instructions expose complementary failures that task success alone cannot see.
The protocol is dataset-agnostic and post-hoc. It can be applied to any continual task checkpoint without changing the training loop.