
The Training Data That Helps Robots Keep Their Skills
A new Stanford and Toyota Research Institute study shows how carefully selected examples can help robots learn new tasks while retaining earlier skills. For data teams, the important question is which distinctions those examples preserve.

Research paper: Memory Anchors for Continual Robot Learning
A new Stanford and Toyota Research Institute study shows how carefully selected examples can help robots learn new tasks while retaining earlier skills. For data teams, the important question is which distinctions those examples preserve.
Two jars have the same shape but different labels and lid textures. One opens clockwise; the other opens counterclockwise. A robot can learn either task on its own, yet struggle to retain both when taught them in sequence.
That is one of the challenges explored in Memory Anchors for Continual Robot Learning (link above). The researchers investigate a practical question: when a robot learns something new, which examples from earlier tasks should remain in its training mix?
Their answer centers on situations the model represents similarly but that require different actions. Earlier examples from these regions can be particularly important for preserving existing skills. The authors call them memory anchors.
Why a new task can disrupt an old one
The robots in this study learn from demonstrations that pair observations with actions. Training on a new task updates the model, but those updates can also weaken previously learned behaviors, a problem known as catastrophic forgetting.
A common safeguard is experience replay: mix a selection of earlier training examples back in while learning the new task. This selection is called a replay buffer. It gives the model continued practice on existing skills without reusing the entire previous dataset.
The paper shows why the choice of examples matters. In the LIBERO simulation benchmark, one task requires a robot to put cream cheese in a bowl. Another requires it to put the bowl on a plate. The scene is similar, and both instructions mention the bowl, but the robot needs to pick up a different object.
After learning the second task, the robot sometimes reaches for the bowl when it should reach for the cream cheese. Success on the earlier task drops by an average of 24 percentage points across the tested training orders.
The researchers connect this interference to overlap in the model’s internal representations: situations that require different behaviors have been encoded too similarly. Replaying earlier examples from those regions can help preserve the distinction.
How memory anchors are selected
Before training on a new task, the method uses the current model to identify where its existing behavior conflicts with the new demonstrations. It then retrieves relevant examples from the earlier data.
The process has three steps:
Find similar situations. Identify new-task examples that lie close to earlier examples in the model’s internal representation.
Find conflicting actions. Within that overlap, identify where the demonstrated action differs substantially from what the current model predicts.
Retrieve the earlier examples. Select old-task data closest to those conflict regions and give it a deliberate place in replay.
These are observation-action samples drawn from demonstrations, rather than a selection of complete demonstrations. In the jar experiment, the selected anchors cluster around the approach to the lid, where the robot must choose how to make contact for the correct rotation. The method concentrates replay on a distinction the new task could disrupt.
An example’s value for retention depends on what the model already knows and what it is learning next. Anchors are selected afresh for each new task, using that relationship to guide which earlier data gets priority.
The same replay budget, different results
The strongest evidence comes from changing the replay selection while holding its size constant.
First, the researchers excluded the highest-ranked 10% of anchor candidates from the earlier data before randomly selecting 1,000 replay examples. Across four LIBERO task suites, average forgetting increased more than fourfold. The buffer still contained the same number of examples; it had lost access to particularly useful ones.
They then tested the reverse approach with a smaller replay budget, equivalent to 1% of the earlier data. Reserving 10% of that buffer for anchors, with the remainder sampled randomly, reduced the average performance drop on the most conflicting task pairs by 63%.
That result concerns the six most conflicting pairs in each suite, rather than a 63% improvement across every task. Broader gains were clearest in the suite where tasks shared the same environment and objects.
The benefits also extended to physical robots. Across the final evaluations of the three jar-opening tasks, the anchor-enriched model completed 40 of 60 attempts, compared with 23 of 60 for random replay. Performance improved substantially, although errors remained. The researchers also observed benefits on a sequence of sweater-folding tasks.
A lighting difference nearly hid the problem
An observation in the appendix is especially relevant to data collection.
During an early attempt, lighting differed slightly between the jar tasks, and continual learning worked well even with random replay. Once lighting was held constant, interference between the tasks became apparent.
The authors describe this as anecdotal evidence. It suggests that incidental visual differences had made the tasks easier for the model to separate; it does not establish that lighting was the robot’s only cue.
For a data team, this raises an important testing question: does the collection setup make the task distinction artificially easy? A task-specific background or lighting condition could help a model separate examples without demonstrating that it reliably uses the intended object cues. Testing under shared conditions can help expose that gap.
What this means for human-data work
For people collecting demonstrations, curating datasets, or designing evaluations, we see three practical implications.
Collect useful contrasts. Ask which existing tasks a new task could be confused with. Include examples that share a scene or object type but require different valid behaviors. Make sure the instruction or object feature that determines the correct action is available to the model. Similarity is useful for testing only when there is still a meaningful distinction to learn.
Keep key moments connected to their context. When segmenting or selecting demonstrations, preserve the relationship between the observation, the instruction where relevant, and the action. A moment becomes valuable because of what the robot must distinguish there. Visually similar examples may carry different and important behavioral requirements.
Test earlier skills after each update. New-task success does not establish that previous capabilities survived. Examine which task combinations trigger regressions, and distinguish choosing the wrong behavior from executing the correct behavior poorly. Those failures call for different follow-up work. The paper’s task-pair analysis and real-robot failure breakdown illustrate why both distinctions matter.
Where the findings apply
The clearest case for targeted replay is a constrained replay budget and tasks that are easy to confuse. Diverse examples still matter: the authors report that filling the buffer entirely with anchors harms performance. The study also found that π0.5, a large pretrained robot model, remained sensitive to replay selection under restricted budgets and benefited from anchor enrichment.
The experiments cover sequences of three to ten tasks learned from demonstrations. Selecting anchors requires access to earlier training data and the model’s internal representations, and adds computation. Longer deployments and alternative ways of identifying useful examples remain questions for further work.
The central lesson is that a demonstration can help teach an action and preserve when that action should be used. As robots acquire more skills, those distinctions deserve deliberate attention in the data they revisit.
At Data Gradient, we translate frontier research into practical training for the people who create and evaluate AI data. Explore our courses and applied learning through Data Gradient Learn.
Research: Maximilian Du, Zhanyi Sun, Chen Xu, Paarth Shah, Masha Itkina, and Shuran Song. Memory Anchors for Continual Robot Learning. Stanford University and Toyota Research Institute. Preprint, arXiv:2608.26545, August 27, 2026.