RoboAug:One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation

CoRL 2026Accepted paper

Xinhua Wang1,* Kun Wu1,* Zhen Zhao1,* Hu Cao2 Yinuo Zhao1,3 Zhiyuan Xu1 Meng Li1 Shichao Fan1,4 Di Wu1,5 Yixue Zhang1,6 Ning Liu1 Zhengping Che1,†,✉ Jian Tang1,✉

Beijing Innovation Center of Humanoid Robotics

Affiliations & author notes
  1. 1Beijing Innovation Center of Humanoid Robotics
  2. 2School of Automation, Southeast University
  3. 3City University of Hong Kong
  4. 4The School of Mechanical Engineering and Automation, Beihang University
  5. 5State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University
  6. 6The School of Advanced Manufacturing and Robotics, Peking University

* Equal contribution · † Team leader · ✉ Corresponding authors

Cite this work
RoboAug augments task-relevant demonstrations for robotic generalization
RoboAug enables robust robotic generalization in diverse, unseen scenes.
1reference annotation
35k+real-world rollouts
3robot platforms
170unseen backgrounds

Abstract

Learning beyond
the training scene.

Read on arXiv

Abstract

Enhancing the generalization capability of robotic learning to enable robots to operate effectively in diverse, unseen scenes is a fundamental and challenging problem. Existing approaches often depend on pretraining with large-scale data collection, which is labor-intensive and time-consuming, or on semantic data augmentation techniques that necessitate an impractical assumption of flawless upstream object detection in real-world scenarios. In this work, we propose RoboAug, a novel generative data augmentation framework that significantly minimizes the reliance on large-scale pretraining and the perfect visual recognition assumption by requiring only the bounding box annotation of a single image during training. Leveraging this minimal information, RoboAug employs pre-trained generative models for precise semantic data augmentation and integrates a plug-and-play region-contrastive loss to help models focus on task-relevant regions, thereby improving generalization and boosting task success rates. We conduct extensive real-world experiments on three robots, namely UR-5e, AgileX, and Tien Kung 2.0, spanning over 35k rollouts.

Empirical results demonstrate that RoboAug significantly outperforms state-of-the-art data augmentation baselines. Specifically, when evaluating generalization capabilities in unseen scenes featuring diverse combinations of backgrounds, distractors, and lighting conditions, our method achieves substantial gains over the baseline without augmentation. The success rates increase from 0.09 to 0.47 on UR-5e, from 0.16 to 0.60 on AgileX, and from 0.19 to 0.67 on Tien Kung 2.0. These results highlight the superior generalization and effectiveness of RoboAug in real-world manipulation tasks.

01 The method

From one annotation
to robust policies.

Overview of RoboAug. RoboAug contains three stages: (1) task-relevant region extraction, (2) semantic data augmentation, and (3) region-contrastive policy learning.

The three-stage RoboAug framework: region extraction, semantic augmentation, and region-contrastive learning
01

Task-Relevant Region Extraction

In Step 1, key regions are identified in all videos through a training-free, one-shot matching mechanism that requires only a single manually labeled reference image. Next, these initial annotations are propagated across video frames using integrated segmentation and tracking to generate consistent, dense masks for the entire dataset.

02

Semantic Data Augmentation

In Step 2, novel background scenes are synthesized in their entirety using a pre-trained generative model, thereby circumventing the visual artifacts typically associated with inpainting techniques. The task-relevant foreground is then composited onto these generated backgrounds, preserving critical structural integrity and massively expanding the training dataset.

03

Region-Contrastive Policy Learning

In Stage 3, we train the robotic manipulation policy using a plug-and-play region-contrastive loss. This loss function helps the model focus on task-relevant regions while being invariant to irrelevant background variations, thereby improving generalization to unseen scenes and environments.

02 Open data

RoboAug-D Dataset

We introduce RoboAug-D, a large-scale object detection dataset manually annotated from the perspective of robotic manipulators. The dataset encompasses 33 distinct manipulation tasks, comprising a total of 73,749 keyframes and 366,835 bounding boxes across 46 object categories.

Get the dataset
01

Dataset examples

Annotated robotic manipulation keyframes from the RoboAug-D dataset
Manipulation tasks
33
Keyframes
73,749
Bounding boxes
366,835
Object categories
46
02

Data distribution

RoboAug-D object category distribution

Zero-shot detection evaluation

We evaluate zero-shot object detection on the full RoboAug-D test set without fine-tuning, comparing RoboAug with GroundingDINO, LLMDet, and SAM3 using mAP@0.5. The figure reports the dataset-wide average and results for five representative object categories.

CoRL 2026 Figure 3: mAP@0.5 comparison of GroundingDINO, LLMDet, SAM3, and RoboAug
Figure 3 from the CoRL 2026 paper. Zero-shot detection results on RoboAug-D, including SAM3. View full resolution

03 Real-world validation

Three robots.
Over 35,000 trials.

We evaluated RoboAug across three robot embodiments (Tien Kung 2.0, Single-Arm UR-5e, and AgileX Cobot Magic 2.0), spanning tasks from single-arm pick-and-place to precise bimanual manipulation. Through over 35,000 real-world trials, we rigorously assess its generalization to unseen scenes with diverse backgrounds, distractors, and lighting conditions.

Single-Arm UR-5eAgileX Cobot Magic 2.0Tien Kung 2.0
Robot embodiments, manipulation tasks, and experimental equipment

04 Generalization in action

RoboAug real-world experiments

A

Compositional generalization

We evaluated RoboAug under triple-factor variations: 3 unseen backgrounds, 4 lighting conditions, and 3 distractors.

Compositional generalization results under background, lighting, and distractor changes

Tien Kung 2.0

TK2-WeightApple
TK2-LayPlateBowl
TK2-HeatBread

AgileX Cobot Magic 2.0

AGX-UprightMug
AGX-PutCornPlate
AGX-CloseDrawerCorn

Single-Arm UR-5e

UR-MoveLemon
UR-OpenDrawerCorn
UR-StoreCarrot
B

Dual-factor generalization

Additionally, we evaluated RoboAug under dual-factor variations: 5 unseen backgrounds and 10 task-irrelevant distractors.

Dual-factor generalization results across backgrounds and distractors
UR-MoveLemon
AGX-StackBowl
UR-PutCornPot
C

Single-factor generalization

We also evaluated RoboAug under single-factor variations: up to 170 unseen backgrounds, 20 lighting conditions, and up to 10 task-irrelevant distractors.

UR-PutCornPot with 170 Backgrounds
AGX-CloseDrawerCorn with 20 Lighting
UR-StackBowl with 1,3,5,10 Distractors

05 Side-by-side evaluation

Comparison with baselines

RoboAug was compared against baseline methods including No Aug, GenAug, Roboengine-T, and Roboengine-G. We selected the best baseline (GenAug) for comparative demonstration videos. Demonstration videos are shown below:

AGX-StackBowl
TK2-CollectBall
UR-StackBowl

06 Publication

Cite RoboAug

If you find RoboAug or RoboAug-D useful in your research, please cite our paper.

arXiv · BibTeX
Download .bib
@misc{wang2026roboaugannotationhundredsscenes,
      title={RoboAug: One Annotation to Hundreds of Scenes via Region-Contrastive Data Augmentation for Robotic Manipulation},
      author={Xinhua Wang and Kun Wu and Zhen Zhao and Hu Cao and Yinuo Zhao and Zhiyuan Xu and Meng Li and Shichao Fan and Di Wu and Yixue Zhang and Ning Liu and Zhengping Che and Jian Tang},
      year={2026},
      eprint={2602.14032},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2602.14032},
}