Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation

Prachi Garg, Steve Xing*, Prahit Yaugand*, Saurabh Gupta, Derek Hoiem
University of Illinois Urbana-Champaign
*Equal contribution

Abstract

State-of-the-art VLAs such as π0.5 exhibit strong semantic understanding, instruction following and task behavior. However, when deployed on new robots, even minor mismatches in hardware configuration relative to pretraining causes severe performance drops. Finetuning the VLA on in-domain expert data improves performance on the target task but leads to a loss in the original instruction following capabilities of the VLA. In this paper, we propose a self-supervised method that generates online interaction rollouts from the zero-shot VLA as additional training data for VLA finetuning. Our experiments show this finetuning scheme yields strong multi-task policies that (1) learn new skills from expert data, (2) inherit prior tasks distilled from the zero-shot model, (3) generalize to out-of-distribution objects, and (4) accomplish new tasks specified through text. We demonstrate the success of our approach on (1) a real ALOHA robot and (2) a new simulation benchmark in RoboTwin.

Results

Citation

@article{garg2026finetuning,
  title   = {Fine-Tuning VLAs with Self-Demonstrated Generative Control for Multi-Task Manipulation},
  author  = {Garg, Prachi and Xing, Steve and Yaugand, Prahit and Gupta, Saurabh and Hoiem, Derek},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}