Yucheng Hu

Hi! I am a 2th-year PhD student in Computer Science at Tsinghua University, advised by Prof. Jianyu Chen. I received my bachelor's degree in EE also from Tsinghua University in 2024 . I have previously interned at Seed-Robotics, Robotera, and Shanghai AI Laboratory.

My research focuses on embodied foundation models and generative models , aiming to maximize data utilization toward general-purpose embodied intelligence. To this end, My work includes World Action Model[7,9], UAM (Unified Action Model)[10,14,15] and VLA[3,6,8,13] Model.

Besides research, I enjoy playing video games and traveling. I'm currently preparing to launch a startup. Anyone interested in embodied intelligence is welcome to get in touch with us.

My email: h790762279@gmail.com

Phone and WeChat: 17347316224

Email  /  Scholar  /  Github

profile photo

Selected Research (* indicates equal contribution)

World Action Model:

[9] Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Yucheng Hu*, Yanjiang Guo*, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, Jianyu Chen
International Conference on Machine Learning (ICML), 2025   (Spotlight, 2.6%)
project page / code / arXiv / twitter / 机器之心 / 量子位

We finetune a general-purpose video diffusion model into manipulation-focused video prediction model to guide policy learning. VPP was the earliest WAM work.
[7] Prediction with Action: Visual Policy Learning via Joint Denoising Process
Yanjiang Guo*, Yucheng Hu*, Jianke Zhang, Yen-Jen Wang, Xiaoyu Chen, Chaochao Lu, Jianyu Chen
Advances in Neural Information Processing Systems (NeurIPS), 2024
project page / code / arXiv

We jointly predict future images and robot actions in a unified DiT network, transfering physical knowledge from internet video data to robots.

Unified Action Models:

[14] BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
Yucheng Hu*, Jianke Zhang*, Yuanfei Luo*, Yanjiang Guo, Xiaoyu Chen, Xinshu Sun, Kun Feng, Qingzhou Lu, Sheng Chen, Yangang Zhang, Wei Li, Jianyu Chen
Robotics: Science and Systems (RSS), 2026
project page / arXiv

We propose a unified VLA framework that integrates understanding, prediction, and execution.

[15] UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning
Jianke Zhang*, Yucheng Hu*, Yanjiang Guo, Xiaoyu Chen, Yichen Liu, Wenna Chen, Chaochao Lu, Jianyu Chen
International Conference on Machine Learning (ICML), 2026
arXiv

We introduced JEPA into VLA.

[10] UP-VLA: A Unified Understanding and Prediction Model for Embodied Agent
Jianke Zhang*, Yanjiang Guo*, Yucheng Hu*, Xiaoyu Chen, Jianyu Chen
International Conference on Machine Learning (ICML), 2025
arXiv / code

We incoperate both multi-modal understanding (MMU) and future prediction into VLA model, enhancing both high-level semantic knowledge and low-level visual dynamics.

Other Paper:

[13] VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
Jianke Zhang, Xiaoyu Chen, Yanjiang Guo, Yucheng Hu, Jianyu Chen
International Conference on Learning Representations (ICLR), 2026
project page / arXiv / code

We systematically evaluate how the base VLM affect the performance of VLA.

[8] Improving Vision-Language-Action Model with Online Reinforcement Learning
Yanjiang Guo*, Jianke Zhang*, Xiaoyu Chen*, Xiang Ji, Yen-Jen Wang, Yucheng Hu, Jianyu Chen
International Conference on Robotics and Automation (ICRA), 2025
arXiv / twitter1 / twitter2

We make some initial exploration on leveraging online RL to improve the VLA model! We notice that online RL for VLA can be extremely unstable and thus we adopted a iterative approach.

[6] HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
Jianke Zhang*, Yanjiang Guo*, Xiaoyu Chen, Yen-Jen Wang, Yucheng Hu, Chengming Shi, Jianyu Chen
Conference on Robot Learning (CoRL), 2024
arXiv / twitter / 机器之心

We finetune pretrained VLM into VLA models with hierarchical transformers, keeping the generalization ability but also much higher control frequency.

[3] Robix: A Unified Model for Robot Interaction, Reasoning and Planning
Huang Fang, Mengxi Zhang, Heng Dong, Wei Li, Zixuan Wang, Qifeng Zhang, Xueyun Tian, Yucheng Hu, Hang Li
arXiv, 2025
arXiv / project page

We introduce Robix, a unified model that integrates robot reasoning, task planning, and natural language interaction within a single vision-language architecture.


Source code from Jon Barron.