Hi! I am a 2th-year PhD student in Computer Science at Tsinghua University, advised by Prof. Jianyu Chen.
I received my bachelor's degree in EE also from Tsinghua University in 2024 .
I have previously interned at Seed-Robotics, Robotera, and Shanghai AI Laboratory.
My research focuses on embodied foundation models and generative models , aiming to maximize data utilization toward general-purpose embodied intelligence. To this end, My work includes World Action Model[7,9], UAM (Unified Action Model)[10,14,15] and VLA[3,6,8,13] Model.
Besides research, I enjoy playing video games and traveling. I'm currently preparing to launch a startup. Anyone interested in embodied intelligence is welcome to get in touch with us.
We incoperate both multi-modal understanding (MMU) and future prediction into VLA model, enhancing both high-level semantic knowledge and low-level visual dynamics.
We make some initial exploration on leveraging online RL to improve the VLA model! We notice that online RL for VLA can be extremely unstable and thus we adopted a iterative approach.
We introduce Robix, a unified model that integrates robot reasoning, task planning, and natural language interaction within a single vision-language architecture.