I am currently a Ph.D. candidate in Tianjin Key Laboratory of Visual Computing and Intelligent Perception (VCIP) and Media Computing Lab (MCLab) at the College of Computer Science, Nankai University, supervised by Prof. Ming-Ming Cheng and Prof. Qibin Hou. Prior to this, I completed seven years of undergraduate and master's studies at Dalian University of Technology (DUT).
I am currently a Qingyun Program Intern at Tencent Hunyuan (March 2026–present), working on efficient embodied foundation models and physical-world agents. Previously, I was a Research Intern at ByteDance's Volcano Engine Multimedia Laboratory (November 2024–March 2026), focusing on reinforcement post-training for video MLLMs and their deployment in VOD and live-streaming applications.
My current research interests focus on multimodal large language models, reinforcement-learning post-training, adaptive agents, and open-world perception.
I am dedicated to contributing to open-source projects, and my work can be found in HVision-NKU. Additionally, I maintain a list of Awesome Open-Vocabulary Semantic Segmentation resources.
If you're interested in my research or have any research-related questions, please feel free to contact me via email at yunhengli [at] mail.nankai.edu.cn or yunheng.li.21 [at] gmail.com.