Ph.D. Candidate
School of Intelligence Science and Technology [Google Scholar] [GitHub] [Twitter] [WeChat] |
|
Hi there๐, my name is Chenguo Lin (in Chinese: ๆ็ๆ). I am a final-year Ph.D. candidate of computer science and artificial intelligence at
Peking University, China, supervised by
Prof. Yadong Mu.
My research focuses on Multi-modal World Models: intelligent systems that can ๐ perceive, ๐ง reason, and ๐ฆพ act across both the physical world and the digital world. I view world models as the substrate for recursive self-improvement (RSI): they turn interaction into verifiable experience, enabling agents to continuously expand both their understanding of the world and their ability to act within it. I am particularly interested in: (1) Physical World Models: controllable video/3D/action generation with spatial intelligence, for understanding, simulating, and interacting with the visual/physical environment. (2) Digital World Models: GUI agents and computer-using agents (CUA) that perceive screens, reason over interfaces, and interact through tools, extending world modeling to dynamic and interactive digital environments.
I've spent wonderful time as a research intern at
ByteDance Seedance for multi-modal video generation,
ByteDance PICO for 3D generation and spatial intelligence,
Microsoft Research Asia (MSRA) for self-supervised learning with Dr. Zhirong Wu and Dr. Stephen Lin, and as a research assistant at
HKU with Prof. Ping Luo, and
KAIST with Dr. Chaoning Zhang and Prof. In So Kweon.
I'm always open to research discussions and collaborations. Feel free to ๐ฒcontact me if you are interested.
*: equal contribution; โ : project lead
|
arXiv 2026
|
|
|
arXiv 2026
|
|
|
ECCV 2026
|
|
|
CVPR 2026
|
|
|
CVPR 2026
|
|
|
CVPR 2026
|
|
|
NeurIPS 2025
|
|
|
NeurIPS 2025
|
|
|
ICLR 2025
|
|
|
ICLR 2025
|
|
|
NeurIPS 2024
|
|
|
T-PAMI 2025
|
|
|
ICLR 2024
|
|
|
TMLR 2024
|
|
ยฉ Chenguo Lin