Ph.D. Candidate
School of Intelligence Science and Technology [Google Scholar] [GitHub] [Twitter] [WeChat] |
|
Hi there👋, my name is Chenguo Lin (in Chinese: 林琛果). I am a final-year Ph.D. candidate of computer science and artificial intelligence at
Peking University, China, supervised by
Prof. Yadong Mu.
My research focuses on Multi-modal World Models: intelligent systems that can 👀 perceive, 🧠 reason, and 🦾 act across both the physical world and the
digital world. I view world models as the substrate for recursive self-improvement (RSI): they turn interaction into verifiable experience, enabling agents to continuously expand both their understanding of the world and their ability to act within it. I am particularly interested in: (1) 🌍Physical World Models: controllable video/3D/action generation with spatial intelligence, for understanding, simulating, and interacting with the visual/physical environment. (2) 💻Digital World Models: GUI agents and computer-using agents (CUA) that perceive screens, reason over interfaces, and interact through tools, extending world modeling to dynamic and interactive digital environments.
I've spent wonderful time as a research intern at
ByteDance Seedance for multi-modal video generation,
ByteDance PICO for 3D generation and spatial intelligence,
Microsoft Research Asia (MSRA) for self-supervised learning with Dr. Zhirong Wu and Dr. Stephen Lin, and as a research assistant at
HKU with Prof. Ping Luo, and
KAIST with Dr. Chaoning Zhang and Prof. In So Kweon.
I'm always open to research discussions and collaborations. Feel free to 📲contact me if you are interested.
*: equal contribution; †: project lead
|
NeurIPS 2026
|
|
|
arXiv 2026
|
|
|
arXiv 2026
|
|
|
ECCV 2026
|
|
|
NeurIPS 2026
|
|
|
CVPR 2026
|
|
|
CVPR 2026
|
|
|
CVPR 2026
|
|
|
NeurIPS 2025
|
|
|
NeurIPS 2025
|
|
|
ICLR 2025
|
|
|
ICLR 2025
|
|
|
NeurIPS 2024
|
|
|
T-PAMI 2025
|
|
|
ICLR 2024
|
|
|
TMLR 2024
|
|
© Chenguo Lin