
Yunhao Fang 方云浩
Research Scientist at Physical Intelligence
I build multimodal agents that perceive, memorize, interact, and learn from experience. My research spans multimodal understanding and generation, long-context processing, and real-time virtual and physical interaction.
Previously, I worked at ByteDance Seed and NVIDIA, and studied at UC San Diego and Zhejiang University.
Research arc01 — 04
The long view
From multimodal perception to agents that memorize, interact, and learn.
- 01Perceive
Unify multimodal understanding and generation.
- 02Memorize
Process long contexts without losing what matters.
- 03Interact
Operate in virtual and physical worlds in real time.
- 04Learn
Distill agent experience into model weights.
Publication indexSelected
01
Multimodal perception
Understanding & generation
- 2025Seed1.5-VL Technical ReportTechnical report
- 2025WorldModelBench: Judging Video Generation Models As World Models*NeurIPS Datasets & Benchmarks
- 2025NVILA: Efficient Frontier Visual Language ModelsCVPR
- 2025VILA-U: A Unified Foundation Model Integrating Visual Understanding and GenerationICLR
- 2024VILA²: VILA Augmented VILA*ICCVW Oral
- 2023Distilling Large Vision-Language Model with Out-of-Distribution Generalizability*ICCV
02
Memory
Long-context processing
03
Real-time interaction
Virtual- and physical-world agents
- 2026π₀.₇: A Steerable Generalist Robotic Foundation Model with Emergent CapabilitiesTechnical report
- 2026EgoVLA: Learning Vision-Language-Action Models from Egocentric Human VideosIROS
- 2025π*₀.₆: A VLA That Learns From ExperienceTechnical report
- 2025Lumine: An Open Recipe for Building Generalist Agents in 3D Open Worlds*Technical report
04
Learn from experience
Experience distillation
Have a hard problem at the edge of models and the physical world?
Let’s chat.

