Ph.D. researcher in diffusion-based controllable generation at UIUC, advised by Yu-Xiong Wang, with research internships at Meta and SpreeAI.
At SpreeAI, my work progressed from high-fidelity, controllable video try-on with Dress&Dance to identity-consistent, minute-scale generation with Virtual Fitting Room(NeurIPS 2025).
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang
NeurIPS 2025 · First author
Anchored autoregressive generation for identity-consistent, minute-scale try-on videos trained from short clips; evaluated up to 90 seconds and compared with FramePack and Kling 2.0.
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang
WACV 2027, under review · First author
Multimodal conditioning and multi-stage training for garment, identity, pose, and motion control. Reduced FVD by 20–59% versus a Kling 1.6 try-on-and-animation pipeline on two evaluation sets.
Publications appear under Jun-Kun Chen; earlier papers may use Junkun Chen or J. Chen. * denotes equal contribution.
Current Research
Computer-use visual world model
Image diffusion · mixture of transformers
Developed task-specific supervision and diffusion post-training for next-screen prediction on SenseNova MoT. Combined teacher-derived consequence supervision and agent-interaction data with separate adaptation of reasoning and generation.
Physical-event world model
Video diffusion · multimodal reasoning
Adapted Cosmos3 video diffusion for physical-event prediction from RGB history. Implemented online teacher distillation using generated-versus-target video feedback to train the language Reasoner and diffusion Generator.
Robot-interaction world model
Motion-conditioned video diffusion
Built an initial synthetic video-diffusion prototype predicting robot–object interactions from an initial scene and prescribed motion. Designed visual-proxy and paired-correspondence conditioning to represent motion across robot–world setups.
Experience
Meta Reality Labs
Research Scientist Intern
May–Dec. 2025 · Bay Area, CA, U.S. Part-time continuation: Aug.–Dec. 2025
Scaled a 3D/video-generation training project to 128–256 GPUs across multi-node clusters, including parallel training workflows.
May–Aug. 2023 · Zurich, Switzerland
Designed context-rich multi-view conditioning, 3D-consistent structured noise, and self-supervised consistency training for ConsistDreamer (CVPR 2024).
SpreeAI
AI Research Intern
May 2024–May 2025 · Remote, U.S.
Dress&Dance: multimodal conditioning, paired-triplet data construction, and multi-stage training for high-fidelity video try-on with garment, identity, pose, and motion control.
Subsequent work — Virtual Fitting Room (NeurIPS 2025): anchored autoregressive generation extended this research to identity-consistent, minute-scale video, trained from short clips.
Earlier Research
Mila – Quebec AI Institute (2020): knowledge-graph reasoning for RNNLogic (ICLR 2021). Baidu NLP (2019–2020): scene-aware dialogue generation for the DSTC8 Workshop at AAAI 2020.