Ph.D. researcher in diffusion-based controllable generation and visual world models, with first- or co-first-author papers at NeurIPS, CVPR, and 3DV and two Meta Reality Labs internships. Research spans image and video synthesis, long-horizon generation, multimodal conditioning, and visual outcome prediction; current work explores bidirectional LLM–diffusion reasoning and multimodal post-training.
Research direction: train language and visual generation components around predicted outcomes, connecting reasoning and generation across computer-use, physical-event, and robot-interaction settings.
Computer-use visual world model
Image diffusion · mixture of transformers
Built an action-conditioned next-screen prediction pipeline on the existing SenseNova MoT backbone using structured consequence supervision and separate Reasoner and Generator post-training.
Preliminary evaluation across 2,850 cases found teacher-refined conditions improved 400 next-screen predictions versus 130 regressions, with 2,320 unchanged.
Physical-event world model
Video diffusion · multimodal reasoning
Built and post-trained a Cosmos3 video-diffusion system that predicts 69 future frames from a 49-frame RGB history; completed reproducible fixed inference and evaluation suites.
Implemented online multimodal-teacher post-training with sequential Reasoner and diffusion-Generator updates; physical-consistency gains remain under evaluation.
Robot-interaction world model
Motion-conditioned video diffusion
Built an initial synthetic video-diffusion prototype that predicts robot–object interaction video from an initial scene and prescribed robot motion, providing a controlled testbed for visual forward modeling.
The prototype combines robot-only visual proxies with target-system correspondence examples for controlled motion transfer and interaction prediction in the synthetic setting.
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang
NeurIPS 2025 · First author
Anchored autoregressive generation for identity-consistent, minute-scale try-on videos trained from short clips; evaluated up to 90 seconds and against FramePack and Kling 2.0.
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang
arXiv 2025 · Under review at WACV 2027 · First author
A unified conditioning mechanism for garment, identity, text, layout, pose, and motion control, together with paired-triplet data generation and multi-stage training.
Publications appear under Jun-Kun Chen; earlier papers may use Junkun Chen or J. Chen. * denotes equal contribution.
Experience
Meta Reality Labs
Research Scientist Intern
May–Aug. 2023 · Zurich, Switzerland May–Dec. 2025 · Bay Area, CA
2025: Scaled a separate generative-model training program to 128–256 GPUs across multi-node clusters, supporting faster model and data iteration.
2023: Created context-rich multi-view conditioning, 3D-consistent structured noise, and self-supervised consistency training for ConsistDreamer (CVPR 2024).
SpreeAI
AI Research Intern
May 2024–May 2025 · Remote, U.S.
Dress&Dance (WACV 2027, under review): multimodal conditioning, paired-triplet data construction, and multi-stage training for garment, identity, pose, and motion control.
Virtual Fitting Room (NeurIPS 2025): anchored autoregressive generation for identity-consistent, minute-scale try-on video trained from short clips.
Earlier Research
Mila – Quebec AI Institute (2020): knowledge-graph reasoning for RNNLogic (ICLR 2021). Baidu NLP (2019–2020): scene-aware dialogue generation for the DSTC8 Workshop at AAAI 2020.