Ph.D. researcher in diffusion-based controllable generation and visual world models, with first- or co-first-author papers at NeurIPS, CVPR, and 3DV and two Meta Reality Labs internships. Research spans image and video synthesis, long-horizon generation, multimodal conditioning, and visual outcome prediction; current work explores bidirectional LLM–diffusion reasoning and multimodal post-training.
Research vision: connect the complementary priors and capabilities of language models and visual generative models bidirectionally, so generated visual outcomes can inform subsequent reasoning instead of remaining one-way decoded outputs.
Computer-use visual world model
Action-conditioned image diffusion · mixture of transformers
Building a SenseNova mixture-of-transformers (MoT) image-diffusion world model for action-conditioned visual forward prediction: given the current UI and a candidate action, it predicts one next UI screenshot for computer-use agent look-ahead.
The work spans diffusion post-training and Reasoner–Generator interface and training design, including structured consequence supervision, an interpretable text path, and experimental cache-preserving inference. Current research explores outcome-verified post-training and joint-training strategies that preserve Reasoner–Generator compatibility for candidate-action evaluation.
Physical-event world model
Video diffusion · multimodal reasoning
Post-training a Cosmos3 video-diffusion world model to predict physical events and outcomes from RGB history.
Current work explores a teacher-guided Reasoner–Generator loop that compares generated and target outcomes to adapt video reasoning and generation.
Robot-interaction world model
Motion-conditioned video diffusion
Motion-conditioned video-diffusion world model for predicting robot–object interactions from an initial observation and prescribed robot motion.
The synthetic prototype studies query-time proxy-to-target correspondence across robot embodiments and evaluates trajectories, contacts, and task outcomes.
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang
NeurIPS 2025 · First author
Anchored autoregressive generation for identity-consistent, minute-scale try-on videos trained from short clips; evaluated up to 90 seconds and against FramePack and Kling 2.0.
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang
arXiv 2025 · Under review at WACV 2027 · First author
A unified conditioning mechanism for garment, identity, text, layout, pose, and motion control, together with paired-triplet data generation and multi-stage training.
Publications appear under Jun-Kun Chen; earlier papers may use Junkun Chen or J. Chen. * denotes equal contribution.
Experience
Meta Reality Labs
Research Scientist Intern
May–Aug. 2023 · Zurich, Switzerland May–Dec. 2025 · Bay Area, CA
2025: Scaled a separate generative-model training program to 128–256 GPUs across multi-node clusters, supporting faster model and data iteration.
2023: Created context-rich multi-view conditioning, 3D-consistent structured noise, and self-supervised consistency training for ConsistDreamer (CVPR 2024).
SpreeAI
AI Research Intern
May 2024–May 2025 · Remote, U.S.
Dress&Dance (WACV 2027, under review): multimodal conditioning and data-efficient multi-stage training for garment, identity, pose, and motion control.
Virtual Fitting Room (NeurIPS 2025): anchored autoregressive generation for identity-consistent, minute-scale try-on video trained from short clips.
Earlier Research
Mila – Quebec AI Institute (2020): knowledge-graph reasoning for RNNLogic (ICLR 2021). Baidu NLP (2019–2020): scene-aware dialogue generation for the DSTC8 Workshop at AAAI 2020.