About

Ph.D. researcher in diffusion-based controllable generation and visual world models, with first- or co-first-author papers at NeurIPS, CVPR, and 3DV and two Meta Reality Labs internships. Research spans image and video synthesis, long-horizon generation, multimodal conditioning, and visual outcome prediction; current work explores bidirectional LLM–diffusion reasoning and multimodal post-training.

Advised by Prof. Yu-Xiong Wang at the University of Illinois Urbana-Champaign.

Current Research

Research direction: train language and visual generation components around predicted outcomes, connecting reasoning and generation across computer-use, physical-event, and robot-interaction settings.

Computer-use visual world model

Image diffusion · mixture of transformers

Built an action-conditioned next-screen prediction pipeline on the existing SenseNova MoT backbone using structured consequence supervision and separate Reasoner and Generator post-training.

Preliminary evaluation across 2,850 cases found teacher-refined conditions improved 400 next-screen predictions versus 130 regressions, with 2,320 unchanged.

Physical-event world model

Video diffusion · multimodal reasoning

Built and post-trained a Cosmos3 video-diffusion system that predicts 69 future frames from a 49-frame RGB history; completed reproducible fixed inference and evaluation suites.

Implemented online multimodal-teacher post-training with sequential Reasoner and diffusion-Generator updates; physical-consistency gains remain under evaluation.

Robot-interaction world model

Motion-conditioned video diffusion

Built an initial synthetic video-diffusion prototype that predicts robot–object interaction video from an initial scene and prescribed robot motion, providing a controlled testbed for visual forward modeling.

The prototype combines robot-only visual proxies with target-system correspondence examples for controlled motion transfer and interaction prediction in the synthetic setting.

Selected Publications

Google Scholar

Dress&Dance: Dress up and Dance as You Like It

Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang

arXiv 2025 · Under review at WACV 2027 · First author

A unified conditioning mechanism for garment, identity, text, layout, pose, and motion control, together with paired-triplet data generation and multi-stage training.

Additional publications

Publications appear under Jun-Kun Chen; earlier papers may use Junkun Chen or J. Chen. * denotes equal contribution.

Experience

Meta Reality Labs

Research Scientist Intern

May–Aug. 2023 · Zurich, Switzerland
May–Dec. 2025 · Bay Area, CA

2025: Scaled a separate generative-model training program to 128–256 GPUs across multi-node clusters, supporting faster model and data iteration.

2023: Created context-rich multi-view conditioning, 3D-consistent structured noise, and self-supervised consistency training for ConsistDreamer (CVPR 2024).

SpreeAI

AI Research Intern

May 2024–May 2025 · Remote, U.S.

Dress&Dance (WACV 2027, under review): multimodal conditioning, paired-triplet data construction, and multi-stage training for garment, identity, pose, and motion control.

Virtual Fitting Room (NeurIPS 2025): anchored autoregressive generation for identity-consistent, minute-scale try-on video trained from short clips.

Earlier Research

Mila – Quebec AI Institute (2020): knowledge-graph reasoning for RNNLogic (ICLR 2021). Baidu NLP (2019–2020): scene-aware dialogue generation for the DSTC8 Workshop at AAAI 2020.

Education

University of Illinois Urbana-Champaign

Ph.D. Candidate in Computer Science · Expected Dec. 2026
Advisor: Yu-Xiong Wang

Tsinghua University

B.Eng. in Computer Science, Yao Class · 2017–2021
Major GPA: 3.94/4.0

Selected Honors

  • ICPC EC-Final — Gold Medal, 4th Place (2018)
  • ICPC Qingdao Regional — Gold Medal, 2nd Place (2018)
  • CCPC Harbin Regional — Gold Medal (2017)
  • National Olympiad in Informatics — Gold Medal (2016)