Ph.D. researcher in diffusion-based controllable generation at UIUC, advised by Yu-Xiong Wang. I previously worked as a research intern at Meta and SpreeAI.
My research centers on learning new generative capabilities through supervision and training design. At SpreeAI, this led to new visual controls in Dress&Dance and, subsequently, identity-consistent, minute-scale generation in Virtual Fitting Room(NeurIPS 2025).
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang
NeurIPS 2025 · First author
Reference-conditioned training on ordinary short-clip pairs enables generated 360° anchors to guide identity-consistent try-on videos up to 90 seconds. Evaluated against D&D+FramePack and D&D+Kling 2.0 pipelines.
Jun-Kun Chen, Samuel Rota Bulò, Norman Müller, Lorenzo Porzi, Peter Kontschieder, Yu-Xiong Wang
CVPR 2024 · First author
Cross-view self-supervision and staged diffusion adaptation for consistent scene editing, with losses throughout the denoising trajectory. An asynchronous multi-GPU pipeline overlaps diffusion training and image generation with scene fitting.
Jun-Kun Chen, Aayush Bansal, Minh Phuoc Vo, Yu-Xiong Wang
WACV 2027, under review · First author
Target-steered supervision and staged training teach pretrained video diffusion new garment, identity, and motion controls. Generates 1152×720, 24-FPS try-on videos, with 20–59% lower FVD than a Kling 1.6 try-on-and-animation pipeline on Internet and captured benchmarks.
Publications appear under Jun-Kun Chen; earlier papers may use Junkun Chen or J. Chen. * denotes equal contribution.
Current Research
Computer-use visual world model
Image diffusion · mixture of transformers
Designed and ran diffusion post-training on SenseNova MoT for next-screen prediction from a screenshot and action. Formulated teacher-grounded supervision for action effects and content preservation; incorporated interaction traces and failure cases into Reasoner training.
Physical-event world model
Video diffusion · multimodal reasoning
Adapted Cosmos3 video diffusion for physical-event prediction from RGB history. Implemented online teacher distillation using generated-versus-target video feedback to train the language Reasoner and diffusion Generator.
Robot-interaction world model
Motion-conditioned video diffusion
Built an initial synthetic video-diffusion prototype predicting robot–object interactions from an initial scene and prescribed motion. Designed visual-proxy and paired-correspondence conditioning to represent motion across robot–world setups.
Experience
Meta Reality Labs
Research Scientist Intern
May–Dec. 2025 · Bay Area, CA, U.S. Part-time continuation: Aug.–Dec. 2025
Designed and implemented a GPU-role-separated pipeline for data processing and model training, supporting 3D/video-generation research at 128–256-GPU scale on Meta's distributed infrastructure.
May–Aug. 2023 · Zurich, Switzerland
For ConsistDreamer (CVPR 2024), carried the research from method design through implementation and evaluation: cross-view self-supervision, staged diffusion adaptation, and asynchronous scheduling of diffusion training, image generation, and scene fitting.
SpreeAI
AI Research Intern
May 2024–May 2025 · Remote, U.S.
Dress&Dance: designed and implemented target-steered supervision, synthetic-condition/real-target training triplets, and staged image/video curricula to teach pretrained diffusion new visual controls. Evaluated garment and identity fidelity, motion, and video quality.
Subsequent work — Virtual Fitting Room (NeurIPS 2025): designed training around arbitrary pairs of short clips to learn reference appearance and prefix continuation, without requiring canonical 360° training anchors. At inference, a generated 360° reference and anchored autoregression support minute-scale consistency; targeted augmentation trains a refiner to remove characteristic generation artifacts.
Earlier Research
Mila – Quebec AI Institute (2020): knowledge-graph reasoning for RNNLogic (ICLR 2021). Baidu NLP (2019–2020): scene-aware dialogue generation for the DSTC8 Workshop at AAAI 2020.