I am a PhD student at Korea Advanced Institute of Science and Technology (KAIST), advised by Prof. Jinwoo Shin. Prior to this, I received B.S. in Mathematical Science and Computer Science at KAIST in 2023. My research interests lie in representation learning and generative models, with a focus on robot foundation models that generalize across diverse and complex tasks. Recently, I have been particularly focused on developing generalist robot policies. As part of this, I had the opportunity to co-lead the development of RLDX-1, a robot foundation model trained at scale. My current work focuses on representation learning for video model-based robot policies (i.e., World Action Models), including improving the representations learned by video diffusion models and leveraging human video data to learn representations that transfer from human behavior to robotic manipulation.
[T1] RLDX-1 Technical Report
Technical Report, 2026
[C11] Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models
CoRL 2026
[C10] Dual-stream diffusion for world-model augmented vision-language-action model
ICML 2026
[P2] ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
Arxiv preprint, 2025
[C8] SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
ECCV 2026
NeurIPSW-SpaVLE 2025
[C7] Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
CVPR 2025
[C6] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
ICLR 2025, Oral Presentation (207/11672=1.77%)
[C3] Visual Representation Learning with Stochastic Frame Prediction
ICML 2024
[C2] Modality-agnostic Self-supervised Learning with Meta-learned Masked Auto-encoder
NeurIPS 2023
[C1] Unsupervised Meta-learning via Few-shot Pseudo-supervised Contrastive Learning
ICLR 2023, Spotlight Presentation (280/4956=5.6%)
NeurIPSW-MetaLearn 2022
[T1] RLDX-1 Technical Report
Technical Report, 2026
[C11] Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models
CoRL 2026
[P3] RoboAlign: Learning Test-Time Reasoning for Language-Action Alignment in Vision-Language-Action Models
Arxiv preprint, 2026
[C10] Dual-stream diffusion for world-model augmented vision-language-action model
ICML 2026
[P2] ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
Arxiv preprint, 2025
[C9] Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
NeurIPS 2025
[C8] SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
ECCV 2026
NeurIPSW-SpaVLE 2025
[C7] Efficient Long Video Tokenization via Coordinate-based Patch Reconstruction
CVPR 2025
[C6] Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
ICLR 2025, Oral Presentation (207/11672=1.77%)
[C5] TrackIME: Enhanced Video Point Tracking via Instance Motion Estimation
NeurIPS 2024, Spotlight Presentation (326/15671=2%)
[C4] Adversarial Robustification via Text-to-Image Diffusion Models
ECCV 2024, Oral Presentation (200/8585=2.3%)
[C3] Visual Representation Learning with Stochastic Frame Prediction
ICML 2024
[C2] Modality-agnostic Self-supervised Learning with Meta-learned Masked Auto-encoder
NeurIPS 2023
[C1] Unsupervised Meta-learning via Few-shot Pseudo-supervised Contrastive Learning
ICLR 2023, Spotlight Presentation (280/4956=5.6%)
NeurIPSW-MetaLearn 2022
[P1] AltUB: Alternating Training Method to Update Base Distribution of Normalizing Flow for Anomaly Detection
Arxiv preprint, 2022
Korea Advanced Institute of Science and Technology (KAIST)Mar. 2023 - Present
PhD. Student in Artificial Intelligence
Korea Advanced Institute of Science and Technology (KAIST)Mar. 2019 - Feb. 2023
B.S. in Mathematical Science and Computer Science
AI Research InternOct. 2022 - Oct. 2023
i-SENSConference Reviewer, IJCAI'23; NeurIPS'24-26; ICLR'25-26; ICML'25-26; CVPR'25-26; ICCV'25; ECCV'26; AISTATS'25
Journal Reviewer, IJCV
Travel Award, International Conference on Machine Learning (ICML) 2024Jul. 2024
Travel Award, Conference on Neural Information Processing Systems (NeurIPS) 2023 Dec. 2023
Travel Award, International Conference on Learning Representations (ICLR) 2023 May. 2023
Recipient, Google Conference Scholarships (APAC)May. 2023
Modality-Agnostic Self-Supervised Learning with Meta-Learned Masked Auto-Encoder
Samsung Electronics Device Solution (DS)Oct. 2024
Samsung AI Forum (SAIF) 2023Nov. 2023
Samsung Advanced Institue of Technology (SAIT)Jun. 2023
Unsupervised Meta-learning via Few-shot Pseudo-supervised Contrastive Learning
International Conference on Learning Representations (ICLR)May. 2023