I am currently a PhD student in the Department of Data Science and AI at Monash University, luckily advised by Prof. Tien-Tsin Wong and Prof. Jianfei Cai. I received B.S. and M.Phil. in Computer Science from Beijing Jiaotong University.

I am interested in Image/Video Generation and World Models.

πŸ“– Education

  • 2024.11 - now, Doctor of Philosophy, Department of Data Science and AI, Monash University, Melbourne.
  • 2020.06 - 2024.06, Master of Computer Science, School of Computer and Information Technology, Beijing Jiaotong University, Beijing.
  • 2016.09 - 2020.06, Bachelor of Computer Science, School of Computer and Information Technology, Beijing Jiaotong University, Beijing.

πŸ”₯ News

  • 2026.10: Β πŸ”₯ We release Oneira, an open-world interaction video world model.
  • 2026.06: Β πŸŽ‰ One paper ATA is accepted by ECCV 2026.
  • 2025.09: Β πŸ”₯ We release the HunyuanImage 3.0 Technical Report.
  • 2025.06: Β πŸŽ‰ One paper VLIPP is accepted by ICCV 2025.
  • 2024.01: Β πŸŽ‰ One paper Neural Field Classifier is accepted by ICLR 2024.
  • 2023.08: Β πŸŽ‰ SDFStudio has supported S3IM.
  • 2023.08: Β πŸ”₯ We release S3IM(⭐️200+).
  • 2023.07: Β πŸŽ‰ One paper S3IM is accepted by ICCV 2023.

πŸ“ Publications and Manuscripts

arXiv 2026
sym

Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models.

Xindi Yang, Baolu Li, Liam Lee, Zhenfei Yin, Songxin Zhang, Zhuoyang Song, Xu Jia, Jianfei Cai, Tien-Tsin Wong, Bingyi Jing, Mengyue Yang.

Project Code

  • Oneira is an interactive video world model built around an explicit world state managed by a coding agent. It enables instance-level interaction with objects that emerge during open-world exploration, and the consequences of each interaction persist across long horizons.
ECCV 2026
sym

Attribute Token Arithmetic: Disentangled and Continuous Semantic Control for Visual Autoregressive Models.

Xindi Yang, Yicheng Wu, Cheng Zhang, Jianfei Cai, Tien-Tsin Wong.

arXiv

  • ATA finds attribute directions (e.g., aging, fatness, emotion) in the latent space of a pretrained visual autoregressive model from a single reference image, enabling disentangled, continuous and composable attribute control without retraining.
Technical Report
sym

HunyuanImage 3.0 Technical Report.

Core Contributor

arXiv Code

  • Hunyuan Foundation Model.
ICCV 2025
sym

VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior.

Xindi Yang*, Baolu Li*, Yiming Zhang, Zhenfei Yin, Lei Bai, Liqian Ma, Zhiyong Wang, Jianfei Cai, Tien-Tsin Wong, Huchuan Lu, Xu Jia (*equal contribution).

Project Code

  • VLIPP is a two-stage image-to-video generation framework that explicitly incorporates physics with vision and language informed physical prior.
ICLR 2024
sym

Neural Field Classifiers via Target Encoding and Classification Loss.

Xindi Yang, Zeke Xie, Xiong Zhou, Boyu Liu, Buhua Liu, Yi Liu, Haoran Wang, Yunfeng Cai, Mingming Sun.

  • Neural Field Classifiers via Target Encoding and Classification Loss can significantly outperform the standard regression-based neural field counterparts.
ICCV 2023
sym

S3IM: Stochastic Structural SIMilarity and Its Unreasonable Effectiveness for Neural Fields.

Zeke Xie*, Xindi Yang*, Yujie Yang, Qi Sun, Yixiang Jiang, Haoran Wang, Yunfeng Cai, Mingming Sun (*equal contribution).

Project

  • S3IM is a plug-and-play loss, effective and robust in various difficult tasks.
  • Academic Impact: Our work has been featured by 4+ media and forums, such as ηŸ₯乎, ζžδΈ–εΉ³ε°

πŸŽ– Honors and Awards

  • 2023, Outstanding Intern of the Year, Baidu Research
  • 2016-2022, Model Student of Academic Records of Beijing Jiaotong University
  • 2018, National Contemporary Undergraduate Mathematical Contest IN Modeling in China, First Prize in Beijing region

πŸ“ Academic Service

  • Journal Review: IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), IEEE Transactions on Visualization and Computer Graphics (TVCG), Computer Graphics Forum (CGF)
  • Conference Review: ICLR, NeurIPS, ICCV, ECCV, CVPR