I am Chufan Shi, a PhD student at the University of Southern California since 2025, fortunate to be advised by Prof. Xuezhe Ma.
Before USC, I received my master's degree from Tsinghua University, advised by Prof. Yujiu Yang, and my bachelor's degree from the University of Electronic Science and Technology of China.

My research aims to build foundation models that scale efficiently and reason reliably. Specifically, I work on:
Selected publications (* equal contribution); see my Google Scholar profile for the full list.
I am a core contributor to IFM's K2 Horizon models and the xLLM pre-training infrastructure. I also explore new designs for mixture-of-experts, attention, and looped models toward more efficient scaling.

K2 Horizon: Frontier Performance, Radically Open IFM K2 Team (Chufan Shi, Core Contributor) Technical Blog 2026 [blog] [models]

xLLM: Light-weight Infrastructure for Express LLM Pre-training IFM K2 Team (Chufan Shi, Core Contributor) Open-Source Library 2026 [code]

EMO: Frustratingly Easy Progressive Training of Extendable MoE Linghao Jin, Chufan Shi, Huijuan Wang, Nuan Wen, Zhengzhong Liu, Eric Xing, Xuezhe Ma Preprint 2026 [paper]

SoftDelta Attention: The Write Way to Attend Shicheng Wen, Chufan Shi, Linghao Jin, Zhengzhong Liu, Eric Xing, Xuezhe Ma Blog 2026 [blog]

Towards Looped Models Done Right — Part I: Topology, Input Injection, Recurrent-State Design Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu, Eric Xing, Xuezhe Ma Blog 2026 [blog]

Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast Chufan Shi, Cheng Yang, Xinyu Zhu, Jiahao Wang, Taiqiang Wu, Siheng Li, Deng Cai, Yujiu Yang, Yu Meng NeurIPS 2024 [paper] [code]
I study whether multimodal models truly ground their reasoning in what they see, and whether unified models generate in line with what they reason about. I develop post-training methods to close these gaps, toward multimodal agents that act reliably in complex environments.

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination Chufan Shi, Cheng Yang, Yaokang Wu, Linghao Jin, Bo Shui, Taylor Berg-Kirkpatrick, Xuezhe Ma ICML 2026 (Oral) [paper] [code]

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs Xin Gao*, Cheng Yang*, Chufan Shi*, Taylor Berg-Kirkpatrick ICML 2026 [paper] [code]

UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models Cheng Yang*, Chufan Shi*, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick EMNLP 2026 (Oral) [paper] [project]