👋 Hi!

I am Chufan Shi, a PhD student at the University of Southern California since 2025, fortunate to be advised by Prof. Xuezhe Ma.

Before USC, I received my master's degree from Tsinghua University, advised by Prof. Yujiu Yang, and my bachelor's degree from the University of Electronic Science and Technology of China.

Snipaste_2026-06-26_16-12-46.png

Email Twitter

My research aims to build foundation models that scale efficiently and reason reliably. Specifically, I work on:


📜 Publications

Selected publications (* equal contribution); see my Google Scholar profile for the full list.

Scalable Architectures

I am a core contributor to IFM's K2 Horizon models and the xLLM pre-training infrastructure. I also explore new designs for mixture-of-experts, attention, and looped models toward more efficient scaling.

k2_horizon.png

K2 Horizon: Frontier Performance, Radically Open IFM K2 Team (Chufan Shi, Core Contributor) Technical Blog 2026 [blog] [models]

xllm_pretraining.png

xLLM: Light-weight Infrastructure for Express LLM Pre-training IFM K2 Team (Chufan Shi, Core Contributor) Open-Source Library 2026 [code]

emo_overview.png

EMO: Frustratingly Easy Progressive Training of Extendable MoE Linghao Jin, Chufan Shi, Huijuan Wang, Nuan Wen, Zhengzhong Liu, Eric Xing, Xuezhe Ma Preprint 2026 [paper]

softdelta_attention.png

SoftDelta Attention: The Write Way to Attend Shicheng Wen, Chufan Shi, Linghao Jin, Zhengzhong Liu, Eric Xing, Xuezhe Ma Blog 2026 [blog]

looped_models.png

Towards Looped Models Done Right — Part I: Topology, Input Injection, Recurrent-State Design Benhao Huang, Chufan Shi, Junlin Chen, Shicheng Wen, Zhengzhong Liu, Eric Xing, Xuezhe Ma Blog 2026 [blog]

image.png

Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast Chufan Shi, Cheng Yang, Xinyu Zhu, Jiahao Wang, Taiqiang Wu, Siheng Li, Deng Cai, Yujiu Yang, Yu Meng NeurIPS 2024 [paper] [code]

Multimodal Intelligence

I study whether multimodal models truly ground their reasoning in what they see, and whether unified models generate in line with what they reason about. I develop post-training methods to close these gaps, toward multimodal agents that act reliably in complex environments.

image.png

Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination Chufan Shi, Cheng Yang, Yaokang Wu, Linghao Jin, Bo Shui, Taylor Berg-Kirkpatrick, Xuezhe Ma ICML 2026 (Oral) [paper] [code]

image.png

Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs Xin Gao*, Cheng Yang*, Chufan Shi*, Taylor Berg-Kirkpatrick ICML 2026 [paper] [code]

Snipaste_2026-06-26_16-39-48.png

UReason: Benchmarking Reasoning-to-Generation Alignment in Unified Multimodal Models Cheng Yang*, Chufan Shi*, Bo Shui, Yaokang Wu, Muzi Tao, Huijuan Wang, Ivan Yee Lee, Yong Liu, Xuezhe Ma, Taylor Berg-Kirkpatrick EMNLP 2026 (Oral) [paper] [project]