Biography

I am a researcher at Alibaba Tongyi Lab, where I lead Z-Image's reinforcement learning alignment. My work mainly focuses on foundation model post-training, spanning SFT, RL, and training infrastructure, while also contributing to pre-training for advanced model capabilities. My research centers on multimodal generation and understanding. Earlier, I worked on multimodal large language models at Tencent Hunyuan, with a focus on improving their advanced reasoning capabilities.

News

  • 2026 New paper on pixel-space text-to-image diffusion models released on arXiv.
  • 2026 Released Z-Reward, a reasoning-internalized teacher-student reward modeling framework for text-to-image post-training.
  • 2026 Released Z-Image (Base), the high-quality foundation model specializing in rich aesthetics and controllability. The series is now the #2 most popular Chinese model on Hugging Face and ranks 2nd in global Image Generation API usage on fal.ai.
  • 2025 Released Z-Image-Turbo, an efficient 6B foundation model specializing in photorealistic image generation. It ranked #1 among open-source text-to-image models on the Artificial Analysis Image Arena, while achieving sub-second inference latency.
  • 2024 Released MM-IQ, a new benchmark for assessing the core reasoning capabilities of large multimodal models.
  • 2024 New paper on Self-Correction in LLMs released on arXiv.
  • 2024 Paper System-2 Mathematical Reasoning accepted to TMLR.
  • 2023 The extended paper Tri-token Equipped Transformer Model for Image Matting released on arXiv.
  • 2022 TransMatting accepted to ECCV.
  • 2021 Won the 2nd Place Award in NTIRE 2021 Challenge on Multi-modal Aerial View Object Classification at CVPR 2021.

Publications

Selected work

More publications

5 papers
Showing 2 of 5Showing all 5 Show 3 more publicationsShow fewer publications