Zongyang (Zane) Qiu|邱宗扬

MPhil Student in Artificial Intelligence @ HKUST(GZ)

zane.zy.qiu [at] gmail.com

I'm Zongyang Qiu, an MPhil student at HKUST(GZ), advised by Prof. Hui Xiong, and co-advised by Prof. Zeyu Wang. I received my B.S. degree in Computer Science from Fudan University.

Currently, I am a student researcher at Shanghai Innovation Institute (SII), supervised by Yutao Sun, working on world models, specifically VLMs for embodied models. Previously, I was a research intern at InternLM Group, Shanghai AI Lab, supervised by Dr. Bin Fu, focusing on building the next-generation DLLM and Visual Agents. Before that, I worked with Dr. Zeyu Wang at CIS Lab, HKUST(GZ), and with Prof. Wenhan Luo at C4 Group, HKUST.

I also work closely with Yijiang Li and Nuno Vasconcelos at UCSD in the field of video reasoning.

I am deeply passionate about extensive topics in Computer Vision and Computer Graphics. I aim to explore the construction of biological-like visual intelligence, which should be capable of autonomous understanding, decision-making, and generation within both physical and digital worlds. My current research focuses on multimodal learning of VLMs and reasoning ability in visual generative models.

  • Multimodal Learning
  • Vision-Language Models
  • Visual Generative Models
  • Computer Vision & Graphics

News

[2026/04] Honored to be named a Shanghai Outstanding Graduate (Top 5%).

[2026/01] Delighted to attend AAAI 2026 in Singapore and give an oral presentation.

Publications

InternLumina-U2: A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding and Image Generation teaser

InternLumina-U2: A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding and Image Generation

Intern Lumina U2 Team, Shanghai AI Laboratory

Technical report (in preparation)

InternLumina-U2 unifies language, images, video, and 3D within a single multi-codebook diffusion large language model, sharing one MoE diffusion backbone across text question answering, image understanding and editing, image generation, video understanding, and 3D understanding. A multi-codebook visual tokenizer represents visual content as discrete VQ levels, processing spatial positions in parallel while sequentially decoding codebooks, enabling omni-visual understanding and generation within one model. The technical report is forthcoming; code is available on GitHub.
Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models teaser

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong

arXiv preprint

Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat. Joint-training studies cannot settle the disagreement: with overlapping supervision, a gain cannot be attributed to the architecture rather than the data. To further investigate the relationship between the two directions in UMMs, we separate them by construction. A novel visual entity, a rendered 3D asset paired with a pseudo-word screened for absence from the frozen model’s behavior, is bound through exactly one task direction, and the untrained direction is then measured. We find that the channel is real in both directions, but the directions differ in kind: generation training installs a name the model can only match among candidates; understanding training installs one it can also produce. What governs cross-task usability is where the binding enters the shared computation. An alignment probe predicts export across 36 configurations (Spearman ρ = +0.68). That objective’s alignment term, maximized in closed form over activations with every weight frozen, makes a concept drawable when injected at layer 7 of 28 and is indistinguishable from the base model from layer 14 on, while the weight-based version of the same edit peaks at layers 10-14. In an observational series of four models, this window appears only where the understanding pathway is a semantic vision encoder, suggesting that unified weights are not enough: the two directions must share a semantic format at the entry point. Exploiting the rule, a mid-stack alignment objective acquires the concept for a 0.1% relative loss of the model’s general text-to-image ability, against 41% for the standard generative route.
EmoSpace: Fine-Grained Emotion Prototype Learning for Immersive Affective Content Generation teaser

EmoSpace: Fine-Grained Emotion Prototype Learning for Immersive Affective Content Generation

Bingyuan Wang, Xingbei Chen, Zongyang Qiu, Linping Yuan, Zeyu Wang

arXiv preprint

Emotion is important for creating compelling virtual reality (VR) content. Although some generative methods have been applied to lower the barrier to creating emotionally rich content, they fail to capture the nuanced emotional semantics and the fine-grained control essential for immersive experiences. To address these limitations, we introduce EmoSpace, a novel framework for emotion-aware content generation that learns dynamic, interpretable emotion prototypes through vision-language alignment. We employ a hierarchical emotion representation with rich learnable prototypes that evolve during training, enabling fine-grained emotional control without requiring explicit emotion labels. We develop a controllable generation pipeline featuring multi-prototype guidance, temporal blending, and attention reweighting that supports diverse applications, including emotional image outpainting, stylized generation, and emotional panorama generation for VR environments. Our experiments demonstrate the superior performance of EmoSpace over existing methods in both qualitative and quantitative evaluations. Additionally, we present a comprehensive user study investigating how VR environments affect emotional perception compared to desktop settings. Our work facilitates immersive visual content generation with fine-grained emotion control and supports applications like therapy, education, storytelling, artistic creation, and cultural preservation. Code and models will be made publicly available.
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation teaser

EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation

Zongyang Qiu, Bingyuan Wang, Xingbei Chen, Yingqing He, Zeyu Wang

AAAI 2026 (Oral)

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual domain, the video community lacks dedicated resources to bridge emotion understanding with generative tasks, particularly for stylized and non-realistic contexts. To address this gap, we introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for creative media, which includes cartoon animations, movie clips, and animated stickers. Each video is annotated with emotion labels, visual attributes (brightness, colorfulness, hue), and text captions. Through systematic analysis, we uncover spatial and temporal patterns linking visual features to emotional perceptions across diverse video forms. Building on these insights, we develop an emotion-conditioned video generation technique by fine-tuning the Wan2.1 model. The results show a significant improvement in both quantitative metrics and the visual quality of generated videos for text-to-video and image-to-video tasks. EmoVid establishes a new benchmark for affective video computing. Our work not only offers valuable insights into visual emotion analysis in artistically styled videos, but also provides practical methods for enhancing emotional expression in video generation.

Experience

Education

Fudan University (复旦大学)

2022.09 – 2026.06

B.S. in Computer Science · GPA 3.74/4.0

Coursework: Programming (A), Artificial Intelligence (Honors, A), Digital Image Processing (A), Set and Graph Theory (Honors, A)

Awards and Honors

2026

Shanghai Outstanding Graduate (Top 5%)

2025

National Scholarship (Top 1%, ranked 1st in the school)

2025

Meritorious Winner, Mathematical Contest in Modeling (Top 7%)

2024

Gold Award, China International College Student Innovation Competition

Service

Reviewer

AAAI Conference on Artificial Intelligence (AAAI 2026)
IEEE Transactions on Affective Computing