CV

Summary

Currently employed at HKUST(GZ). An MPhil student at HKUST(GZ). B.S. in Computer Science from Fudan University.

Education

  • Artificial Intelligence
    HKUST(GZ)
  • Computer Science
    2026-06
    Fudan University
    GPA: 3.73

Publications

  • Paper Title Number 1
    2009
    Journal 1
    This paper is about the number 1. The number 2 is left for future work.
  • Paper Title Number 2
    2010
    Journal 1
    This paper is about the number 2. The number 3 is left for future work.
  • Paper Title Number 3
    2015
    Journal 1
    This paper is about the number 3. The number 4 is left for future work.
  • Paper Title Number 4
    2024
    GitHub Journal of Bugs
    This paper is about fixing template issue #693.
  • Paper Title Number 5, with math $$E=mc^2$$
    2024
    GitHub Journal of Bugs
    This paper is about a famous math equation, $$E=mc^2$$
  • EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
    2025
    AAAI 2026 (Oral)
    We introduce EmoVid, the first multimodal, emotion-annotated video dataset specifically designed for creative media. It enables emotion-conditioned video generation by fine-tuning the Wan2.1 model.
  • EmoSpace: Fine-Grained Emotion Prototype Learning for Immersive Affective Content Generation
    2026
    arXiv preprint
    EmoSpace learns dynamic, interpretable emotion prototypes through vision-language alignment, enabling fine-grained emotional control for immersive VR content generation, including emotional image outpainting, stylized generation, and emotional panorama generation.
  • Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models
    2026
    arXiv preprint
    We isolate generation-to-understanding and understanding-to-generation transfer in unified multimodal models by binding a novel visual concept through exactly one task direction, finding that cross-task usability is governed by where the binding enters the shared computation.
  • InternLumina-U2: A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding and Image Generation
    2026
    Technical report (in preparation)
    InternLumina-U2 is a unified diffusion large language model that shares a single MoE diffusion backbone with a multi-codebook visual tokenizer for text question answering, image understanding and editing, image generation, video understanding, and 3D understanding.

Presentations

  • Talk 1 on Relevant Topic in Your Field
    2012
    UC San Francisco, Department of Testing
    San Francisco, CA, USA
  • Tutorial 1 on Relevant Topic in Your Field
    2013
    UC-Berkeley Institute for Testing Science
    Berkeley, CA, USA
  • Talk 2 on Relevant Topic in Your Field
    2014
    London School of Testing
    London, UK
  • Conference Proceeding talk 3 on Relevant Topic in Your Field
    2014
    Testing Institute of America 2014 Annual Conference
    Los Angeles, CA, USA

Teaching

  • Teaching experience 1
    2014
    University 1, Department
    Role: Undergraduate course
  • Teaching experience 2
    2015
    University 1, Department
    Role: Workshop

Portfolio

  • Portfolio item number 1
    Portfolio
    Short description of portfolio item number 1