InternLumina-U2: A Multi-Codebook Diffusion Large Language Model for Omni-Visual Understanding and Image Generation

Published in Technical report (in preparation), 2026

Abstract: InternLumina-U2 unifies language, images, video, and 3D within a single multi-codebook diffusion large language model, sharing one MoE diffusion backbone across text question answering, image understanding and editing, image generation, video understanding, and 3D understanding. A multi-codebook visual tokenizer represents visual content as discrete VQ levels, processing spatial positions in parallel while sequentially decoding codebooks, enabling omni-visual understanding and generation within one model. The technical report is forthcoming; code is available on GitHub.