Multimodality must be native.
Vision, video, documents, audio, interfaces, and space belong inside the model’s reasoning loop—not as an attachment added after the text model is built.
MULTIMODEL-NATIVE INTELLIGENCE
Lead, Qwen-VL & Visual Intelligence · Alibaba Qwen Team
Building multimodal-native intelligence that can understand the world, use tools, and act across digital and physical environments.
RESEARCH PERSPECTIVE
My work connects foundation models, agent systems, and the environments where intelligence is ultimately useful.
Across perception, reasoning, tool use, and action, the goal is a coherent system rather than a collection of isolated capabilities.
Vision, video, documents, audio, interfaces, and space belong inside the model’s reasoning loop—not as an attachment added after the text model is built.
A multimodal agent should perceive precisely, decide when to use tools, interact with environments, inspect the result, and recover when reality disagrees.
The same foundation should move across browsers, documents, software, video, 3D and CAD—and ultimately from digital action to physical intelligence.
MODEL EVOLUTION
Perception becomes useful when it is fused with reasoning, coding, environment interaction, and long-horizon execution.
QWEN 3.5
Multimodality became part of the general model itself: understand visual inputs, reason, code, and use tools in one agentic foundation.
RELEASE NOTESQWEN 3.6
The focus moved from isolated capabilities to agents that can work through longer, messier, real-world tasks.
RELEASE NOTESQWEN 3.7 PLUS
Stronger planning, environment interaction, and multimodal execution pushed Qwen further into the agent era.
RELEASE NOTESQWEN 3.8 MAX
A 2.4T-scale foundation for coding and cowork—where visual context, tool use, and long-horizon execution converge.
RELEASE NOTESNative multimodal models for the agent era
Bringing visual intelligence into Qwen’s general agent foundation—connecting perception with reasoning, coding, cowork, and tool use.
Qwen‑VL · Qwen2‑VL · Qwen2.5‑VL · Qwen3‑VL
A sustained program in visual understanding, grounding, documents, long video, spatial reasoning, computer use, and visual agents.
Vision · Language · Action
Extending multimodal intelligence beyond screens into action across tasks, environments, and robot embodiments.
Understand · Generate · Edit
Unifying visual understanding and generation so a model can move naturally from interpreting the world to depicting and editing it.
Make any agent harness multimodal-native
A portable capability layer for images, video, documents, web research, long-video memory, editing, 3D, CAD, and education workflows.
SELECTED COLLABORATIONS
Qwen2.5‑Omni · Qwen3‑Omni · Qwen3.5‑Omni
End-to-end text, image, audio, and video understanding with real-time natural speech interaction.
Native text rendering · generation · editing
An image foundation model connecting strong visual semantics with complex text rendering and precise editing.
One architecture · many tasks · many modalities
An early unified sequence-to-sequence framework for vision, language, grounding, and generation.
Selected systems, not a complete publication list.
FULL PUBLICATION LIST ON GOOGLE SCHOLAR PAPERS ON HUGGING FACETHE CAPABILITY LAYER
Qwen-MM-Plugins turns multimodal intelligence into portable, usable capabilities across the agent tools people already work with.
EXPLORE QWEN-MM-PLUGINSWe welcome exceptional interns, new graduates, experienced researchers, and collaborators working on multimodal models, agents, computer use, video, documents, 3D, world models, and VLA.
Signal over pedigree. Show what you built, what you discovered, or the question you cannot stop thinking about.
START A CONVERSATIONCONNECT
GLOBAL VISITORS
This map shows approximate daily visits by world region. One visit is counted per browser each day; only the visitor's country region is used.