MULTIMODEL-NATIVE INTELLIGENCE

Shuai Bai / 白帅

Lead, Qwen-VL & Visual Intelligence · Alibaba Qwen Team

Building multimodal-native intelligence that can understand the world, use tools, and act across digital and physical environments.
FOCUS Multimodal models & agents SCOPE Digital ↔ Physical
FOUNDATION MODELS MULTIMODAL AGENTS TOOL USE DIGITAL + PHYSICAL INTELLIGENCE
01 / PERSPECTIVE MODEL · AGENT · WORLD

RESEARCH PERSPECTIVE

My work connects foundation models, agent systems, and the environments where intelligence is ultimately useful.

Across perception, reasoning, tool use, and action, the goal is a coherent system rather than a collection of isolated capabilities.

01 MODEL

Multimodality must be native.

Vision, video, documents, audio, interfaces, and space belong inside the model’s reasoning loop—not as an attachment added after the text model is built.

02 AGENT

Understanding must lead to action.

A multimodal agent should perceive precisely, decide when to use tools, interact with environments, inspect the result, and recover when reality disagrees.

03 WORLD

One intelligence, many environments.

The same foundation should move across browsers, documents, software, video, 3D and CAD—and ultimately from digital action to physical intelligence.

02 / MODEL EVOLUTION FROM NATIVE MULTIMODALITY TO COWORK

MODEL EVOLUTION

Qwen 3.5 → Qwen 3.8 Max

Perception becomes useful when it is fused with reasoning, coding, environment interaction, and long-horizon execution.

03 / WORK MODEL · SYSTEM · ENVIRONMENT

CORE PROGRAMS

Selected programs

MORE WORK

SELECTED COLLABORATIONS

Collaborative & participating work

Selected systems, not a complete publication list.

FULL PUBLICATION LIST ON GOOGLE SCHOLAR PAPERS ON HUGGING FACE
04 / OPEN SYSTEM QWEN-MM-PLUGINS

THE CAPABILITY LAYER

A multimodal capability layer for agent systems.

Qwen-MM-Plugins turns multimodal intelligence into portable, usable capabilities across the agent tools people already work with.

EXPLORE QWEN-MM-PLUGINS
CAPABILITY MAP 08 MODULES
01 Vision + documents
02 Audio + video understanding
03 Web + image search
04 Long-video memory
05 Video editing
06 Blender / 3D
07 FreeCAD / engineering
08 Education agents
05 / JOIN BEIJING · CHINA
WE'RE HIRING

Join the Qwen multimodal team.

We welcome exceptional interns, new graduates, experienced researchers, and collaborators working on multimodal models, agents, computer use, video, documents, 3D, world models, and VLA.

Signal over pedigree. Show what you built, what you discovered, or the question you cannot stop thinking about.

START A CONVERSATION

CONNECT

Contact

06 / VISITOR MAP GLOBAL · APPROXIMATE · COOKIELESS

GLOBAL VISITORS

Ideas travel.

This map shows approximate daily visits by world region. One visit is counted per browser each day; only the visitor's country region is used.

VISITS SINCE AUG 2026
REGIONAL FOOTPRINT CONNECTING
COUNTRY-LEVEL LOOKUP · REGIONAL COUNTS ONLY · NO COOKIE CC0 MAP · NATURAL EARTH / WIKIMEDIA