Qwen-Drive

Qwen-Drive-1.0

by Qwen · Current open-weight research release

An open Apache 2.0 Qwen model for autonomous-driving research, combining multimodal scene understanding with 3D BEV perception, visual question answering, and five-second ego-vehicle trajectory generation.

Text Actions Reasoning Coding
Qwen-Drive-1.0 extends a Qwen3.5-4B vision-language model with dedicated perception and planning components for driving scenes. It can interpret visual and textual inputs, produce 3D perception results, answer questions about road environments, and generate future vehicle trajectories. The downloadable release is aimed at research and experimentation rather than hosted API use or safety-certified vehicle control.
Outputs

What Qwen-Drive-1.0 can produce

Text Actions
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

6/10 Reasoning
2/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Qwen-Drive
Model type Multimodal
Context window tokens
Maximum output tokens
Release date 2026-08-31
Status Current open-weight research release
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was published in the reviewed official model card, repository, technical report, or announcement.

Model notes

Qwen-Drive-1.0 is a downloadable model and inference-code release, not a separately priced hosted API model. The released Qwen-Drive-1.0-4B package uses a Qwen3.5-4B vision-language model plus separate perception, planner-sft, and planner-rl components. The planning expert is a separate 1.0B-parameter network. The documented planner predicts 50 future waypoints at 10 Hz, corresponding to a five-second trajectory. The model produces text, perception predictions, and trajectory/action outputs, but does not generate images, audio, or video. Editorial scores reflect its specialized autonomous-driving role rather than general LLM performance.

Model guide

Qwen-Drive-1.0: An Open Vision-Language Model for Autonomous Driving

Qwen-Drive-1.0 is an open Apache 2.0 vision-language foundation model for autonomous-driving research. Built around Qwen3.5-4B, it combines visual question answering, 3D bird's-eye-view perception, semantic occupancy prediction, map segmentation, and future ego-vehicle trajectory planning in one research-oriented system.

What is Qwen-Drive-1.0?

Qwen-Drive-1.0 is an open vision-language foundation model designed specifically for autonomous-driving research. It is provided by the Qwen team and Huazhong University of Science and Technology, and is released under the Apache 2.0 license. Rather than serving as a general-purpose chat model or a complete self-driving stack, it brings visual understanding, 3D scene perception, and motion planning into a shared research framework.

The model is built around Qwen3.5-4B, a vision-language model that can process visual and textual information. Qwen-Drive-1.0 adds task-specific components around that base: a bird's-eye-view perception head for interpreting the 3D driving environment and a separate Planning Expert for predicting the future motion of the ego vehicle, meaning the vehicle carrying the system.

This design makes Qwen-Drive-1.0 useful for experiments in embodied AI, driving-scene understanding, visual driving assistants, simulation, and planning evaluation. Its outputs are not limited to text. Depending on the operating mode, it can produce text answers, structured perception predictions, and future trajectory or action signals.

How the model is structured

Qwen-Drive-1.0 keeps the pretrained Qwen3.5-4B vision-language architecture unchanged and attaches additional modules for autonomous-driving tasks. This separation is important: the base model supplies multimodal understanding, while the perception and planning components expose more specialized driving capabilities.

The shared pathway can work with driving images, multi-view images, temporal image sequences, and general visual inputs. A bird's-eye-view, or BEV, perception head converts the shared representations into a top-down description of the surrounding scene. In practical terms, BEV outputs can help represent where objects are located around the vehicle, which areas are occupied, and how the road layout is organized.

The perception head jointly supports three documented tasks:

  • 3D object detection: identifying objects and their positions in three-dimensional driving space.
  • Semantic occupancy prediction: estimating which regions around the vehicle are occupied and what they represent.
  • BEV map segmentation: labeling road and map-related elements from a bird's-eye-view perspective.

Planning is handled by a separate Planning Expert. The release includes an imitation-trained planner-sft component and a reward-optimized planner-rl component. The documented Planning Expert is a separate 1.0-billion-parameter network and uses flow matching to generate future ego-vehicle trajectories. Textual reasoning can be used as an optional planning condition.

What can Qwen-Drive-1.0 do?

Visual question answering

In visual question-answering mode, the base Qwen3.5-4B vision-language model answers free-form questions about images or driving scenes. A user could, for example, ask what objects are visible, whether a road area appears occupied, or what is happening in a particular scene. This mode is closer to multimodal scene interpretation than to direct vehicle control.

The model can also handle textual reasoning associated with a visual input. However, the supplied research does not establish a fixed reasoning-token limit, a formal reasoning benchmark, or a guaranteed chain-of-thought interface. Its reasoning capability should therefore be understood as task-oriented visual and textual inference rather than as a separately measured general reasoning product.

3D perception

The perception mode exposes driving-specific 3D information through object detection, semantic occupancy, and BEV map segmentation. This is a stronger fit for autonomous-driving experiments than a general image-language model that only describes a picture in prose.

Because these outputs depend on the perception head and the supplied scene data, users should not assume that the base Qwen3.5-4B model alone provides all of the documented 3D functions. The released package separates the base vision-language model from the perception and planning directories.

Motion planning

The planning mode generates future ego-vehicle trajectories. The documented configuration predicts 50 future waypoints at 10 Hz, corresponding to a five-second trajectory. Each waypoint represents position and heading information, giving the system a sequence that can be evaluated in a simulator or used in research pipelines.

The two planner variants serve different training approaches. The supervised planner-sft component supports direct and reasoning-based planning, while the reinforcement-optimized planner-rl component is intended for reasoning-conditioned planning. These names describe the released components; they do not imply that either planner is ready to control a road vehicle without additional safety, validation, and systems engineering.

Supported inputs and outputs

Qwen-Drive-1.0 is multimodal on input. The supplied documentation describes support for visual and textual inputs, including driving images, multi-view images, and temporal image sequences. The model is not documented as an audio or video-generation system, and the available research does not verify native audio input or video-file processing as a separate modality.

Its output types are more specialized than those of a standard text-only language model:

  • Text: answers to visual and driving-related questions, including optional textual reasoning.
  • Perception predictions: 3D detections, semantic occupancy results, and BEV map segmentation.
  • Trajectory and action outputs: future ego-vehicle waypoints containing position and heading information.

It does not natively generate images, audio, music, or video. The research also does not verify a general-purpose function-calling interface, web search, a standardized JSON mode, or a hosted tool-use API. Perception and trajectory outputs are task-specific model results rather than evidence of a general agent framework.

Limits, pricing, and deployment

No authoritative context-window limit or maximum text-output-token limit is provided in the supplied model information. Users should not assume that Qwen-Drive-1.0 has the same limits as a hosted Qwen API product. Input capacity will also depend on the selected inference implementation, image or sequence representation, available memory, and the task-specific component being used.

Qwen-Drive-1.0 is distributed as downloadable weights and inference code, not as a separately priced metered hosted API model. Consequently, there is no verified input price, output price, monthly subscription, or per-request billing rate to report. The Hugging Face package contains an approximately 9.1 GB base vision-language model together with separate planner and perception directories. That download size is not an operating-cost estimate: local inference still depends on hardware, memory, storage, and runtime configuration.

The official project documents workflows involving Transformers, vLLM, and SGLang-compatible serving, alongside the project's own inference code. These options may help users integrate the model into research infrastructure, but the supplied information does not establish a managed Qwen-Drive endpoint, service-level agreement, or production support commitment.

Strengths and trade-offs

The clearest strength of Qwen-Drive-1.0 is its focus. It is not merely a vision-language model asked to describe road images. Its released components explicitly address 3D perception and future motion planning, allowing researchers to study how visual-language representations connect with driving-specific outputs.

The open Apache 2.0 release is another practical advantage for research teams that need downloadable weights and inference code. It permits local experimentation without a per-call API price, although teams remain responsible for compute and deployment costs. The modular structure also makes it easier to distinguish the base multimodal model from the perception and planning components.

There are important trade-offs. A specialized driving model is a poor substitute for a general coding assistant, general-purpose reasoning model, or conventional hosted API when those are the primary requirements. The editorial assessment supplied with the research rates its coding suitability low because coding is outside its central purpose. Its speed and cost also depend heavily on local hardware and the selected planner or perception pipeline; there is no hosted latency or price specification to use for a direct API comparison.

The model's outputs are likewise conditional on camera configuration, calibration data, scene representation, runtime hardware, and task-specific heads. The project documentation identifies consistency between textual reasoning and generated trajectories as an area for further improvement. This is especially relevant when interpreting a model that can explain a scene and also propose a trajectory: plausible text does not by itself validate the safety or correctness of the corresponding motion plan.

When to choose Qwen-Drive-1.0

Choose Qwen-Drive-1.0 when the project needs an open model centered on autonomous-driving perception and planning rather than a general chatbot or hosted multimodal API. It is particularly suitable for:

  • research on driving-scene visual question answering;
  • experiments with 3D object detection, occupancy, or BEV map segmentation;
  • simulation-based evaluation of future ego-vehicle trajectories;
  • investigations into reasoning-conditioned motion planning;
  • embodied-AI projects that need downloadable multimodal weights and task-specific outputs.

Another type of option may be more appropriate when the priority is general coding, broad tool integration, predictable hosted latency, simple per-request billing, or production support. A general multimodal API is likely easier for ordinary image-question answering and application development, while a specialized robotics or autonomous-driving stack may provide more mature interfaces, calibration handling, safety controls, and vehicle-integration tooling.

Qwen-Drive-1.0 should not be treated as a complete production-ready autonomous-driving system. It is also not safety-certified vehicle-control software. Any use involving real vehicles would require extensive validation, redundant sensing and control systems, operational safeguards, and compliance work beyond the model release itself.

Practical assessment

Qwen-Drive-1.0 occupies a specialized position in the Qwen lineup: an open, research-oriented driving model built on Qwen3.5-4B and extended with perception and planning experts. Its value comes from combining visual-language understanding with explicit 3D and trajectory outputs, not from general-purpose language or coding performance.

For researchers who can run downloadable weights and want to explore the connection between scene understanding, BEV perception, reasoning, and motion planning, the release provides a focused starting point. For users seeking a simple hosted service, a fixed context and output contract, or a production autonomous-driving controller, the available information does not support those expectations.


Answers to Frequently Asked Questions

Is Qwen-Drive-1.0 a production-ready autonomous-driving system?
No. Qwen-Drive-1.0 is a research-oriented model rather than a complete or safety-certified autonomous-driving stack. Real-vehicle use would require extensive validation, redundant sensing and control systems, operational safeguards, calibration, and regulatory compliance. Plausible textual explanations or generated trajectories do not by themselves establish safety.
What inputs and outputs does Qwen-Drive-1.0 support?
The model accepts visual and textual inputs, including driving images, multi-view images, and temporal image sequences. It can produce text answers, 3D perception predictions such as detections and occupancy maps, BEV map segmentation, and future trajectory waypoints. It is not documented as a native audio, image-generation, or video-generation system.
How does Qwen-Drive-1.0 generate motion plans?
Qwen-Drive-1.0 uses a separate 1.0-billion-parameter Planning Expert with flow matching to generate future ego-vehicle trajectories. The documented configuration predicts 50 waypoints at 10 Hz, representing a five-second trajectory with position and heading information. The release includes supervised planner-sft and reward-optimized planner-rl variants.
What is Qwen-Drive-1.0?
Qwen-Drive-1.0 is an open vision-language foundation model for autonomous-driving research. Built on Qwen3.5-4B, it combines multimodal scene understanding with specialized modules for 3D perception and ego-vehicle motion planning. It is released under the Apache 2.0 license.
What autonomous-driving tasks does Qwen-Drive-1.0 support?
It supports visual question answering, 3D object detection, semantic occupancy prediction, BEV map segmentation, and future ego-vehicle trajectory generation. Its perception head produces bird's-eye-view driving-scene outputs, while the Planning Expert predicts future motion.


Sources 6
Provider

About Qwen