Qwen-Drive-1.0-4B is a 4B parameter vision-language model that unifies 3D perception, visual question and answering, and trajectory planning. Developed with Huazhong University of Science and Technology, the system uses a Qwen3.5-4B backbone with a BEV perception head and a flow-based Planning Expert. It posted a 69.43 average driving VQA score, beating the 58.15 score of the three times larger 12B Gemma4 model.

Performance metrics include a 90.7 score on NAVSIM and 43.95 mAP on nuScenes. The release includes two planner versions, with a reinforcement learning variant that reduces the off-road rate by 50% by adopting more cautious weights. Alibaba released the code, weights, and demo data under an Apache 2.0 license.

Sign in to suggest edits

Key sources

  1. SOURCE@techtechchina“unifies 3D perception, visual Q&A and motion planning while retaining the original VLM architecture”x.com
  2. SUPPORT@askalphaxiv“using a BEV perception head and flow-based Planning Expert”x.com
  3. SUPPORT@eyishazyer“The RL version cut the off-road rate in half”x.com
Markdown