Qwen/Qwen3.5-4B 已完整同步
Qwen3.5-4B
Note
This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.
These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency.
Qwen3.5 Highlights
Qwen3.5 features the following enhancement:
-
Unified Vision-Language Foundation: Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.
-
Efficient Hybrid Architecture: Gated Delta Networks combined with sparse Mixture-of-Experts deliver high-throughput inference with minimal latency and cost overhead.
-
Scalable RL Generalization: Reinforcement learning scaled across million-agent environments with progressively complex task distributions for robust real-world adaptability.
-
Global Linguistic Coverage: Expanded support to 201 languages and dialects, enabling inclusive, worldwide deployment with nuanced cultural and regional understanding.
-
Next-Generation Training Infrastructure: Near-100% multimodal training efficiency compared to text-only training and asynchronous RL frameworks supporting massive-scale agent scaffolds and environment orchestration.
Benchmark Results
For more details, please refer to our blog post Qwen3.5.
Model Overview
- Type: Causal Language Model with Vision Encoder
- Training Stage: Pre-training & Post-training
- Language Model
- Number of Parameters: 4B
- Hidden Dimension: 2560
- Token Embedding: 248320 (Padded)
- Number of Layers: 32
- Hidden Layout: 8 × (3 × (Gated DeltaNet → FFN) → 1 × (Gated Attention → FFN))
- Gated DeltaNet:
- Number of Linear Attention Heads: 32 for V and 16 for QK
- Head Dimension: 128
- Gated Attention:
- Number of Attention Heads: 16 for Q and 4 for KV
- Head Dimension: 256
- Rotary Position Embedding Dimension: 64
- Feed Forward Network:
- Intermediate Dimension: 9216
- LM Output: 248320 (Tied to token embedding)
- MTP: trained with multi-steps
- Context Length: 262,144 natively and extensible up to 1,010,000 tokens.
Benchmark Results
Language
| GPT-OSS-120B | GPT-OSS-20B | Qwen3-Next-80B-A3B-Thinking | Qwen3-30BA3B-Thinking-2507 | Qwen3.5-9B | Qwen3.5-4B | |
|---|---|---|---|---|---|---|
| Knowledge & STEM | ||||||
| MMLU-Pro | 80.8 | 74.8 | 82.7 | 80.9 | 82.5 | 79.1 |
| MMLU-Redux | 91.0 | 87.8 | 92.5 | 91.4 | 91.1 | 88.8 |
| C-Eval | 76.2 | 71.4 | 89.7 | 87.4 | 88.2 | 85.1 |
| SuperGPQA | 54.6 | 48.5 | 60.8 | 56.8 | 58.2 | 52.9 |
| GPQA Diamond | 80.1 | 71.5 | 77.2 | 73.4 | 81.7 | 76.2 |
| Instruction Following | ||||||
| IFEval | 88.9 | 88.2 | 88.9 | 88.9 | 91.5 | 89.8 |
| IFBench | 69.0 | 65.1 | 61.5 | 51.5 | 64.5 | 59.2 |
| MultiChallenge | 45.3 | 40.1 | 51.3 | 46.5 | 54.5 | 49.0 |
| Long Context | ||||||
| AA-LCR | 50.7 | 30.7 | 51.7 | 49.0 | 63.0 | 57.0 |
| LongBench v2 | 48.2 | 45.6 | 48.0 | 44.8 | 55.2 | 50.0 |
| Reasoning & Coding | ||||||
| HMMT Feb 25 | 90.0 | 76.7 | 73.7 | 63.1 | 83.2 | 74.0 |
| HMMT Nov 25 | 90.0 | 81.8 | 81.2 | 73.8 | 82.9 | 76.8 |
| LiveCodeBench v6 | 82.7 | 74.6 | 68.7 | 66.0 | 65.6 | 55.8 |
| OJBench | 41.5 | 36.3 | 29.7 | 25.1 | 29.2 | 24.1 |
| General Agent | ||||||
| BFCL-V4 | -- | -- | 49.7 | 42.4 | 66.1 | 50.3 |
| TAU2-Bench | -- | -- | 57.4 | 41.9 | 79.1 | 79.9 |
| VITA-Bench | -- | -- | 29.5 | 14.1 | 29.8 | 22.0 |
| DeepPlanning | -- | -- | 0.4 | 4.9 | 18.0 | 17.6 |
| Multilingualism | ||||||
| MMMLU | 78.2 | 69.7 | 81.3 | 78.4 | 81.2 | 76.1 |
| MMLU-ProX | 74.5 | 67.3 | 73.6 | 69.1 | 76.3 | 71.5 |
| NOVA-63 | 51.1 | 48.7 | 53.3 | 52.5 | 55.9 | 54.3 |
| INCLUDE | 74.0 | 65.3 | 78.3 | 74.4 | 75.6 | 71.0 |
| Global PIQA | 84.1 | 79.8 | 83.5 | 80.2 | 83.2 | 78.9 |
| PolyMATH | 54.0 | 30.9 | 62.4 | 52.6 | 57.3 | 51.1 |
| WMT24++ | 74.4 | 67.8 | 57.4 | 69.3 | 72.6 | 66.6 |
| MAXIFE | 83.7 | 80.1 | 79.9 | 77.4 | 83.4 | 78.0 |
* TAU2-Bench: we follow the official setup except for the airline domain, where all models are evaluated by applying the fixes proposed in the Claude Opus 4.5 system card.
* MMLU-ProX: we report the averaged accuracy on 29 languages.
* WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL.
* MAXIFE: we report the accuracy on English + multilingual original prompts (totally 23 settings).
* Empty cells (--) indicate scores not yet available or not applicable.
Vision Language
| GPT-5-Nano-2025-08-07 | Gemini-2.5-Flash-Lite | Qwen3-VL-30B-A3B | Qwen3.5-9B | Qwen3.5-4B | |
|---|---|---|---|---|---|
| STEM and Puzzle | |||||
| MMMU | 75.8 | 73.4 | 76.0 | 78.4 | 77.6 |
| MMMU-Pro | 57.2 | 59.7 | 63.0 | 70.1 | 66.3 |
| MathVision | 62.2 | 52.1 | 65.7 | 78.9 | 74.6 |
| Mathvista(mini) | 71.5 | 72.8 | 81.9 | 85.7 | 85.1 |
| We-Math | 62.5 | 32.1 | 70.0 | 75.2 | 75.4 |
| DynaMath | 78.0 | 69.9 | 80.1 | 83.6 | 83.3 |
| ZEROBench | 1.0 | 1.0 | 0.0 | 3.0 | 3.0 |
| ZEROBench_sub | 22.2 | 19.2 | 23.7 | 31.1 | 26.3 |
| VlmsAreBlind | 66.7 | 68.4 | 72.5 | 93.7 | 92.6 |
| BabyVision | 14.4 | 17.5 | 18.6 | 28.6/25.8 | 16.0/19.1 |
| General VQA | |||||
| RealWorldQA | 71.8 | 72.2 | 77.4 | 80.3 | 79.5 |
| MMStar | 68.6 | 69.1 | 75.5 | 79.7 | 78.3 |
| MMBenchEN-DEV-v1.1 | 80.3 | 82.7 | 88.9 | 90.1 | 89.4 |
| SimpleVQA | 46.0 | 54.1 | 54.3 | 51.2 | 43.4 |
| HallusionBench | 58.4 | 64.5 | 66.0 | 69.3 | 65.0 |
| Text Recognition and Document Understanding | |||||
| OmniDocBench1.5 | 55.9 | 79.4 | 86.8 | 87.7 | 86.2 |
| CharXiv(RQ) | 50.1 | 56.1 | 56.6 | 73.0 | 70.8 |
| MMLongBench-Doc | 31.8 | 46.5 | 47.4 | 57.7 | 54.2 |
| CC-OCR | 58.9 | 72.9 | 77.8 | 79.3 | 76.7 |
| AI2D_TEST | 81.9 | 85.7 | 86.9 | 90.2 | 89.6 |
| OCRBench | 75.3 | ||||
14 个文件
浏览文件数据集版权信息
本数据集的许可证为 Apache License 2.0。如有违反相关条款,请联系 WEHUB,我们将及时处理。 查看许可证
通过 WeHub CLI 下载当前数据集快照。下列命令会固定为当前页面展示的数据版本(如果页面提供版本)。文件字节由本机直连存储下载,浏览器不会签发或保存下载链接。
前置要求
需要 Node.js 18 及以上,以及 npm(或 npx)。
1. 安装 CLI
npm install -g wehub-cli@latest
2. 下载此数据集
wehub datasets download ds_ext_3302_44693f7232 --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --output ./ds_ext_3302_44693f7232
若中断或部分失败,在同一目录重新执行同一命令即可续传。默认会校验 SHA-256。
免全局安装
npx --yes wehub-cli@latest datasets download ds_ext_3302_44693f7232 --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --output ./ds_ext_3302_44693f7232
高级选项
以下为 wehub datasets download 已支持的参数示例:
强制重新下载,不复用已校验的本地文件
wehub datasets download ds_ext_3302_44693f7232 --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --output ./ds_ext_3302_44693f7232 --overwrite
仅包含匹配路径
wehub datasets download ds_ext_3302_44693f7232 --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --output ./ds_ext_3302_44693f7232 --include "*.jsonl"
排除匹配路径
wehub datasets download ds_ext_3302_44693f7232 --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --output ./ds_ext_3302_44693f7232 --exclude "*.md"
提高并发下载数
wehub datasets download ds_ext_3302_44693f7232 --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --output ./ds_ext_3302_44693f7232 --jobs 8
完整帮助:wehub datasets download --help