数据集 / Nanbeige/Nanbeige4.1-3B

Nanbeige/Nanbeige4.1-3B 已完整同步

Introduction

Nanbeige4.1-3B is built upon Nanbeige4-3B-Base and represents an enhanced iteration of our previous reasoning model, Nanbeige4-3B-Thinking-2511, achieved through further post-training optimization with supervised fine-tuning (SFT) and reinforcement learning (RL). As a highly competitive open-source model at a small parameter scale, Nanbeige4.1-3B illustrates that compact models can simultaneously achieve robust reasoning, preference alignment, and effective agentic behaviors.

Specifically, Nanbeige4.1-3B exhibits the following key strengths:

  • Strong Reasoning: Nanbeige4.1-3B is capable of solving complex, multi-step problems through sustained and coherent reasoning within a single forward pass, and reliably produces correct final answers on challenging tasks such as LiveCodeBench-Pro, IMO-Answer-Bench, and AIME 2026 I.
  • Robust Preference Alignment: Nanbeige4.1-3B achieves solid alignment performance, outperforming not only same-scale models such as Qwen3-4B-2507 and Nanbeige4-3B-2511, but also substantially larger models including Qwen3-30B-A3B and Qwen3-32B on Arena-Hard-v2 and Multi-Challenge.
  • Agentic Capability: Nanbeige4.1-3B is the first general small model to natively support deep-search tasks and reliably sustain complex problem solving involving more than 500 rounds of tool invocations. It fills a long-standing gap in the small-model ecosystem where models are typically optimized for either general reasoning or agentic scenarios, but rarely excel at both.

Technical Report: Link

Performances

We evaluate Nanbeige4.1-3B across a broad and diverse set of benchmarks covering general reasoning, and deep-search capabilities.

General Reasoning Tasks

On general reasoning tasks including code, math, science, alignment, and tool-use benchmarks, Nanbeige4.1-3B not only significantly outperforms same-scale models such as Qwen3-4B, but also demonstrates overall superior performance compared to larger models including Qwen3-30B-A3B-2507 and Qwen3-32B.

Benchmark Qwen3-4B-2507 Qwen3-8B Qwen3-14B Qwen3-32B Qwen3-30B-A3B-2507 Nanbeige4-3B-2511 Nanbeige4.1-3B
Code
Live-Code-Bench-V6 57.4 49.4 55.9 55.7 66.0 46.0 76.9
Live-Code-Bench-Pro-Easy 40.2 41.2 33.0 42.3 60.8 40.2 81.4
Live-Code-Bench-Pro-Medium 5.3 3.5 1.8 3.5 3.5 5.3 28.1
Math
AIME 2026 I 81.46 70.42 76.46 75.83 87.30 84.1 87.40
HMMT Nov 68.33 48.33 56.67 57.08 71.25 66.67 77.92
IMO-Answer-Bench 48.00 36.56 41.81 43.94 54.34 38.25 53.38
Science
GPQA

16 个文件

浏览文件