数据集 / Fineweb-Edu-Chinese-V2.2

Fineweb-Edu-Chinese-V2.2 已完整同步

Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train)

Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs

Chinese Fineweb Edu Dataset V2.2 is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain.

This project aims to solve the core pain point of "scarcity of high-quality educational corpora" in the Chinese open-source community. Building on the massive pre-training data of V2.1, the V2.2 version leverages the powerful text understanding capabilities of DeepSeek V3.2 to distill 1.43 million high-quality Q&A pairs from the top 0.1% of high-quality corpora, providing a standardized Post-training dataset for the community.


Why Do We Need This Dataset?

In current LLM research and development, the "scarcity of high-quality post-training data" has become the biggest bottleneck restricting the leap in model intelligence.

1. The Trap of "Models Taking Shortcuts"

Current open-source SFT data (such as early Alpaca, ShareGPT) allows models to learn dialogue formats but often sacrifices factual accuracy.

Core Arguments & Evidence:

  • LIMA Hypothesis (Less Is More for Alignment): Meta AI research shows that the primary role of fine-tuning is "format alignment" rather than "learning new knowledge." Just 1,000 carefully selected high-quality samples can outperform 50,000 ordinary samples. This proves that data purity is far more important than quantity.

  • The False Promise of Imitation Learning: Research from UC Berkeley points out that training models with large amounts of low-quality SFT data only allows them to "imitate the style of proprietary models" without acquiring their logical reasoning capabilities. This results in models that are "giants in style, but dwarfs in fact."

2. The Quality Crisis of Synthetic Data

As more data is generated by AI, models will degrade if strict quality control is lacking.

Core Arguments & Evidence:

  • Model Collapse: Rice University research found that if models are trained recursively on low-quality synthetic data, "Model Collapse" occurs, losing tail information of the distribution and leading to a loss of creativity and diversity. The only way to avoid collapse is to use highly pure, textbook-quality synthetic sources.

  • Lessons from AlpaGasus: Researchers filtered out 90% of low-quality Alpaca data and trained a model with only 9,000 samples, which outperformed the model trained on the full dataset in various metrics.

Strategy of V2.2

Addressing the above industry pain points, V2.2 insists on Quality Over Quantity:

  1. **Reject Low

3 个文件

浏览文件