数据集 / inclusionAI/Ming-UniAudio-16B-A3B

inclusionAI/Ming-UniAudio-16B-A3B 已完整同步

Ming-UniAudio

📑 Technical Report📖Project Page🤗 Hugging Face🤖 ModelScope

Introduction

Ming-UniAudio is a novel framework that unifies speech understanding, generation, and editing. Its core is a unified continuous speech tokenizer that effectively unifies semantic and acoustic features within an end-to-end model. We developed a speech language model that strikes a balance between generation and understanding capabilities based on the unified continuous audio tokenizer. Leveraging this foundational model, which exhibits robust performance in both domains, we further trained a dedicated speech editing model built upon Ming-Lite-Omni. Crucially, Ming-UniAudio is the first to enable universal, free-form speech editing guided solely by natural language instructions, handling complex semantic and acoustic modifications without manual region specification.

  • 🔥 First unified continuous speech tokenizer for both understanding and generation tasks: MingTok-Audio
  • 🔥 First Speech LLM with unifed continuous tokenizer for both understanding and generation: Ming-UniAudio
  • 🔥 First universal free-form speech editing model for various semantic and acoustic editing task without timestamp condition: Ming-UniAudio-Edit
  • 🔥 First benchmark for free-form speech editing: Ming-Freeform-Audio-Edit-Benchmark

📌 Updates

  • [2025.09.30] 🔥 We release Ming-UniAudio with significant improvements across speech understanding, generation, and free-form editing tasks.

Key Features

Ming-UniAudio features key optimizations as follows, compared to other audio-assisted LLMs:

  • Unified Continuous Speech Tokenizer: Ming-UniAudio proposes a unified continuous speech tokenizer MingTok-Audio based on a VAE framework with a causal Transformer architecture, the first continuous speech tokenizer to effectively integrate semantic and acoustic features, and enables a closed-loop system with LLMs through hierarchical feature representations, makes it suitable for both understanding and generation tasks

  • Unified Speech Language Model for Generation and Understanding: We pretrain an end-to-end unified speech language model with a single LLM backbone for both understanding and generation tasks, enhanced with a Diffusion Head to ensure high-fidelity speech synthesis.

  • Instruction-Guided Free-Form Speech Editing: We introduce the first instruction-guided, free-form speech editing framework that supports comprehensive semantic and acoustic edits without requiring explicit edit regions, along with Ming-Freeform-Audio-Edit, the first open-source evaluation set for such tasks.

Evaluation

In various benchmark tests, Ming-UniAudio demonstrates highly competitive results compared to industry-leading models of similar scale.

Speech Understanding

ASR performance comparison on various audio benchmark datasets. The best results are in bold.
Datasets Mo

12 个文件

浏览文件