数据集 / FunAudioLLM/Fun-CosyVoice3-0.5B-2512

FunAudioLLM/Fun-CosyVoice3-0.5B-2512 已完整同步

SVG Banners

👉🏻 CosyVoice 👈🏻

Fun-CosyVoice 3.0: Demos; Paper; Modelscope; Huggingface; CV3-Eval

CosyVoice 2.0: Demos; Paper; Modelscope; HuggingFace

CosyVoice 1.0: Demos; Paper; Modelscope; HuggingFace

Highlight🔥

Fun-CosyVoice 3.0 is an advanced text-to-speech (TTS) system based on large language models (LLM), surpassing its predecessor (CosyVoice 2.0) in content consistency, speaker similarity, and prosody naturalness. It is designed for zero-shot multilingual speech synthesis in the wild.

Key Features

  • Language Coverage: Covers 9 common languages (Chinese, English, Japanese, Korean, German, Spanish, French, Italian, Russian), 18+ Chinese dialects/accents (Guangdong, Minnan, Sichuan, Dongbei, Shan3xi, Shan1xi, Shanghai, Tianjin, Shandong, Ningxia, Gansu, etc.) and meanwhile supports both multi-lingual/cross-lingual zero-shot voice cloning.
  • Content Consistency & Naturalness: Achieves state-of-the-art performance in content consistency, speaker similarity, and prosody naturalness.
  • Pronunciation Inpainting: Supports pronunciation inpainting of Chinese Pinyin and English CMU phonemes, providing more controllability and thus suitable for production use.
  • Text Normalization: Supports reading of numbers, special symbols and various text formats without a traditional frontend module.
  • Bi-Streaming: Support both text-in streaming and audio-out streaming, and achieves latency as low as 150ms while maintaining high-quality audio output.
  • Instruct Support: Supports various instructions such as languages, dialects, emotions, speed, volume, etc.

Roadmap

  • 2025/12

    • release Fun-CosyVoice3-0.5B-2512 base model, rl model and its training/inference script
    • release Fun-CosyVoice3-0.5B modelscope gradio space
  • 2025/08

    • Thanks to the contribution from NVIDIA Yuekai Zhang, add triton trtllm runtime support and cosyvoice2 grpo training support
  • 2025/07

    • release Fun-CosyVoice 3.0 eval set
  • 2025/05

    • add CosyVoice2-0.5B vllm support
  • 2024/12

    • 25hz CosyVoice2-0.5B released
  • 2024/09

    • 25hz CosyVoice-300M base model
    • 25hz CosyVoice-300M voice conversion function
  • 2024/08

    • Repetition Aware Sampling(RAS) inference for llm stability
    • Streaming inference mode support, including kv cache and sdpa for rtf optimization
  • 2024/07

    • Flow matching training support
    • WeTextProcessing support when ttsfrd is not available
    • Fastapi server and client

Evaluation

Model Open-Source Model Size test-zh
CER (%) ↓
test-zh
Speaker Similarity (%) ↑
test-en
WER (%) ↓
test-en
Speaker Similarity (%) ↑
test-hard
CER (%) ↓
test-hard
Speaker Similarity (%) ↑
Human - - 1.26 75.5 2.14 73.4 - -
Seed-TTS - 1.12 79.6 2.25 76.2 7.59 77.6
MiniMax-Speech - 0.83 78.3 1.65 69.2 - -
F5-TTS 0.3B 1.52 74.1 2.00 64.7 8.67 71.3
Spark TTS

20 个文件

浏览文件