数据集 / openbmb/VoxCPM1.5

openbmb/VoxCPM1.5 已完整同步

🎙️ VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning

Project Page Technical ReportLive Playground Samples

  • VoxCPM1.5

Hugging Face ModelScope

VoxCPM Logo

🎉 VoxCPM1.5 Updates

Release Date: December 5, 2025

VoxCPM1.5 brings improvements in audio quality and efficiency:

Feature VoxCPM VoxCPM1.5
Audio VAE Sampling Rate 16kHz 44.1kHz
LM Token Rate 12.5Hz 6.25Hz
Patch Size 2 4
SFT Support
LoRA Support

Key Improvements:

  • 🔊 Higher Quality: 44.1kHz sampling rate preserves more high-frequency details for better voice cloning
  • More Efficient: Reduced token rate (6.25Hz) lowers computational cost while maintaining performance
  • 🎓 Fine-tuning Support: Train personalized voice models with SFT or LoRA

Note: Output quality depends on the prompt speech quality. VoxCPM-0.5B remains fully supported with backward compatibility.

📚 Model Overview

VoxCPM is a novel tokenizer-free Text-to-Speech (TTS) system that redefines realism in speech synthesis. By modeling speech in a continuous space, it overcomes the limitations of discrete tokenization and enables two flagship capabilities: context-aware speech generation and true-to-life zero-shot voice cloning.

Unlike mainstream approaches that convert speech to discrete tokens, VoxCPM uses an end-to-end diffusion autoregressive architecture that directly generates continuous speech representations from text. Built on MiniCPM-4 backbone, it achieves implicit semantic-acoustic decoupling through hierachical language modeling and FSQ constraints, greatly enhancing both expressiveness and generation stability.

VoxCPM Model Architecture

🚀 Key Features

  • Context-Aware, Expressive Speech Generation - VoxCPM comprehends text to infer and generate appropriate prosody, delivering speech with remarkable expressiveness and natural flow. It spontaneously adapts speaking style based on content, producing highly fitting vocal expression trained on a massive 1.8 million-hour bilingual corpus.
  • True-to-Life Voice Cloning - With only a short reference audio clip, VoxCPM performs accurate zero-shot voice cloning, capturing not only the speaker’s timbre but also fine-grained characteristics such as accent, emotional tone, rhythm, and pacing to create a faithful and natural replica.
  • High-Efficiency Synthesis - VoxCPM supports streaming synthesis with a Real-Time Factor (RTF) as low as 0.17 on a consumer-grade NVIDIA RTX 4090 GPU, making it possible for real-time applications.

Quick Start

🔧 Install from PyPI

pip install voxcpm

1. Model Download (Optional)

By default, when you first run the script, the model will be downloaded automatically, but you can also download the model in advance.

  • Download VoxCPM1.5
    from huggingface_hub import snapshot_download
    snapshot
    

10 个文件

浏览文件