数据集 / meituan-longcat/LongCat-Video-Avatar

meituan-longcat/LongCat-Video-Avatar 已完整同步

LongCat-Video-Avatar


🚀 Model Introduction

We are excited to announce the release of LongCat-Video-Avatar, a unified model that delivers expressive and highly dynamic audio-driven character animation, supporting native tasks including Audio-Text-to-Video, Audio-Text-Image-to-Video, and Video Continuation with seamless compatibility for both single-stream and multi-stream audio inputs.

Key Features

  • 🌟 Support Multiple Generation Modes: One unified model can be used for audio-text-to-video (AT2V) generation, audio-text-image-to-video (ATI2V) generation, and Video Continuation.
  • 🌟 Natural Human Dynamics: The disentangled unconditional guidance is designed to effectively decouple speech signals from motion dynamics for natural behavior.
  • 🌟 Avoid Repetitive Content: The reference skip attention is adopted to​ strategically incorporates reference cues to preserve identity while preventing excessive conditional image leakage.
  • 🌟 Alleviate Error Accumulation from VAE: Cross-Chunk Latent Stitching is designed to eliminates redundant VAE decode-encode cycles to reduce pixel degradation in long sequences.

For more detail, please refer to the comprehensive LongCat-Video-Avatar Technical Report.

🌀 Preview Gallery

The following videos showcase example generations from our model and have been compressed for easier viewing.

</td>
<td>
  
  
</td>
</td>
<td>
  
  
</td>
</td>
<td>
  <video src="https://github.com/user-attachments/assets/03cca3e0-86ed-4a0a-a14f-663

34 个文件

浏览文件