VibeVoice: Expressive, Longform Conversational Speech Synthesis
VibeVoice is a cutting-edge framework designed for generating expressive, long-form, multi-speaker conversational audio from text. This community-maintained fork preserves the original codebase and introduces additional functionalities, including unofficial training and fine-tuning implementations.
Key Features:
- Long-form Synthesis: Generate speech up to 90 minutes long with up to 4 distinct speakers.
- Continuous Speech Tokenizers: Utilizes Acoustic and Semantic tokenizers for high audio fidelity and computational efficiency.
- Next-token Diffusion Framework: Leverages a Large Language Model (LLM) for understanding textual context and dialogue flow.
- Fine-tuning Support: Adapt VibeVoice to new languages or voices, enhancing versatility.
Benefits:
- Scalability: Overcomes limitations of traditional Text-to-Speech systems.
- Natural Turn-taking: Improves conversational flow and speaker consistency.
- Community-driven: Maintained by a community of contributors, ensuring ongoing development and support.
Highlights:
- Open Source: Licensed under MIT, promoting accessibility and collaboration.
- Active Community: Join the unofficial Discord community for support and sharing experiences.
- Innovative Technology: Addresses challenges in TTS systems, making it a powerful tool for content creators and developers.




