LogoAISecKit
  • Search
  • Collection
  • Category
  • Tag
  • Blog
  • Pricing
  • Submit
LogoAISecKit

Newsletter

Join the Community

Subscribe to our newsletter for the latest news and updates

LogoAISecKit

Curated directory of 1700+ AI tools, models, frameworks, MCP servers, and cybersecurity resources

GitHub
Product
  • Search
  • Collection
  • Category
  • Tag
Resources
  • Blog
  • Pricing
  • Submit
Company
  • About Us
  • Privacy Policy
  • Terms of Service
  • Sitemap
Copyright © 2026 All Rights Reserved.
Sponsored Resources
  1. Home
  2. Category
  3. VibeVoice
icon of VibeVoice

VibeVoice

VibeVoice is a community-maintained fork for expressive, longform conversational speech synthesis.

Visit Website
image for VibeVoice
Visit Website

Introduction

VibeVoice: Expressive, Longform Conversational Speech Synthesis

VibeVoice is a cutting-edge framework designed for generating expressive, long-form, multi-speaker conversational audio from text. This community-maintained fork preserves the original codebase and introduces additional functionalities, including unofficial training and fine-tuning implementations.

Key Features:
  • Long-form Synthesis: Generate speech up to 90 minutes long with up to 4 distinct speakers.
  • Continuous Speech Tokenizers: Utilizes Acoustic and Semantic tokenizers for high audio fidelity and computational efficiency.
  • Next-token Diffusion Framework: Leverages a Large Language Model (LLM) for understanding textual context and dialogue flow.
  • Fine-tuning Support: Adapt VibeVoice to new languages or voices, enhancing versatility.
Benefits:
  • Scalability: Overcomes limitations of traditional Text-to-Speech systems.
  • Natural Turn-taking: Improves conversational flow and speaker consistency.
  • Community-driven: Maintained by a community of contributors, ensuring ongoing development and support.
Highlights:
  • Open Source: Licensed under MIT, promoting accessibility and collaboration.
  • Active Community: Join the unofficial Discord community for support and sharing experiences.
  • Innovative Technology: Addresses challenges in TTS systems, making it a powerful tool for content creators and developers.
Back

Information

  • Publisher
    AISecKit
  • Websitegithub.com
  • Published date2025/10/18

Categories

  • AI Models
  • AI Application Platforms
  • AI Audio Tools

Tags

  • Fine-tuning
  • Voice Assistants
  • Text-to-Audio

More Products

image of Nano Bananary
AI ModelsAI Application PlatformsAI Video Tools
Visit Website
icon of Nano Bananary

Nano Bananary

Nano Bananary is an AI batch image and video generator with 142 effects.

Text-to-VideoGenerative AI
image of Twocast
AI Application PlatformsAI Productivity ToolsAI Audio Tools
Visit Website
icon of Twocast

Twocast

AI Podcast Generator for bilingual episodes, supporting multiple languages and alternative to NotebookLLM.

Content Creation
image of ZCF
AI Application PlatformsAI Productivity ToolsAI Development Frameworks
Visit Website
icon of ZCF

ZCF

Zero-Config Code Flow for Claude code & Codex, enabling seamless integration and configuration for AI development.

Open SourceClaude