1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
Voice cloning and fine-tuning from a minute of audio, with a web UI.
- 言語
- Python
- ライセンス
- MIT
- スター
- 62k
- 最新リリース
- 20250606v2pro
- 最終プッシュ
- 2026-10-06
Hosted text-to-speech and voice cloning service, with an API. ElevenLabsのクローズドな製品です。
| ツール | 置き換えの度合い | スター | ライセンス | 利用条件 | セルフホスト | 言語 | 最新リリース | 最終プッシュ |
|---|---|---|---|---|---|---|---|---|
| GPT-SoVITS | 部分的 | 62k | MIT | オープンソース | 可 | Python | 20250606v2pro | 2026-10-06 |
| Fish Speech | 部分的 | 33k | Other | ソースアベイラブル | 可 | Python | v1.5.1 | 2026-10-05 |
| Chatterbox | 部分的 | 27k | MIT | オープンソース | 可 | Python | v0.1.2 | 2026-07-21 |
| IndexTTS | 部分的 | 24k | Other | ソースアベイラブル | 可 | Python | v2.5.0 | 2026-09-29 |
| CosyVoice | 部分的 | 24k | Apache-2.0 | オープンソース | 可 | Python | v2.0 | 2026-05-25 |
| ebook2audiobook | 部分的 | 20k | Apache-2.0 | オープンソース | 可 | Python | v26.10.1 | 2026-10-06 |
| F5-TTS | 部分的 | 15k | MIT | オープンソース | 可 | Python | 1.1.22 | 2026-09-21 |
| Piper | 部分的 | 5.8k | GPL-3.0 | オープンソース | 可 | C++ | v1.8.0 | 2026-09-28 |
| Kokoro-FastAPI | 部分的 | 5.5k | Apache-2.0 | オープンソース | 可 | Python | v0.9.0 | 2026-10-05 |
代替9件
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
Voice cloning and fine-tuning from a minute of audio, with a web UI.
SOTA Open Source TTS
Multilingual TTS with voice cloning, under a non-commercial licence.
SoTA open-source TTS
An open TTS model with zero-shot voice cloning and emotion control.
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Zero-shot TTS with voice cloning and control over duration and emotion.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Multilingual TTS with zero-shot voice cloning and streaming output.
Generate audiobooks from e-books, voice cloning & 1158+ languages!
Converts ebooks into audiobooks with local TTS models, including cloned voices.
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
Voice cloning from a short reference clip, with a Gradio app.
Fast and local neural text-to-speech engine
Fast offline voices that run on a Raspberry Pi, without voice cloning.
Dockerized OpenAI-compatible wrapper for Kokoro-82M text-to-speech w/multiplatform CPU, AMD, NVIDIA GPU PyTorch; multi-speaker, clone-tuning, caption timestamps, SSML, optional readalong web UI
An OpenAI-compatible speech API around the Kokoro model, on CPU or GPU, without voice cloning.