Good Audio Generation space, model, dataset
Good Audio Generation space, model, dataset collection
-
Audio-to-Audio • 0.6B • Updated • 85.5k • 118 -
KittenML/kitten-tts-nano-0.1
Updated • 7.05k • 516 -
FunAudioLLM/ThinkSound
Video-to-Video • Updated • 54 -
ThinkSound
🔊321Generate audio for a silent video using text prompts
-
Higgs Audio Demo
🎤405Higgs Audio Demo
-
bosonai/higgs-tts-2-3b-base
Text-to-Speech • 6B • Updated • 525k • 693 -
Song Prep
🎵751Transcribe song lyrics with timestamps
-
Hibiki Samples
🤗54Translate speech in real-time with high fidelity
-
kyutai/moshiko-pytorch-bf16
8B • Updated • 151k • 255 -
kyutai/mimi
Feature Extraction • 96.2M • Updated • 638k • • 323 -
maya-research/Veena
Text-to-Speech • 4B • Updated • 14.1k • 239 -
MiniMax Speech Tech Report
🎙109Generate natural speech in any voice from text
-
google/magenta-realtime
Updated • 235 • 556 -
PlayDiffusion
🎨119Generate modified audio from text and voice
-
Qwen2.5 Omni 7B Demo
🏆375Chat with text, audio, images, and video, get spoken replies
-
Open ASR Leaderboard
🏆1.44kExplore and compare speech‑recognition models by WER and speed
-
Open NotebookLM
🎙143Generate a podcast to discuss the topic of your choice!
-
Voila Demo
💻44Chat with a voice-clone AI
-
Voice Clone
🗣2.67kClone a voice and generate speech from text
-
moonshotai/Kimi-Audio-7B-Instruct
Text-to-Speech • 10B • Updated • 27.7k • 417 -
moonshotai/Kimi-Audio-7B
Text-to-Speech • 10B • Updated • 213 • 96 -
Dia 1.6B
👯1.79kGenerate realistic dialogue from a script, using Dia!
-
nari-labs/Dia-1.6B
Text-to-Speech • 2B • Updated • 31.2k • • 2.91k -
ByteDance/MegaTTS3
Text-to-Speech • Updated • 518 • 418 -
Di♪♪Rhythm
🎶689Blazingly Fast and Embarrassingly Simple Song Generation
-
Gemini Audio Video
♊35Gemini understands audio and video!
-
nvidia/diar_sortformer_4spk-v1
Automatic Speech Recognition • 0.1B • Updated • 174k • 151 -
ACE Step
😻678A Step Towards Music Generation Foundation Model
-
ACE-Step/ACE-Step-v1-3.5B
Text-to-Audio • Updated • 743 -
stepfun-ai/Step-Audio-2-mini
Any-to-Any • 8B • Updated • 2.92k • 264 -
neuphonic/neutts-air
Text-to-Speech • 0.7B • Updated • 5.3k • 885 -
NeuTTS-Air
☁321Clone a voice and generate custom speech
-
KaniTTS
😻115Generate expressive speech from your text in seconds
-
microsoft/UserLM-8b
Text Generation • 8B • Updated • 2.6k • 388 -
pipecat-ai/smart-turn-v3
Voice Activity Detection • Updated • 197 -
meituan-longcat/LongCat-Audio-Codec
Updated • 43 -
Qwen3 TTS Voice Design
📈115Generate custom speech from text and voice description
-
Qwen TTS Clone Demo
👀64Create a custom voice and synthesize speech from text
-
ResembleAI/chatterbox-turbo
Text-to-Speech • Updated • • 681 -
Chatterbox Turbo Demo
⚡525Chatterbox Turbo Demo
-
zai-org/GLM-TTS
Text-to-Speech • Updated • 294 • 351 -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.54M • 1.93k -
Qwen3-TTS Demo
🎙2.19kGenerate speech from text with voice design, cloning, or presets
-
Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice
Text-to-Speech • 0.9B • Updated • 1.19M • 183 -
FlashLabs/Chroma-4B
Any-to-Any • 6B • Updated • 164 • 383 -
FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning
Paper • 2601.11141 • Published • 23 -
MOSS Transcribe Diarize: Accurate Transcription with Speaker Diarization
Paper • 2601.01554 • Published • 66 -
FunAudioLLM/Fun-Audio-Chat-8B
Any-to-Any • 9B • Updated • 471 • 188 -
OpenMOSS-Team/MOSS-TTS-Nano-100M
Text-to-Speech • Updated • 33.8k • 237 -
KittenTTS Demo
😻93Generate natural‑sounding speech from typed text
-
HF Realtime Voice
🎙546Voice chat over WebSocket against a HF speech-to-speech