cjwbw/voicecraft
Performs zero-shot speech editing and text-to-speech synthesis using audio input and text transcripts. Supports four mai...
Found 145 models (showing 41-60)
Performs zero-shot speech editing and text-to-speech synthesis using audio input and text transcripts. Supports four mai...
Synthesize speech from text in a cloned voice using a reference audio sample. Provide a text prompt and speaker referenc...
Generate expressive multilingual speech from text. Accept a text prompt and a language selection (ar, da, de, el, en, es...
Generate English speech from text with zero-shot voice cloning from a 5–10s reference clip. Provide text, a short refere...
Generate speech from text with optional voice cloning from a reference voice sample. Accepts a text prompt and optional...
Convert text to expressive speech, with optional speaker style cloning from a short reference audio. Accepts text input...
Generate Spanish speech from text by cloning the voice from a reference audio. Provide Spanish text, a reference audio s...
Generate speech audio from text across 25 languages. Accepts text and a language selection; returns spoken audio. Suppor...
Synthesizes speech from text using a reference audio file and its transcript to clone the speaker's voice. Takes text in...
Generate long-form, multi-speaker conversational speech from a text script. Accepts a script and up to four selected voi...
Generate speech from text and convert voices. Use zero_shot voice cloning to synthesize speech in the style of a prompt_...
Clone voices from audio samples with ultra-low latency streaming synthesis. Supports zero-shot voice cloning, cross-ling...
Generate multilingual speech from text with zero-shot voice cloning. Provide a short reference audio clip and its transc...
Converts text to speech using a speaker reference audio file to clone the voice characteristics and speaking style of th...
Generate natural, conversational speech and two-speaker dialogues from text. Choose from preset voices (Angelo, Arsenio,...
Converts text to speech using the Kokoro model with 82 million parameters, based on StyleTTS2 architecture. Supports mul...
Generate spoken audio from text. Clone a target voice by providing a prompt audio sample (voice_cloning mode), or synthe...
Generates speech from text input using a reference speaker audio and corresponding text. Takes text to synthesize, a ref...
Clone a voice from a short reference audio and synthesize speech from text. Provide at least 6 seconds of speaker audio...
Convert Persian (Farsi) text to speech. Accepts a Persian text string and returns spoken audio. Runs fast on low-resourc...