minimax/speech-02-hd
Generate speech audio from text with multilingual voices, emotion control, and voice cloning. Accepts text (up to 10,000...
Found 145 models (showing 21-40)
Generate speech audio from text with multilingual voices, emotion control, and voice cloning. Accepts text (up to 10,000...
Translate speech and text across 100+ languages, returning text and optionally synthesized speech. Accept audio or text...
Processes audio input to perform multiple audio tasks including speech-to-text transcription, audio captioning, emotion...
Generate expressive, natural speech from text prompts with unique emotion control and instant voice cloning capabilities...
Generate speech from text using a reference speaker audio to clone the speakerβs voice. Accepts a text prompt and a spea...
Converts text into spoken audio using a speaker reference audio file to clone the voice characteristics and speaking sty...
Generate speech from text with a cloned voice. Provide a text prompt, a target language (en, zh, es, ja, ko, fr), and a...
Convert text to speech with zero-shot voice cloning from a reference audio sample. Accepts text and a voice sample and o...
Generate conversational speech from text input. Convert text to spoken audio with Sesameβs CSM 1B; choose between two sp...
Clones voices and generates speech in multiple languages from text input and a reference audio sample. Takes text and an...
Generate speech from text while cloning a target voice from a reference audio sample. Input text and a speaker reference...
Generate conversational speech for phone-call applications from text input. Select from multiple preset voices (male_voi...
Generate speech audio from text input with low latency. Select from 700+ multilingual voices and accents, or use a voice...
Generate expressive speech from text. Accepts a text prompt and returns spoken audio with controllable emotion and proso...
Convert text to speech with optional zero-shot voice cloning from a short reference audio clip. Accepts text and an opti...
Generates speech and audio from text prompts using a transformer-based model. Supports multilingual speech generation wi...
Generate speech, music, background noise, and simple sound effects from a text prompt. Output an audio file, with an opt...
Clone a voice from a short reference clip and generate speech from text. Accepts text and a reference audio sample; outp...
Generate speech audio from text with selectable multilingual voices. Accepts text input, a preset voice, and a speed mul...
Generate expressive speech from text. Accepts text input and outputs spoken audio, with preset voices (tara, dan, josh,...