Text-to-Speech (TTS)
Models that convert text into natural-sounding spoken audio — the counterpart to speech-to-text transcription — used for voice agents, narration, and voice cloning.
Modern TTS models generate speech with natural prosody and, in many vendors' products, can clone a specific voice from a short reference sample. Buyers in this category typically care about latency (critical for real-time voice-agent use cases, less so for pre-recorded narration), language and accent coverage, and voice-cloning consent/licensing terms given the potential for misuse. Speech-to-text (transcription, the reverse direction) is often bundled by the same vendors, and full voice-agent products layer both directions plus an LLM on top to handle a live conversation rather than one-way conversion.
Last verified: