◖
Models by output type

AI Audio Output: What Models Generate and When You Need It

AI audio output is the ability of a model or model endpoint to produce sound that an application can play, save, stream, or process. The most common examples are text-to-speech systems and realtime voice models, but the category can also include other forms of generated audio depending on how a provider defines its capabilities. This distinction matters because a model that understands audio is not necessarily able to generate it. A text model may also appear to “speak” because a separate text-to-speech service is used by the surrounding application. This guide explains what audio-output models produce, how they differ from transcription and audio understanding, what they are useful for, and which technical and practical criteria matter when comparing them.
What this means

Audio output is the broader category for generated sound. Speech and music are shown separately where the model has a more specific documented capability.

Audio models

104 models currently match this capability.

View all models →
◎
Amazon

Amazon Nova 2 Sonic

Amazon Nova 2

Real-time voice assistants, customer-service automation, telephony, interactive learning, multilingual conversations, and tool-enabled speech agents.

Audio Multimodal 1,000,000 ctx Audio input Tool use Structured output
View model →
◎
Amazon

Amazon Nova Sonic

Amazon Nova

Real-time voice assistants, customer-service automation, interactive education, language learning, and speech-enabled enterprise workflows

Audio Multimodal 300,000 ctx Audio input Tool use Streaming
View model →
◎
Tencent

AuK

AuK

Open-source text-to-speech, reference-voice generation, speech and lyric editing, emotion and timbre transformation, speech enhancement, and source separation

Audio Other Audio input
View model →
◎
Baidu

Baidu MuseSteamer 2.0

MuseSteamer 2.0

Chinese image-to-video generation, audiovisual storytelling, marketing videos, multi-person dialogue, synchronized speech, sound effects, and cinematic short-form content

Audio Multimodal Image input
View model →
◎
OpenAI

ChatGPT-4o

GPT-4o

Fast general-purpose conversations, vision, voice interactions, coding, and everyday productivity

Audio Multimodal 128,000 ctx Image input Audio input Video input
View model →
◎
Z.ai

CogVideoX-3

CogVideoX

Short-form text-to-video, image animation, start-and-end-frame transitions, advertising, marketing, realistic scenes, and 3D-style video generation

Audio Video Generation Image input
View model →
◎
NVIDIA

Cosmos3-Nano

Cosmos 3

Physical AI, robotics, autonomous-vehicle simulation, multimodal world generation, future-state prediction, action reasoning, and synthetic training data

Audio Multimodal Image input Audio input Video input
View model →
◎
NVIDIA

Cosmos3-Super

Cosmos 3

High-quality Physical AI simulation, synthetic-data generation, robotics and autonomous-vehicle research, multimodal world modeling, and teacher-model distillation

Audio Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 2.5 Flash Live

Gemini 2.5

Real-time voice and video agents, speech-to-speech assistants, interactive customer support, tutoring, coaching, and multimodal Live API applications

Audio Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 2.5 Flash TTS

Gemini 2.5

Low-latency controllable text-to-speech, voice assistants, narration, read-aloud features, and multi-speaker audio generation

Audio Other 8,192 ctx
View model →
◎
Google DeepMind

Gemini 2.5 Pro TTS

Gemini 2.5

High-fidelity single-speaker and multi-speaker narration, audiobooks, podcasts, professional voiceovers, and scripted creative audio

Audio Other 8,192 ctx
View model →
◎
Google DeepMind

Gemini 3.1 Flash Live Preview

Gemini 3.1

Low-latency voice agents, real-time dialogue, multimodal live sessions, and interactive audio applications

Audio Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.1 Flash TTS

Gemini 3.1 Flash Audio

Controllable expressive speech, narration, accessibility, scripted audio, and multi-speaker TTS prototypes

Audio Other 8,192 ctx Streaming
View model →
◎
Google DeepMind

Gemini 3.5 Live Translate

Gemini 3.5 Audio

Low-latency, real-time speech-to-speech translation for calls, meetings, travel, customer support, and multilingual voice applications

Audio Multimodal 131,072 ctx Audio input Streaming
View model →
◎
Google DeepMind

Gemini 3.8 Flash TTS

Gemini 3.8

Studio-quality narration, audiobooks, expressive voice acting, complex multi-speaker dialogue, regional accents, difficult pronunciations, long-form narration, voice design, and voice replication.

Audio Other 8,192 ctx Streaming
View model →
◎
Google DeepMind

Gemini 3.8 Flash-Lite TTS

Gemini 3.8

High-volume text-to-speech production, low-latency voice-agent cascades, read-aloud applications, voice replication, and everyday single-speaker speech

Audio Other 8,192 ctx Streaming
View model →
◎
Google DeepMind

Gemini 3.8 Live

Gemini 3.8

Low-latency voice agents, real-time audio-to-audio dialogue, multimodal assistants, and interactive tool-using applications

Audio Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.8 Live Extended Thinking

Gemini 3.8 Audio

Complex real-time voice agents, multi-step problem solving, asynchronous tool workflows, technical support, travel coordination, and spoken STEM or coding tutoring

Audio Reasoning 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Omni Flash

Gemini Omni

Fast text-to-video, image-to-video, conversational video editing, video extension, interpolation, marketing content, and short-form cinematic production

Audio Multimodal 1,048,576 ctx Image input Video input Streaming
View model →
◎
OpenAI

GPT-4o Audio

GPT-4o

Voice assistants, spoken conversational agents, audio-enabled customer service, and applications requiring direct audio understanding and speech generation

Audio Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-4o Mini Audio

GPT-4o

Lower-cost audio understanding, conversational voice interfaces, and applications requiring text and spoken-audio input/output

Audio Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-4o Mini Realtime

GPT-4o

Low-cost realtime voice assistants, speech-to-speech interfaces, interactive audio applications, and conversational prototypes

Audio Lightweight 16,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-4o Mini TTS

GPT-4o Mini

Fast, controllable text-to-speech for narration, voice interfaces, customer service, accessibility, and realtime audio applications.

Audio Other 2,000 ctx Streaming
View model →
◎
OpenAI

GPT-4o Realtime

GPT-4o

Low-latency voice assistants, speech-to-speech applications, live translation, language learning, and interactive customer support

Audio Multimodal 32,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-Audio

GPT-Audio

Audio-enabled chat applications, voice interfaces, spoken assistants, and applications requiring direct audio understanding and generation through Chat Completions.

Audio Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Audio Mini

GPT-Audio

Cost-sensitive, turn-based audio conversations, voice assistants, and audio-enabled applications using function calling

Audio Multimodal 128,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-Audio-1.5

GPT-Audio

Audio-in, audio-out conversational applications using the Chat Completions API, including voice assistants and tool-enabled spoken interfaces.

Audio Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Live 1

GPT-Live

Natural low-latency voice agents, customer support, conversational workflows, live assistance, and applications requiring interruption-aware speech interaction

Audio Multimodal Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Realtime

GPT-Realtime

Low-latency speech-to-speech voice agents, realtime customer support, education, accessibility, and conversational applications with function calling

Audio Realtime 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime Mini

GPT-Realtime

Cost-sensitive realtime voice agents, speech-to-speech applications, interactive assistants, and multimodal interfaces

Audio Realtime 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-1.5

GPT-Realtime

Low-latency speech-to-speech voice agents, customer support, realtime assistants, and audio applications that need function calling.

Audio Realtime Audio 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2

GPT-Realtime

Reasoning voice agents, speech-to-speech applications, customer support, live assistants, tool-driven workflows, and long conversational sessions

Audio Multimodal 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2.1

GPT-Realtime

Low-latency speech-to-speech agents, customer-service voice workflows, realtime tool use, telephony, and multimodal assistants with image input

Audio Realtime 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2.1 Mini

GPT-Realtime-2.1

Lower-cost, low-latency realtime voice agents, speech-to-speech assistants, and tool-enabled conversational applications

Audio Lightweight 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-Translate

GPT-Realtime

Low-latency spoken translation, multilingual calls, live interpretation, broadcasts, meetings, lessons, video rooms, captions, and translated audio experiences.

Audio Other 16,000 ctx Audio input Streaming
View model →
◎
xAI

Grok Imagine Video 1.5

Grok Imagine Video

Short-form text-to-video, image-to-video, reference-guided video, cinematic prototyping, marketing clips, and audiovisual creative workflows

Audio Other Image input Audio input
View model →
◎
xAI

Grok Voice Think Fast 2.0

Grok Voice Think Fast

Realtime voice agents, customer support, telephony, sales, multilingual conversations, and tool-enabled spoken workflows

Audio Multimodal Audio input Tool use Web search
View model →
◎
fal

H3 Max

MiniMax H3

Fast short-form text-to-video, image-to-video, reference-guided generation, and synchronized-audio production

Audio Multimodal Image input Audio input Video input
View model →
◎
NAVER

HyperCLOVA X SEED 8B Omni

HyperCLOVA X SEED

Korean-first any-to-any multimodal assistants, speech and vision applications, multimodal research, and self-hosted deployments

Audio Multimodal 32,768 ctx Image input Audio input Video input
View model →
◎
Moonshot AI

Kimi-Audio-7B

Kimi-Audio

Fine-tuning and research on speech recognition, audio understanding, audio classification, audio question answering, and speech-audio generation

Audio Multimodal 8,192 ctx Audio input Streaming
View model →
◎
Moonshot AI

Kimi-Audio-7B-Instruct

Kimi-Audio

Self-hosted speech recognition, audio understanding, audio question answering, audio captioning, and spoken conversational agents

Audio Multimodal Audio input Streaming
View model →
◎
Google DeepMind

Lyria 3 Clip

Lyria 3

Generating short music or audio clips

Audio Other
View model →
◎
Google DeepMind

Lyria 3 Pro

Lyria 3

Full-length AI music, soundtrack creation, songwriting, structured compositions, advertising audio, games, and creative production workflows

Audio Other Image input
View model →
◎
Google DeepMind

Lyria 3.5

Lyria

Full-length AI-generated songs, vocal music, instrumental arrangements, songwriting experiments, soundtracks, and image-inspired music creation

Audio Other 131,072 ctx Image input
View model →
◎
Google DeepMind

Lyria RealTime

Lyria

Interactive instrumental music generation, live musical improvisation, prompt-driven DJ tools, MIDI-controlled experiences, and real-time creative audio applications

Audio Other Streaming
View model →
◎
NVIDIA

Magpie TTS Multilingual

Magpie TTS

Multilingual voice agents, accessibility, narration, audiobooks, dubbing, localization and interactive speech applications

Audio Other Streaming
View model →
◎
Microsoft

MAI-Voice-2

MAI-Voice

Expressive long-form narration, audiobooks, podcasts, educational content, voice-over, accessibility, and high-fidelity branded audio

Audio Other Audio input
View model →
◎
Microsoft AI

MAI-Voice-2-Flash

MAI-Voice-2

Low-latency expressive speech for voice agents, assistants, call centers, IVR systems, and interactive multilingual applications

Audio Other Streaming
View model →
◎
Xiaomi MiMo

MiMo-Audio-7B-Base

MiMo-Audio

Few-shot audio-language research, speech continuation, voice and style conversion, speech translation, speech editing, and audio-text experimentation

Audio Multimodal 8,192 ctx Audio input
View model →
◎
Xiaomi

MiMo-Audio-7B-Instruct

MiMo-Audio

Local audio understanding, speech-to-text dialogue, spoken conversational agents, and controllable text-to-speech research

Audio Multimodal 8,192 ctx Audio input
View model →
◎
Xiaomi

MiMo-V2.5-TTS

MiMo-V2.5-TTS

Expressive text-to-speech, audiobooks, podcasts, dubbing, character dialogue, voice interfaces, narrated content, and stylized speech or singing.

Audio Other 8,000 ctx Streaming
View model →
◎
Xiaomi MiMo

MiMo-V2.5-TTS-VoiceClone

MiMo-V2.5-TTS

Zero-shot voice cloning, expressive narration, character dialogue, personalized speech, and custom-voice audio production

Audio Other 8,000 ctx Audio input Streaming
View model →
◎
Xiaomi

MiMo-V2.5-TTS-VoiceDesign

MiMo-V2.5-TTS

Custom synthetic voices for narration, characters, podcasts, ASMR, games, assistants, and creative audio production

Audio Other 8,192 ctx Streaming
View model →
◎
MiniMax

MiniMax H3

MiniMax H3

Multimodal commercial video generation, reference-based editing, product and advertising content, short cinematic clips, and locally deployed 768p workflows

Audio Multimodal Image input Audio input Video input
View model →
◎
MiniMax

MiniMax Music 2.0

MiniMax Music

Generating complete songs with expressive vocals, lyrics, melodies, instrumental arrangements, duets, a cappella passages, and cinematic musical soundscapes.

Audio Other
View model →
◎
MiniMax

MiniMax Music 2.6

MiniMax Music

Text-to-music, instrumental generation, game and video scoring, detailed musical direction, and genre reinterpretation with Cover mode

Audio Other Audio input
View model →
◎
MiniMax

MiniMax Music 3.0

MiniMax Music

Local generation of complete songs from lyrics and structured musical descriptions

Audio Other 5,000 ctx
View model →
◎
MiniMax

MiniMax Music Cover

MiniMax Music

Reinterpreting existing songs in new genres, vocal styles, arrangements, and production directions while preserving the source melody

Audio Other Audio input
View model →
◎
MiniMax

MiniMax Speech 2.6 Turbo

Speech 2.6

Real-time voice agents, conversational assistants, customer-service automation, interactive characters, multilingual speech, and low-latency text-to-speech

Audio Other Streaming
View model →
◎
NVIDIA

NVIDIA Nemotron 3 VoiceChat

Nemotron 3 VoiceChat

Real-time full-duplex voice agents, interruptible conversational interfaces, speech-to-speech research, and NVIDIA GPU-based enterprise voice applications

Audio Multimodal Audio input Streaming
View model →
◎
Alibaba Cloud

qwen-audio-3.0-realtime-flash

Qwen-Audio-3.0-Realtime

Low-latency voice assistants, real-time customer service, duplex speech conversations, interactive voice agents, and applications requiring streaming audio responses.

Audio Multimodal 40,960 ctx Audio input Tool use Web search
View model →
◎
Alibaba Cloud Model Studio

qwen-audio-3.0-realtime-plus

Qwen-Audio

Low-latency duplex voice assistants, real-time customer service, AI companions, and streamed speech-to-speech applications

Audio Multimodal 40,960 ctx Audio input Tool use Web search
View model →
◎
Alibaba Cloud

Qwen-Audio-3.0-TTS-Plus

Qwen-Audio-TTS

Expressive text-to-speech, audiobooks, film and video dubbing, content creation, premium voice services, multilingual speech, dialect synthesis, and voice cloning

Audio Other Streaming
View model →
◎
Alibaba Cloud Model Studio

qwen-audio-3.1-realtime-plus

Qwen-Audio

Low-latency voice assistants, customer service, AI companions, full-duplex spoken interaction, and voice applications using tools or cloned voices.

Audio Multimodal 262,144 ctx Audio input Tool use Web search
View model →
◎
Alibaba Cloud

Qwen2.5-Omni-7B

Qwen2.5-Omni

Multimodal assistants, audio and video understanding, visual question answering, voice interaction, speech instruction following, and local multimodal AI research

Audio Multimodal 32,768 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud

Qwen3-LiveTranslate-Flash

Qwen3-LiveTranslate

Streaming translation of recorded or uploaded audio and video, multilingual subtitles, translated voice tracks, and applications requiring translated text or synthesized speech.

Audio Other 53,248 ctx Audio input Video input Streaming
View model →
◎
Alibaba Cloud

Qwen3-LiveTranslate-Flash-Realtime

Qwen3-LiveTranslate

Real-time multilingual speech interpretation, live voice translation, conference translation, streaming media, and audiovisual translation with text or synthesized speech output

Audio Other 53,248 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Flash

Qwen3.5-Omni

Fast multimodal analysis, long audio understanding, audiovisual question answering, voice assistants and spoken-response applications

Audio Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Flash-Realtime

Qwen3.5-Omni

Low-latency voice assistants, speech-to-speech applications, realtime multimedia analysis, interactive agents, and multimodal conversations.

Audio Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud

Qwen3.5-Omni-Plus

Qwen3.5-Omni

Multilingual voice assistants, speech-enabled multimodal applications, audio-visual analysis, spoken explanations, accessibility tools, and interactive media workflows.

Audio Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Plus-Realtime

Qwen3.5-Omni

Real-time voice assistants, speech-to-speech applications, multimodal customer service, visual conversational agents, live multimedia analysis, and interactive applications requiring controllable speech output.

Audio Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.8-LiveTranslate-Flash-Realtime

Qwen3.8-LiveTranslate

Real-time speech translation, multilingual meetings, live interpretation, translated voice communication, and audiovisual translation with low latency.

Audio Multimodal 53,248 ctx Image input Audio input Streaming
View model →
◎
Alibaba Cloud Model Studio

Qwen3.8-Omni-Flash-Realtime

Qwen3.8-Omni

Real-time voice assistants, speech-to-speech applications, interactive video agents, live media analysis, multimodal customer service, meeting and collaboration interfaces, and applications requiring tool or MCP integration.

Audio Multimodal 196,608 ctx Audio input Video input Tool use
View model →
◎
Meta

SAM Audio

SAM Audio

Prompted audio separation, speech and noise isolation, instrument and vocal extraction, audiovisual sound segmentation, and audio-editing research

Audio Multimodal Audio input Video input
View model →
◎
Meta

SeamlessM4T-Large v2

SeamlessM4T

Multilingual automatic speech recognition, speech-to-text translation, text translation, text-to-speech translation, and speech-to-speech translation

Audio Multimodal Audio input
View model →
◎
ByteDance Seed

Seed Audio 1.0

Seed Audio

Full-scene audio creation, expressive voice generation, dubbing, dialogue, sound effects, ambience, advertising, games, podcasts, and multilingual audio production

Audio Audio Generation Image input Audio input
View model →
◎
ByteDance Seed

Seed1.5 (Doubao-1.5-pro)

Seed1.5

General-purpose Chinese and multilingual assistance, coding, reasoning, image and document understanding, and voice-interaction applications.

Audio Multimodal 32,768 ctx Image input Audio input Tool use
View model →
◎
ByteDance

Seedance 2.0

Seedance 2.0

Multimodal text-to-video and reference-based video creation, cinematic short clips, video editing and extension, multi-shot storytelling, and synchronized audio-video production.

Audio Multimodal Image input Audio input Video input
View model →
◎
ByteDance

Seedance 2.5

Seedance

Long-form audiovisual storytelling, text-to-video, reference-based video generation, creative production, advertising, education, industrial simulation, and video editing.

Audio Multimodal Image input Audio input Video input
View model →
◎
ByteDance Seed

SeedRealtime

SeedRealtime

Real-time audio-visual assistants, scene-aware guidance, live explanation, interactive learning, accessibility, and proactive multimodal collaboration

Audio Multimodal Image input Audio input Video input
View model →
◎
OpenAI

Sora 2

Sora 2

Rapid video concepting, social clips, image-to-video experiments, prototypes, rough cuts, and audiovisual creative iteration

Audio Multimodal Image input
View model →
◎
OpenAI

Sora 2 Pro

Sora 2

Production-quality text-to-video and image-guided video generation, cinematic prototypes, marketing assets, and high-resolution short clips with synchronized audio.

Audio Video Generation Image input
View model →
◎
MiniMax

Speech-02-HD

MiniMax Speech

High-quality multilingual voiceovers, audiobooks, narration, digital characters, advertising, education, and zero-shot voice cloning

Audio Other 10,000 ctx Audio input Streaming
View model →
◎
MiniMax

Speech-02-Turbo

Speech-02

Low-latency multilingual text-to-speech, streaming voice agents, interactive applications, expressive narration, and voice cloning

Audio Other Audio input Streaming
View model →
◎
MiniMax

Speech-2.6-HD

Speech 2.6

High-quality voiceovers, audiobooks, narration, localization, e-learning, game dialogue, accessibility audio, and production speech

Audio Other
View model →
◎
MiniMax

Speech-2.8-HD

Speech 2.8

High-quality expressive narration, audiobooks, podcasts, advertising, character voices, multilingual speech, and applications prioritizing audio fidelity over the lowest latency

Audio Other Streaming
View model →
◎
MiniMax

Speech-2.8-Turbo

Speech 2.8

Real-time text-to-speech, voice assistants, conversational agents, interactive applications, multilingual narration, gaming characters and expressive voice experiences

Audio Other Streaming
View model →
◎
StepFun

StepAudio 3 Gen

StepAudio 3

Zero-shot text-to-speech, natural-language voice design, singing and vocal generation, music, sound effects, ambience, and complete multi-element audio scenes

Audio Other Audio input
View model →
◎
StepFun

StepAudio 3 Music

StepAudio 3

Song generation, instrumental music, lyric-to-song workflows, accompaniment, cover-style synthesis, and rapid music prototyping

Audio Other Audio input
View model →
◎
StepFun

StepAudio 3 Realtime

StepAudio 3

Natural realtime voice conversation, full-duplex interaction, interruption-aware assistants, emotional audio understanding and voice agents that use tools.

Audio Realtime Audio Audio input Tool use Streaming
View model →
◎
StepFun

StepAudio 3 TTS

StepAudio 3

Controllable multilingual text-to-speech, narration, voice interfaces, localization, and expressive spoken-audio generation

Audio Other Streaming
View model →
◎
NVIDIA

Studio Voice

Studio Voice

Real-time enhancement of speech captured with low-quality microphones in noisy or reverberant environments, including broadcast, conferencing, telecommunications, and media production.

Audio Other Audio input Streaming
View model →
◎
OpenAI

TTS-1

TTS-1

Low-latency text-to-speech, realtime-oriented voice interfaces, narration, accessibility, and automated audio generation

Audio Other Streaming
View model →
◎
OpenAI

TTS-1 HD

TTS-1

High-quality text-to-speech generation, narration, accessibility audio, voice interfaces, and downloadable speech content

Audio Other Streaming
View model →
◎
Allen Institute for AI

Unified-IO 2

Unified-IO

Multimodal research, image understanding and generation, audio and video understanding, spatial prediction, embodied AI and robotic-manipulation experiments, and self-hosted academic prototyping.

Audio Multimodal Image input Audio input Video input
View model →
◎
Google DeepMind

Veo 3.1

Veo 3.1

Cinematic text-to-video and image-to-video generation, short-form storytelling, storyboarding, advertising concepts, visual effects exploration, and creative previsualization with synchronized audio

Audio Multimodal 1,024 ctx Image input Video input
View model →
◎
Google DeepMind

Veo 3.1 Fast

Veo 3.1

Fast, high-volume video generation, creative iteration, social content, advertising concepts, and automated production workflows

Audio Video Generation 1,024 ctx Image input Video input
View model →
◎
Google DeepMind

Veo 3.1 Lite

Veo 3.1

High-volume text-to-video and image-to-video generation, rapid creative iteration, social content, advertising variations, and cost-sensitive production workflows

Audio Other Image input
View model →
◎
Mistral AI

Voxtral TTS

Voxtral TTS

Multilingual voice generation, expressive voice agents, zero-shot voice cloning, custom voice adaptation, and low-latency speech output

Audio Text To Speech Audio input Streaming
View model →
◎
Tencent

WAND-Dubbing-Clone-V1

WAND-Dubbing-Clone

Multilingual video translation, voice-preserving dubbing, subtitle translation, online courses, films, and short-form video localization.

Audio Other Audio input Video input
View model →
◎
Tencent Cloud

WAND-Dubbing-Clone-v2

WAND-Dubbing-Clone

Multilingual video localization, translated online courses, short-form video dubbing, film and media localization, and long-video voice-preserving translation

Audio Other Video input
View model →
◎
xAI

grok-tts

Grok TTS

Expressive speech synthesis, voice agents, narration, podcasts, audiobooks, accessibility, and interactive audio applications

Audio Other Streaming
View model →
◎
Meta

SeamlessExpressive

Seamless

Noncommercial research on expressive multilingual speech-to-speech translation, prosody transfer and voice-style preservation

Audio Multimodal Audio input
View model →
◎
Meta

SeamlessStreaming

Seamless

Real-time multilingual speech recognition, simultaneous translation, speech-to-text translation, and speech-to-speech translation

Audio Multimodal Audio input Streaming
View model →
Learn more

About audio generation models

What audio output means in an AI model

Audio output means that an AI model or endpoint returns digitally represented sound as part of its response. The result may be an audio file, a stream of audio chunks, raw audio frames, or encoded audio data that another system can decode and play.

In current AI systems, audio output most often means synthetic speech. For example, a text-to-speech model may receive a paragraph and return spoken audio in a selected voice. A realtime conversational model may accept a spoken question and return a spoken answer while the conversation is still in progress.

The term is broad rather than perfectly standardized. Some providers use audio output specifically for speech synthesis, while others use it for native speech-to-speech interaction or broader sound-generation capabilities. The exact model, endpoint, response format, and supported modality should therefore be checked rather than inferred from a product's marketing language.

What these models actually produce

An audio-output response can take several forms:

  • Speech audio: spoken words rendered in a synthetic voice.
  • Streamed audio: successive chunks delivered while the response is being generated, useful for low-latency conversations.
  • Raw audio samples: such as linear PCM frames with a specified sample rate and channel layout.
  • Encoded audio files: formats such as MP3, WAV, Opus, AAC, FLAC, PCMU, or PCMA, depending on the service.
  • Audio plus transcript data: some realtime systems provide text transcripts alongside generated sound for captions, logs, or application logic.

A transcript accompanying an audio response is not the same thing as the audio itself. The sound carries information that text does not, including timing, pronunciation, tone, pauses, and vocal delivery.

Audio output does not automatically mean music, sound effects, environmental recordings, or unrestricted sound design. Many catalogued models use the term primarily for spoken-language generation. Music and general sound generation should be verified separately.

Audio input is not audio output

Input and output describe different directions of a model's capabilities:

  • A speech-to-text model accepts audio and returns text.
  • An audio-understanding model analyzes a recording and may return a transcript, summary, classification, or structured data.
  • A text-to-speech model accepts text and returns spoken audio.
  • A speech-to-speech model accepts spoken audio and returns spoken audio, often in a live conversation.
  • A multimodal model may understand audio input while offering no native audio generation.

This distinction is important when evaluating a model catalogue. A model can “listen” without speaking, just as a text-to-speech service can speak without understanding an uploaded recording. An application that combines a text model with a separate speech service may offer a voice interface, but the text model itself may still produce only text.

Native audio generation versus application-level voice features

Native audio output is returned directly by the model or endpoint as one of its declared response modalities. The API may expose an audio file, audio content, or streamed audio events in the same request that generates the response.

In other architectures, the language model produces text or a function call. The application then sends that text to a separate text-to-speech provider. The end user still hears a voice, but the original model did not generate the waveform.

These cases are worth separating:

  • The model directly generates audio.
  • The model generates text that a separate TTS system converts into audio.
  • The model emits a tool or function call requesting audio generation.
  • The surrounding product plays or mixes audio supplied by another service.

The distinction affects latency, pricing, voice controls, privacy, reliability, and integration design. When comparing models, verify whether audio is produced natively by the exact model and endpoint you plan to use.

How audio output works at a useful level

In a conventional text-to-speech pipeline, the system analyzes text for language, pronunciation, phrasing, and prosody. A synthesis model then generates a waveform or an intermediate acoustic representation, which is converted into digital audio and optionally encoded into a requested format.

The input may be plain text or may include SSML, voice settings, speaking-rate controls, pronunciation guidance, pitch, style instructions, or speaker configuration. These controls influence delivery, but they are not always guarantees. A model may still mispronounce an unusual name or interpret a style instruction inconsistently.

Native realtime audio models work differently from file-oriented TTS. They may process an ongoing conversation, manage turn-taking, stream audio while generating it, and stop or adjust output when the user interrupts. The response may include both generated sound and a transcript, but the primary output modality is audio.

The representation has practical consequences. Raw PCM is easy to process but requires more bandwidth than compressed formats. Services may use different sample rates for input and output; for example, a realtime system may accept 16 kHz PCM and return 24 kHz PCM. Telephony integrations may require codecs such as PCMU or PCMA.

What audio output is useful for

Audio-output models are useful when information needs to be heard rather than read, or when spoken interaction is faster and more natural than typing.

Narration and read-aloud content

A text-to-speech model can turn articles, lessons, documentation, announcements, or accessibility content into spoken audio. The user supplies text and voice preferences; the application receives an audio file or stream and plays it through a browser, phone, speaker, or media pipeline. Long-form use requires attention to pronunciation, consistency, file size, and voice licensing.

Voice assistants and conversational agents

A realtime speech model can support hands-free interaction. The user speaks, the system processes the turn, and the model returns spoken audio, often with low-latency streaming. This is useful for customer support, interactive characters, navigation, device control, and other situations where waiting for a complete audio file would make the experience feel slow.

Translation and language learning

Audio output can be used to speak translated text, demonstrate pronunciation, or provide conversational practice. The important evaluation criteria are not just voice quality but also language coverage, accent handling, pronunciation, turn-taking, and whether the spoken content faithfully matches the intended translation.

Media production

Generated speech can provide voiceovers for videos, presentations, podcasts, audiobooks, and training materials. In these workflows, file formats, editing compatibility, speaker consistency, retakes, pacing, and rights for commercial use may matter more than realtime response speed.

Accessibility and operational announcements

Screen readers, alerts, public announcements, and hands-free interfaces can use generated speech to make information available without requiring users to read a screen. For these applications, intelligibility, predictable pronunciation, language support, and dependable output are usually more important than expressive performance.

How to compare audio-output models

The right comparison depends on whether the goal is exact narration, interactive dialogue, translation, or broader sound generation. Useful criteria include:

  • Output type: determine whether the model produces downloadable files, streamed speech, native speech-to-speech responses, music, effects, or another audio category.
  • Input requirements: check whether it accepts text only, SSML, audio, images, video, or a combination of modalities.
  • Time to first audio: for voice interfaces, measure how quickly the first playable audio arrives, not only the total completion time.
  • Streaming and interruption support: realtime applications may need incremental output, barge-in handling, cancellation, and natural turn-taking.
  • Formats and codecs: confirm support for the formats your browser, mobile app, media editor, or telephony system can decode.
  • Sample rate and channels: verify technical compatibility and bandwidth requirements, especially when using raw PCM.
  • Voice and speaker controls: compare available voices, languages, accents, speaker count, custom voices, and voice persistence.
  • Controllability: test speed, pitch, volume, pronunciation, pauses, emotion, and style instructions rather than assuming they are deterministic.
  • Text fidelity: test names, dates, numbers, acronyms, punctuation, specialist terms, and multilingual phrases to see whether the spoken result matches the source.
  • Naturalness and intelligibility: assess clarity and listener comfort separately from expressiveness. A dramatic voice is not necessarily easier to understand.
  • Duration and quotas: check maximum input length, output length, session duration, concurrency, file size, and rate limits.
  • Transcript behavior: determine whether the service returns synchronized or final transcripts for captions, logs, and accessibility features.
  • Cost and infrastructure: account for characters or tokens, audio bandwidth, storage, decoding, buffering, and the cost of additional services.

For a live catalogue, these criteria help narrow the available models without treating every audio-output model as interchangeable. A low-latency conversational model and a high-fidelity narration model may both produce audio but solve very different problems.

Limitations and trade-offs

Pronunciation and textual accuracy

Generated speech can mispronounce names, abbreviations, dates, numbers, foreign words, or technical vocabulary. A natural-sounding recording can still say the wrong thing. Test representative content and review important output before publishing or using it in safety-sensitive settings.

Inconsistent style and prosody

Voice instructions may influence pace, emotion, emphasis, or accent without guaranteeing exact control. Long passages can contain uneven rhythm or changes in delivery. Repeated generations may not be acoustically identical, which matters for serialized narration, character voices, and production workflows.

Latency and network dependence

Streaming reduces perceived waiting time but introduces buffering, connection, and interruption problems. Realtime systems also need reliable session management and careful handling of partial audio. File-based synthesis is often simpler to cache and edit but may be too slow for natural conversation.

Format and integration constraints

Different endpoints can support different codecs, sample rates, channel layouts, maximum lengths, and response representations. Raw audio may require more bandwidth, while compressed audio may introduce encoding delay or quality trade-offs. Confirm compatibility before committing to an implementation.

Voice rights, consent, and privacy

Custom voices and voice-cloning features raise additional questions about consent, impersonation, identity misuse, copyright, and disclosure of generated content. Audio recordings may also contain personal or sensitive information. Review retention, processing, access, and licensing terms before sending private material to a hosted service.

Factual reliability

Audio quality does not validate the information being spoken. A polished voice can deliver an inaccurate answer just as easily as text can. Applications should evaluate the factual content separately from pronunciation and vocal naturalness.

Who needs an audio-output model?

You need this type of model when your application must receive sound directly from an AI system, rather than generating audio through a separate post-processing step. Typical examples include:

  • a voice assistant that must respond aloud in real time;
  • a read-aloud or accessibility feature;
  • automated narration, voiceover, or audiobook production;
  • spoken translation or language-learning practice;
  • interactive characters and game dialogue;
  • telephone or contact-center systems that require streaming speech;
  • announcements or alerts generated from changing text.

You may not need a native audio-output model if your application only needs transcription, audio classification, or a text response. You may also prefer a separate TTS service when the language model's reasoning and the voice system should be managed independently. The key question is not whether the overall product can play sound, but whether the exact model and endpoint return the kind of audio your application needs.

A practical evaluation checklist

  1. Define the task: narration, exact text recitation, realtime conversation, translation, music, or general sound generation.
  2. Verify that the exact model and endpoint return audio directly rather than text or a tool call.
  3. Test names, numbers, abbreviations, technical terms, multilingual phrases, and long passages.
  4. Measure time to first audio, total response time, streaming behavior, and interruption handling.
  5. Confirm formats, codecs, sample rates, channels, maximum lengths, and client compatibility.
  6. Assess intelligibility, naturalness, pronunciation, prosody, voice consistency, and textual fidelity separately.
  7. Check languages, accents, speaker count, customization, consent rules, and commercial-use rights.
  8. Review quotas, pricing, bandwidth, storage, privacy, retention, and safety controls.

Audio output is therefore best understood as a concrete delivery capability, not simply a product label. Once you distinguish native generated sound from transcription, audio understanding, tool calls, and application-level playback, it becomes much easier to choose a model that fits the actual voice or audio workflow.