◆
Models by output type

Multimodal Output in AI Models: What It Means and When It Matters

Multimodal output describes what an AI model produces in its response. A model in this category may return an image, an audio stream, a video file, or a combination of media and text. That is different from a model that merely accepts an image, recording, or video and responds with text. This distinction matters when choosing a model for a real application. The surrounding product may combine several models and tools, so a feature that appears multimodal does not necessarily mean that the underlying language model natively generates media. This guide explains the category, how it differs from related capabilities, what the output looks like, and which practical trade-offs matter.
What this means

Multimodal output means a model can produce multiple output modalities. This should not be confused with multimodal input, which only describes the types of information the model can accept.

Multimodal models

214 models currently match this capability.

View all models →
◎
Yandex

Alice AI ART

Alice AI ART

Text-to-image generation, Russian-language visual content, illustrations, graphic design, advertising assets, presentations, landing pages, and product cards

Multimodal Multimodal 500 ctx Image input Streaming
View model →
◎
Amazon

Amazon Nova 2 Sonic

Amazon Nova 2

Real-time voice assistants, customer-service automation, telephony, interactive learning, multilingual conversations, and tool-enabled speech agents.

Multimodal Multimodal 1,000,000 ctx Audio input Tool use Structured output
View model →
◎
Amazon

Amazon Nova Canvas

Amazon Nova

Enterprise image generation and editing, product visualization, advertising and marketing assets, image variations, background removal, virtual try-on, and brand or subject-consistent visual content.

Multimodal Image Generation Image input
View model →
◎
Amazon

Amazon Nova Reel

Amazon Nova

Short-form advertising, marketing concepts, product visualization, storyboards, social video drafts, and image-guided cinematic clips.

Multimodal Other Image input
View model →
◎
Amazon

Amazon Nova Sonic

Amazon Nova

Real-time voice assistants, customer-service automation, interactive education, language learning, and speech-enabled enterprise workflows

Multimodal Multimodal 300,000 ctx Audio input Tool use Streaming
View model →
◎
Tencent

AuK

AuK

Open-source text-to-speech, reference-voice generation, speech and lyric editing, emotion and timbre transformation, speech enhancement, and source separation

Multimodal Other Audio input
View model →
◎
Baidu

Baidu MuseSteamer 2.0

MuseSteamer 2.0

Chinese image-to-video generation, audiovisual storytelling, marketing videos, multi-person dialogue, synchronized speech, sound effects, and cinematic short-form content

Multimodal Multimodal Image input
View model →
◎
OpenAI

ChatGPT-4o

GPT-4o

Fast general-purpose conversations, vision, voice interactions, coding, and everyday productivity

Multimodal Multimodal 128,000 ctx Image input Audio input Video input
View model →
◎
OpenAI

chatgpt-image-latest

GPT Image

Existing ChatGPT image-generation and image-editing integrations

Multimodal Multimodal Image input
View model →
◎
Z.ai

CogVideoX-3

CogVideoX

Short-form text-to-video, image animation, start-and-end-frame transitions, advertising, marketing, realistic scenes, and 3D-style video generation

Multimodal Video Generation Image input
View model →
◎
NVIDIA

Cosmos-Transfer2.5-2B

Cosmos-Transfer2.5

Controllable video world generation, robotics sim-to-real augmentation, autonomous-vehicle simulation, and Physical AI synthetic-data generation

Multimodal Multimodal Image input Video input
View model →
◎
NVIDIA

Cosmos3-Edge

Cosmos 3

Edge physical AI, robotics, visual reasoning, world simulation, video generation, and action-policy prototyping

Multimodal Multimodal 131,072 ctx Image input Video input
View model →
◎
NVIDIA

Cosmos3-Nano

Cosmos 3

Physical AI, robotics, autonomous-vehicle simulation, multimodal world generation, future-state prediction, action reasoning, and synthetic training data

Multimodal Multimodal Image input Audio input Video input
View model →
◎
NVIDIA

Cosmos3-Super

Cosmos 3

High-quality Physical AI simulation, synthetic-data generation, robotics and autonomous-vehicle research, multimodal world modeling, and teacher-model distillation

Multimodal Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
OpenAI

DALL·E 2

DALL·E

Historical research on text-to-image generation, legacy image workflows, and comparisons with newer OpenAI image models.

Multimodal Other Image input
View model →
◎
OpenAI

DALL·E 3

DALL·E

Historical text-to-image generation, concept art, illustration, visual ideation, marketing imagery, and prompt-following research

Multimodal Other
View model →
◎
Baidu

ERNIE-iRAG-1.0

ERNIE-iRAG

Realistic text-to-image generation, reference-grounded visual creation, commercial-style imagery, and applications needing reduced generative-artificiality.

Multimodal Image Generation
View model →
◎
Baidu

ERNIE-iRAG-Edit

ERNIE-iRAG

Object removal, masked image repainting, image variation, and batch image-editing workflows.

Multimodal Image Editing Image input
View model →
◎
NVIDIA

Eye Contact

NVIDIA Maxine Eye Contact

Gaze correction in video conferencing, telepresence, digital-human applications, and video-processing pipelines

Multimodal Other Image input Video input
View model →
◎
Technology Innovation Institute

Falcon Perception

Falcon Perception

Natural-language object grounding, open-vocabulary detection, promptable instance segmentation, crowded-scene perception, robotics and visual inspection pipelines

Multimodal Multimodal 8,192 ctx Image input Structured output
View model →
◎
Google DeepMind

Gemini 2.5 Flash Live

Gemini 2.5

Real-time voice and video agents, speech-to-speech assistants, interactive customer support, tutoring, coaching, and multimodal Live API applications

Multimodal Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 2.5 Flash TTS

Gemini 2.5

Low-latency controllable text-to-speech, voice assistants, narration, read-aloud features, and multi-speaker audio generation

Multimodal Other 8,192 ctx
View model →
◎
Google DeepMind

Gemini 2.5 Pro TTS

Gemini 2.5

High-fidelity single-speaker and multi-speaker narration, audiobooks, podcasts, professional voiceovers, and scripted creative audio

Multimodal Other 8,192 ctx
View model →
◎
Google DeepMind

Gemini 3.1 Flash Live Preview

Gemini 3.1

Low-latency voice agents, real-time dialogue, multimodal live sessions, and interactive audio applications

Multimodal Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.1 Flash TTS

Gemini 3.1 Flash Audio

Controllable expressive speech, narration, accessibility, scripted audio, and multi-speaker TTS prototypes

Multimodal Other 8,192 ctx Streaming
View model →
◎
Google DeepMind

Gemini 3.5 Live Translate

Gemini 3.5 Audio

Low-latency, real-time speech-to-speech translation for calls, meetings, travel, customer support, and multilingual voice applications

Multimodal Multimodal 131,072 ctx Audio input Streaming
View model →
◎
Google DeepMind

Gemini 3.8 Flash TTS

Gemini 3.8

Studio-quality narration, audiobooks, expressive voice acting, complex multi-speaker dialogue, regional accents, difficult pronunciations, long-form narration, voice design, and voice replication.

Multimodal Other 8,192 ctx Streaming
View model →
◎
Google DeepMind

Gemini 3.8 Flash-Lite TTS

Gemini 3.8

High-volume text-to-speech production, low-latency voice-agent cascades, read-aloud applications, voice replication, and everyday single-speaker speech

Multimodal Other 8,192 ctx Streaming
View model →
◎
Google DeepMind

Gemini 3.8 Live

Gemini 3.8

Low-latency voice agents, real-time audio-to-audio dialogue, multimodal assistants, and interactive tool-using applications

Multimodal Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.8 Live Extended Thinking

Gemini 3.8 Audio

Complex real-time voice agents, multi-step problem solving, asynchronous tool workflows, technical support, travel coordination, and spoken STEM or coding tutoring

Multimodal Reasoning 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Deep Research

Gemini Deep Research

Autonomous market research, due diligence, literature reviews, competitive analysis, source-heavy investigations, and cited research reports

Multimodal Other 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Deep Research Max

Gemini Deep Research

Comprehensive market research, competitive analysis, due diligence, literature reviews, and source-rich investigative reports

Multimodal Reasoning 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Omni Flash

Gemini Omni

Fast text-to-video, image-to-video, conversational video editing, video extension, interpolation, marketing content, and short-form cinematic production

Multimodal Multimodal 1,048,576 ctx Image input Video input Streaming
View model →
◎
Z.ai

GLM-Image

GLM-Image

Text-heavy posters, presentation graphics, educational diagrams, social-media layouts, image editing, and open-weight image-generation research

Multimodal Image Generation Image input
View model →
◎
OpenAI

GPT-4o Audio

GPT-4o

Voice assistants, spoken conversational agents, audio-enabled customer service, and applications requiring direct audio understanding and speech generation

Multimodal Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-4o Mini Audio

GPT-4o

Lower-cost audio understanding, conversational voice interfaces, and applications requiring text and spoken-audio input/output

Multimodal Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-4o Mini Realtime

GPT-4o

Low-cost realtime voice assistants, speech-to-speech interfaces, interactive audio applications, and conversational prototypes

Multimodal Lightweight 16,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-4o Mini TTS

GPT-4o Mini

Fast, controllable text-to-speech for narration, voice interfaces, customer service, accessibility, and realtime audio applications.

Multimodal Other 2,000 ctx Streaming
View model →
◎
OpenAI

GPT-4o Realtime

GPT-4o

Low-latency voice assistants, speech-to-speech applications, live translation, language learning, and interactive customer support

Multimodal Multimodal 32,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-Audio

GPT-Audio

Audio-enabled chat applications, voice interfaces, spoken assistants, and applications requiring direct audio understanding and generation through Chat Completions.

Multimodal Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Audio Mini

GPT-Audio

Cost-sensitive, turn-based audio conversations, voice assistants, and audio-enabled applications using function calling

Multimodal Multimodal 128,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-Audio-1.5

GPT-Audio

Audio-in, audio-out conversational applications using the Chat Completions API, including voice assistants and tool-enabled spoken interfaces.

Multimodal Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Image-1

GPT Image

API-based image generation, image editing, reference-image workflows, inpainting, marketing assets, e-commerce imagery, and visual content production

Multimodal Image Generation Image input
View model →
◎
OpenAI

GPT-Image-1 Mini

GPT Image 1

Cost-sensitive image generation and editing, high-volume variations, rapid ideation, previews, lightweight personalization, and draft creative assets.

Multimodal Multimodal Image input
View model →
◎
OpenAI

GPT-Image-1.5

GPT Image

Production image generation, image editing, branded graphics, ecommerce product imagery, marketing assets, and workflows requiring preservation of important visual details

Multimodal Other Image input
View model →
◎
OpenAI

GPT-Image-2

GPT Image

High-quality text-to-image generation, reference-based image editing, text-heavy visual assets, product imagery, marketing creatives, and production design workflows.

Multimodal Multimodal Image input
View model →
◎
OpenAI

GPT-Image-2.5 Flare

GPT Image 2.5

Fast, high-quality image generation and editing, creator content, product experiences, visual search, rapid prototyping, and high-volume workflows

Multimodal Multimodal Image input
View model →
◎
OpenAI

GPT-Image-2.5 Sunburst

GPT-Image-2.5

Precise image editing, detailed creative work, high-fidelity generation, infographics, layouts, and workflows where fewer retries matter more than minimum latency

Multimodal Multimodal Image input
View model →
◎
OpenAI

GPT-Live 1

GPT-Live

Natural low-latency voice agents, customer support, conversational workflows, live assistance, and applications requiring interruption-aware speech interaction

Multimodal Multimodal Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Realtime

GPT-Realtime

Low-latency speech-to-speech voice agents, realtime customer support, education, accessibility, and conversational applications with function calling

Multimodal Realtime 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime Mini

GPT-Realtime

Cost-sensitive realtime voice agents, speech-to-speech applications, interactive assistants, and multimodal interfaces

Multimodal Realtime 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-1.5

GPT-Realtime

Low-latency speech-to-speech voice agents, customer support, realtime assistants, and audio applications that need function calling.

Multimodal Realtime Audio 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2

GPT-Realtime

Reasoning voice agents, speech-to-speech applications, customer support, live assistants, tool-driven workflows, and long conversational sessions

Multimodal Multimodal 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2.1

GPT-Realtime

Low-latency speech-to-speech agents, customer-service voice workflows, realtime tool use, telephony, and multimodal assistants with image input

Multimodal Realtime 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2.1 Mini

GPT-Realtime-2.1

Lower-cost, low-latency realtime voice agents, speech-to-speech assistants, and tool-enabled conversational applications

Multimodal Lightweight 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-Translate

GPT-Realtime

Low-latency spoken translation, multilingual calls, live interpretation, broadcasts, meetings, lessons, video rooms, captions, and translated audio experiences.

Multimodal Other 16,000 ctx Audio input Streaming
View model →
◎
xAI

Grok Imagine Image

Grok Imagine

API-based text-to-image generation, image editing, visual prototyping, creative applications, and per-image billing

Multimodal Other 1,024 ctx Image input
View model →
◎
xAI

Grok Imagine Image 2.0

Grok Imagine

High-quality text-to-image generation, image editing, multi-reference compositing, marketing graphics, product imagery, typography, and design assets

Multimodal Multimodal Image input
View model →
◎
xAI

Grok Imagine Image Quality

Grok Imagine Image

High-quality image generation and editing, realistic product and marketing imagery, detailed scenes, creative assets, and images requiring stronger text rendering or prompt adherence.

Multimodal Multimodal Image input
View model →
◎
xAI

Grok Imagine Video 1.5

Grok Imagine Video

Short-form text-to-video, image-to-video, reference-guided video, cinematic prototyping, marketing clips, and audiovisual creative workflows

Multimodal Other Image input Audio input
View model →
◎
xAI

Grok Imagine Video 1.5 Lite

Grok Imagine Video 1.5

Low-cost text-to-video and image-to-video drafts, social clips, rapid creative iteration, and high-volume generation

Multimodal Lightweight Image input
View model →
◎
xAI

Grok Voice Think Fast 2.0

Grok Voice Think Fast

Realtime voice agents, customer support, telephony, sales, multilingual conversations, and tool-enabled spoken workflows

Multimodal Multimodal Audio input Tool use Web search
View model →
◎
fal

H3 Max

MiniMax H3

Fast short-form text-to-video, image-to-video, reference-guided generation, and synchronized-audio production

Multimodal Multimodal Image input Audio input Video input
View model →
◎
Tencent

Hunyuan3D 2.0

Hunyuan3D

Local image-to-3D asset generation, textured mesh creation, game and design prototypes, Blender workflows, and research on open 3D generative models

Multimodal Other Image input
View model →
◎
Tencent

Hunyuan3D-2.1

Hunyuan3D

Image-to-3D asset creation, game and virtual-world content, product visualization, design prototyping, and self-hosted 3D generation

Multimodal Other Image input
View model →
◎
Tencent Hunyuan

HunyuanDiT-v1.2

HunyuanDiT

Local English- and Chinese-language text-to-image generation, creative prototyping, ComfyUI workflows, LoRA customization, and ControlNet-based image conditioning

Multimodal Other
View model →
◎
Tencent

HunyuanVideo

HunyuanVideo

Research and production experimentation with high-quality local text-to-video generation

Multimodal Other
View model →
◎
Tencent

HunyuanVideo-1.5

HunyuanVideo

Local text-to-video and image-to-video generation, open-source video research, creative prototyping, and developers needing a comparatively lightweight high-quality video model.

Multimodal Other Image input
View model →
◎
Tencent

HunyuanVideo-I2V

HunyuanVideo

Image-to-video generation, reference-image animation, visual effects, creative video prototyping, and locally hosted open-weight video workflows

Multimodal Other Image input
View model →
◎
Tencent

HunyuanWorld-Mirror

WorldMirror

Multi-view and video-to-3D reconstruction, camera and depth estimation, point-map prediction, surface-normal estimation, novel-view synthesis, and 3D Gaussian Splatting workflows

Multimodal Other Image input Video input
View model →
◎
Tencent

HY-3D-3.0

Hunyuan 3D

Text-to-3D, image-to-3D, sketch-to-3D, rapid game and e-commerce asset creation, 3D printing, and production-oriented asset prototyping.

Multimodal Other Image input
View model →
◎
Tencent

HY-3D-3.1

HY 3D

Text-to-3D, image-to-3D, multi-view reconstruction, game assets, product visualization, digital-human content, e-commerce assets and 3D printing workflows

Multimodal Multimodal Image input
View model →
◎
Tencent

HY-3D-Component

HY-3D

Automated decomposition of FBX 3D models into separate model components.

Multimodal Other
View model →
◎
Tencent

HY-3D-Express

HY-3D

Fast text-to-3D and image-to-3D asset generation, prototyping, and automated 3D content pipelines

Multimodal Other Image input
View model →
◎
Tencent

HY-3D-Format

HY-3D

Automated conversion of existing 3D assets between common interchange and delivery formats

Multimodal Other
View model →
◎
Tencent

HY-3D-Motion

HY-3D

Generating short human character animations from natural-language action descriptions

Multimodal Other
View model →
◎
Tencent

HY-3D-Retopology

HY-3D

Automated retopology, polygon reduction, and preparation of existing 3D meshes for games, rendering, animation, and downstream asset workflows

Multimodal Other
View model →
◎
Tencent

HY-3D-Rigging

HY-3D

Automated rigging and skinning of human or animal 3D characters for animation, games, virtual characters, and asset prototyping

Multimodal Other
View model →
◎
Tencent

HY-3D-Texture

HY-3D

Reference-guided texturing of existing OBJ or GLB meshes, including PBR material generation and automated asset preparation

Multimodal Other Image input
View model →
◎
Tencent

HY-3D-UV

HY-3D

Automated UV unwrapping and preparation of 3D assets for texturing, rendering, game development, and digital-content workflows.

Multimodal Other
View model →
◎
Tencent

Hy-Image-3.0

HunyuanImage

Text-to-image generation, reference-guided image creation, marketing graphics, e-commerce content, visual ideation, and image-generation applications requiring custom aspect ratios

Multimodal Other Image input
View model →
◎
Tencent

Hy-Image-3.5-preview

Hy-Image 3.5

High-resolution text-to-image generation, reference-image creation, multi-turn image editing, posters, UI concepts, marketing assets, and product-image workflows

Multimodal Multimodal 100,000 ctx Image input Web search
View model →
◎
Tencent

Hy-World-2.1-panorama

Hy World

Text-to-panorama and image-to-panorama generation for immersive environments, virtual tours, games, simulations, visualization, and 3D-world pipelines.

Multimodal Other Image input
View model →
◎
Tencent

Hy-World-2.1-scene

Hy-World

Generating explorable 3D environments, Gaussian-splat scenes, point clouds, and collision meshes from text or reference images

Multimodal Other Image input
View model →
◎
NAVER

HyperCLOVA X SEED 8B Omni

HyperCLOVA X SEED

Korean-first any-to-any multimodal assistants, speech and vision applications, multimodal research, and self-hosted deployments

Multimodal Multimodal 32,768 ctx Image input Audio input Video input
View model →
◎
MiniMax

image-01

Image-01

Prompt-based image generation, reference-guided variations, commercial visuals, and batch creative production

Multimodal Other Image input
View model →
◎
DeepSeek

Janus-1.3B

Janus

Local research, image understanding, visual question answering, multimodal prototyping, and lightweight text-to-image experimentation

Multimodal Multimodal 4,096 ctx Image input
View model →
◎
DeepSeek

JanusFlow-1.3B

JanusFlow

Local research, visual question answering, image interpretation, and compact text-to-image experimentation

Multimodal Multimodal 4,096 ctx Image input
View model →
◎
Moonshot AI

Kimi-Audio-7B

Kimi-Audio

Fine-tuning and research on speech recognition, audio understanding, audio classification, audio question answering, and speech-audio generation

Multimodal Multimodal 8,192 ctx Audio input Streaming
View model →
◎
Moonshot AI

Kimi-Audio-7B-Instruct

Kimi-Audio

Self-hosted speech recognition, audio understanding, audio question answering, audio captioning, and spoken conversational agents

Multimodal Multimodal Audio input Streaming
View model →
◎
NVIDIA

LipSync

LipSync

Generative lip dubbing, multilingual video localization, broadcasting, conferencing, and digital-human facial animation

Multimodal Multimodal Image input Audio input Video input
View model →
◎
Google DeepMind

Lyria 3 Clip

Lyria 3

Generating short music or audio clips

Multimodal Other
View model →
◎
Google DeepMind

Lyria 3 Pro

Lyria 3

Full-length AI music, soundtrack creation, songwriting, structured compositions, advertising audio, games, and creative production workflows

Multimodal Other Image input
View model →
◎
Google DeepMind

Lyria 3.5

Lyria

Full-length AI-generated songs, vocal music, instrumental arrangements, songwriting experiments, soundtracks, and image-inspired music creation

Multimodal Other 131,072 ctx Image input
View model →
◎
Google DeepMind

Lyria RealTime

Lyria

Interactive instrumental music generation, live musical improvisation, prompt-driven DJ tools, MIDI-controlled experiences, and real-time creative audio applications

Multimodal Other Streaming
View model →
◎
NVIDIA

Magpie TTS Multilingual

Magpie TTS

Multilingual voice agents, accessibility, narration, audiobooks, dubbing, localization and interactive speech applications

Multimodal Other Streaming
View model →
◎
Microsoft

MAI-Image-2.5

MAI-Image

High-quality text-to-image generation, photorealistic imagery, product and marketing visuals, presentation graphics, and precise image-to-image editing.

Multimodal Other 131,072 ctx Image input
View model →
◎
Microsoft AI

MAI-Image-2.5-Flash

MAI-Image-2.5

Fast, cost-conscious text-to-image generation, image editing, creative production, concept visualization, and high-volume image workflows

Multimodal Multimodal 32,000 ctx Image input
View model →
◎
Microsoft AI

MAI-Image-2.5-Pro

MAI-Image-2.5

High-fidelity text-to-image generation, precise image editing, hero imagery, commercial and photorealistic creative work, accurate in-image typography, and visually dense scenes requiring consistent objects, characters, materials, and spatial relationship

Multimodal Multimodal 131,072 ctx Image input
View model →
◎
Microsoft AI

MAI-Image-2.6

MAI-Image

High-quality text-to-image generation, controlled image editing, commercial imagery, product and branding visuals, photorealistic scenes, and multi-reference creative workflows

Multimodal Multimodal Image input Web search
View model →
◎
Microsoft AI

MAI-Image-2.6-Flash

MAI-Image-2.6

Fast, high-volume text-to-image generation, image editing, marketing assets, product imagery and production design

Multimodal Multimodal 32,000 ctx Image input Web search
View model →
◎
Microsoft

MAI-Voice-2

MAI-Voice

Expressive long-form narration, audiobooks, podcasts, educational content, voice-over, accessibility, and high-fidelity branded audio

Multimodal Other Audio input
View model →
◎
Microsoft AI

MAI-Voice-2-Flash

MAI-Voice-2

Low-latency expressive speech for voice agents, assistants, call centers, IVR systems, and interactive multilingual applications

Multimodal Other Streaming
View model →
◎
Microsoft

MedImageParse

BiomedParse

Text-guided biomedical image segmentation, annotation assistance, organ and tumor delineation, pathology-cell analysis, and research-oriented medical imaging pipelines

Multimodal Other Image input Structured output
View model →
◎
Microsoft

MedImageParse 3D

MedImageParse

Text-prompted segmentation of complete CT or MRI volumes, organ and lesion delineation, volumetry, annotation assistance, and medical-imaging research

Multimodal Other Image input
View model →
◎
Xiaomi MiMo

MiMo-Audio-7B-Base

MiMo-Audio

Few-shot audio-language research, speech continuation, voice and style conversion, speech translation, speech editing, and audio-text experimentation

Multimodal Multimodal 8,192 ctx Audio input
View model →
◎
Xiaomi

MiMo-Audio-7B-Instruct

MiMo-Audio

Local audio understanding, speech-to-text dialogue, spoken conversational agents, and controllable text-to-speech research

Multimodal Multimodal 8,192 ctx Audio input
View model →
◎
Xiaomi

MiMo-V2.5-TTS

MiMo-V2.5-TTS

Expressive text-to-speech, audiobooks, podcasts, dubbing, character dialogue, voice interfaces, narrated content, and stylized speech or singing.

Multimodal Other 8,000 ctx Streaming
View model →
◎
Xiaomi MiMo

MiMo-V2.5-TTS-VoiceClone

MiMo-V2.5-TTS

Zero-shot voice cloning, expressive narration, character dialogue, personalized speech, and custom-voice audio production

Multimodal Other 8,000 ctx Audio input Streaming
View model →
◎
Xiaomi

MiMo-V2.5-TTS-VoiceDesign

MiMo-V2.5-TTS

Custom synthetic voices for narration, characters, podcasts, ASMR, games, assistants, and creative audio production

Multimodal Other 8,192 ctx Streaming
View model →
◎
MiniMax

MiniMax H3

MiniMax H3

Multimodal commercial video generation, reference-based editing, product and advertising content, short cinematic clips, and locally deployed 768p workflows

Multimodal Multimodal Image input Audio input Video input
View model →
◎
MiniMax

MiniMax Hailuo 02

Hailuo

Short text-to-video and image-to-video clips, cinematic experiments, advertising concepts, social content, and scenes with complex motion.

Multimodal Multimodal Image input
View model →
◎
MiniMax

MiniMax Hailuo 2.3

Hailuo

Short cinematic videos, image animation, realistic human motion, stylized scenes, visual effects, advertising concepts, and social-media content.

Multimodal Video Generation Image input
View model →
◎
MiniMax

MiniMax Hailuo 2.3 Fast

Hailuo 2.3

Fast image-to-video generation, high-volume short-form content, social media clips, advertisements, and rapid creative iteration

Multimodal Other Image input
View model →
◎
MiniMax

MiniMax Music 2.0

MiniMax Music

Generating complete songs with expressive vocals, lyrics, melodies, instrumental arrangements, duets, a cappella passages, and cinematic musical soundscapes.

Multimodal Other
View model →
◎
MiniMax

MiniMax Music 2.6

MiniMax Music

Text-to-music, instrumental generation, game and video scoring, detailed musical direction, and genre reinterpretation with Cover mode

Multimodal Other Audio input
View model →
◎
MiniMax

MiniMax Music 3.0

MiniMax Music

Local generation of complete songs from lyrics and structured musical descriptions

Multimodal Other 5,000 ctx
View model →
◎
MiniMax

MiniMax Music Cover

MiniMax Music

Reinterpreting existing songs in new genres, vocal styles, arrangements, and production directions while preserving the source melody

Multimodal Other Audio input
View model →
◎
MiniMax

MiniMax Speech 2.6 Turbo

Speech 2.6

Real-time voice agents, conversational assistants, customer-service automation, interactive characters, multilingual speech, and low-latency text-to-speech

Multimodal Other Streaming
View model →
◎
MiniMax

MiniMax T2V-01

Hailuo Video

Short text-driven video concepts, storyboards, and early Hailuo-style cinematic experiments

Multimodal Other
View model →
◎
Allen Institute for AI

MolmoAct 7B-O

MolmoAct

Open robotics research, visual action reasoning, robot-manipulation experiments, and fine-tuning on custom robot datasets

Multimodal Multimodal Image input
View model →
◎
Allen Institute for AI

MolmoAct-7B-D

MolmoAct

Open research and downstream fine-tuning for vision-guided robotic manipulation, spatial reasoning, trajectory planning, and robot action prediction

Multimodal Multimodal 4,096 ctx Image input
View model →
◎
Allen Institute for AI

MolmoMotion-FM

MolmoMotion

Language-guided 3D trajectory forecasting, robotics planning research, and motion-conditioned video generation

Multimodal Multimodal Image input Video input
View model →
◎
Meta

Muse Image 1.0

Muse Image

Text-to-image generation, precise image editing, multi-image composition, anchored visual series, product imagery, creative assets, and grounded visual content.

Multimodal Multimodal Image input Tool use Web search
View model →
◎
Baidu

MuseSteamer-Air-I2V

MuseSteamer

Low-cost image-to-video generation, short social clips, product animation, marketing assets, and turning still images into dynamic scenes

Multimodal Video Generation Image input
View model →
◎
Baidu

MuseSteamer-Air-Image

MuseSteamer Air

Low-cost text-to-image generation, marketing visuals, creative assets, illustrations, and rapid image prototyping

Multimodal Image Generation
View model →
◎
Google DeepMind

Nano Banana

Gemini 2.5 Flash Image

Fast image generation, conversational image editing, image transformation, and high-volume visual workflows

Multimodal Multimodal 65,536 ctx Image input
View model →
◎
Google DeepMind

Nano Banana 2

Gemini Image

Fast, high-volume image generation and editing, visual iteration, marketing assets, diagrams, infographics, localization, and applications requiring image-search grounding

Multimodal Multimodal 131,072 ctx Image input Video input Web search
View model →
◎
Google DeepMind

Nano Banana 2 Lite

Gemini 3.1 Flash Lite Image

Fast, low-cost 1K image generation and editing, rapid visual prototyping, interactive applications, high-volume image variations, storyboarding, and lightweight creative workflows

Multimodal Multimodal 65,536 ctx Image input Video input
View model →
◎
Google DeepMind

Nano Banana Pro

Gemini Image

Professional image generation and editing, complex compositions, product mockups, infographics, branded creative, multilingual localization, and high-fidelity visual prototyping

Multimodal Multimodal 65,536 ctx Image input Web search
View model →
◎
NVIDIA

NVIDIA Isaac GR00T N1.7

Isaac GR00T N1

Humanoid robot manipulation, cross-embodiment policy learning, robot demonstration fine-tuning, physical AI research, and action-sequence deployment

Multimodal Multimodal Image input Video input
View model →
◎
NVIDIA

NVIDIA Nemotron 3 VoiceChat

Nemotron 3 VoiceChat

Real-time full-duplex voice agents, interruptible conversational interfaces, speech-to-speech research, and NVIDIA GPU-based enterprise voice applications

Multimodal Multimodal Audio input Streaming
View model →
◎
Mistral AI

OCR 3

OCR

High-volume document extraction, scanned forms, handwriting, invoices, complex tables, archival digitization, and document-to-knowledge pipelines.

Multimodal Other Image input Structured output
View model →
◎
Alibaba Cloud

qwen-audio-3.0-realtime-flash

Qwen-Audio-3.0-Realtime

Low-latency voice assistants, real-time customer service, duplex speech conversations, interactive voice agents, and applications requiring streaming audio responses.

Multimodal Multimodal 40,960 ctx Audio input Tool use Web search
View model →
◎
Alibaba Cloud Model Studio

qwen-audio-3.0-realtime-plus

Qwen-Audio

Low-latency duplex voice assistants, real-time customer service, AI companions, and streamed speech-to-speech applications

Multimodal Multimodal 40,960 ctx Audio input Tool use Web search
View model →
◎
Alibaba Cloud

Qwen-Audio-3.0-TTS-Plus

Qwen-Audio-TTS

Expressive text-to-speech, audiobooks, film and video dubbing, content creation, premium voice services, multilingual speech, dialect synthesis, and voice cloning

Multimodal Other Streaming
View model →
◎
Alibaba Cloud Model Studio

qwen-audio-3.1-realtime-plus

Qwen-Audio

Low-latency voice assistants, customer service, AI companions, full-duplex spoken interaction, and voice applications using tools or cloned voices.

Multimodal Multimodal 262,144 ctx Audio input Tool use Web search
View model →
◎
Alibaba Qwen

Qwen-Drive-1.0

Qwen-Drive

Autonomous-driving research, driving-scene VQA, 3D BEV perception, trajectory prediction, and embodied-AI experimentation

Multimodal Multimodal Image input Streaming
View model →
◎
Alibaba Cloud

qwen-image-2.0

Qwen-Image

Text-to-image generation, image editing, text rendering in images, photorealistic scenes, creative design, and producing multiple image variants.

Multimodal Image Generation Image input
View model →
◎
Alibaba Cloud Model Studio

qwen-image-2.0-pro

Qwen-Image 2.0

Professional text-to-image generation, image editing, posters, infographics, multilingual in-image text, photorealistic scenes, and reference-based creative production

Multimodal Other Image input
View model →
◎
Alibaba Cloud

qwen-image-3.0

Qwen Image 3.0

Fast text-to-image generation, text-heavy layouts, image editing, reference-image compositing, marketing graphics, and high-volume image workflows.

Multimodal Other Image input
View model →
◎
Alibaba Cloud

qwen-image-3.0-pro

Qwen Image

Complex text-to-image layouts, multilingual typography, posters, menus, storyboards, interface mockups, product visuals, and image editing with one to three reference images

Multimodal Image Generation 4,500 ctx Image input
View model →
◎
Alibaba Cloud Model Studio

qwen-image-edit

Qwen-Image-Edit

Natural-language single-image editing, bilingual text changes, object insertion or removal, style transfer, pose changes, and image fusion

Multimodal Other Image input
View model →
◎
Alibaba Cloud Model Studio

qwen-image-edit-max

Qwen-Image-Edit

High-quality image editing, multi-image composition, industrial design concepts, geometric transformations, character-consistent edits, and controlled visual revisions

Multimodal Other Image input
View model →
◎
Alibaba Cloud

qwen-image-edit-plus

Qwen Image

Instruction-based image editing, multi-image fusion, character-consistent compositions, object replacement, style transfer, poster and text editing, and product-image variation.

Multimodal Image Editing Image input
View model →
◎
Alibaba Cloud

Qwen-Image-Max

Qwen Image

Realistic text-to-image generation, creative concepts, marketing visuals, editorial artwork, product concepts, and general-purpose image creation.

Multimodal Image Generation
View model →
◎
Alibaba Cloud Model Studio

qwen-image-plus

Qwen-Image

Cost-sensitive text-to-image generation, posters, marketing graphics, illustrations, and images containing Chinese or English text

Multimodal Image Generation
View model →
◎
Alibaba Cloud

Qwen2.5-Omni-7B

Qwen2.5-Omni

Multimodal assistants, audio and video understanding, visual question answering, voice interaction, speech instruction following, and local multimodal AI research

Multimodal Multimodal 32,768 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud

Qwen3-LiveTranslate-Flash

Qwen3-LiveTranslate

Streaming translation of recorded or uploaded audio and video, multilingual subtitles, translated voice tracks, and applications requiring translated text or synthesized speech.

Multimodal Other 53,248 ctx Audio input Video input Streaming
View model →
◎
Alibaba Cloud

Qwen3-LiveTranslate-Flash-Realtime

Qwen3-LiveTranslate

Real-time multilingual speech interpretation, live voice translation, conference translation, streaming media, and audiovisual translation with text or synthesized speech output

Multimodal Other 53,248 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Flash

Qwen3.5-Omni

Fast multimodal analysis, long audio understanding, audiovisual question answering, voice assistants and spoken-response applications

Multimodal Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Flash-Realtime

Qwen3.5-Omni

Low-latency voice assistants, speech-to-speech applications, realtime multimedia analysis, interactive agents, and multimodal conversations.

Multimodal Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud

Qwen3.5-Omni-Plus

Qwen3.5-Omni

Multilingual voice assistants, speech-enabled multimodal applications, audio-visual analysis, spoken explanations, accessibility tools, and interactive media workflows.

Multimodal Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Plus-Realtime

Qwen3.5-Omni

Real-time voice assistants, speech-to-speech applications, multimodal customer service, visual conversational agents, live multimedia analysis, and interactive applications requiring controllable speech output.

Multimodal Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.8-LiveTranslate-Flash-Realtime

Qwen3.8-LiveTranslate

Real-time speech translation, multilingual meetings, live interpretation, translated voice communication, and audiovisual translation with low latency.

Multimodal Multimodal 53,248 ctx Image input Audio input Streaming
View model →
◎
Alibaba Cloud Model Studio

Qwen3.8-Omni-Flash-Realtime

Qwen3.8-Omni

Real-time voice assistants, speech-to-speech applications, interactive video agents, live media analysis, multimodal customer service, meeting and collaboration interfaces, and applications requiring tool or MCP integration.

Multimodal Multimodal 196,608 ctx Audio input Video input Tool use
View model →
◎
NVIDIA

Relighting

AI4M Relighting

Real-time video relighting, virtual production, media effects, HDR-based lighting changes, and foreground/background compositing.

Multimodal Other Image input Video input Streaming
View model →
◎
Meta

SAM 3.1

Segment Anything

Open-vocabulary object detection, pixel-level image segmentation, and multi-object video tracking

Multimodal Other Image input Video input Structured output
View model →
◎
Meta

SAM 3D Objects

SAM 3D

Single-image reconstruction of textured 3D objects from natural scenes, 3D computer-vision research, Gaussian-splat workflows, and rapid asset prototyping

Multimodal Other Image input
View model →
◎
Meta

SAM Audio

SAM Audio

Prompted audio separation, speech and noise isolation, instrument and vocal extraction, audiovisual sound segmentation, and audio-editing research

Multimodal Multimodal Audio input Video input
View model →
◎
Meta

SeamlessM4T-Large v2

SeamlessM4T

Multilingual automatic speech recognition, speech-to-text translation, text translation, text-to-speech translation, and speech-to-speech translation

Multimodal Multimodal Audio input
View model →
◎
ByteDance Seed

Seed Audio 1.0

Seed Audio

Full-scene audio creation, expressive voice generation, dubbing, dialogue, sound effects, ambience, advertising, games, podcasts, and multilingual audio production

Multimodal Audio Generation Image input Audio input
View model →
◎
ByteDance Seed

Seed GR-3

Seed GR

Embodied robotics research, long-horizon manipulation, bimanual control, dexterous object handling, and adapting robot policies to new objects and tasks

Multimodal Other Image input
View model →
◎
ByteDance Seed

Seed1.5 (Doubao-1.5-pro)

Seed1.5

General-purpose Chinese and multilingual assistance, coding, reasoning, image and document understanding, and voice-interaction applications.

Multimodal Multimodal 32,768 ctx Image input Audio input Tool use
View model →
◎
ByteDance

Seedance 2.0

Seedance 2.0

Multimodal text-to-video and reference-based video creation, cinematic short clips, video editing and extension, multi-shot storytelling, and synchronized audio-video production.

Multimodal Multimodal Image input Audio input Video input
View model →
◎
ByteDance

Seedance 2.5

Seedance

Long-form audiovisual storytelling, text-to-video, reference-based video generation, creative production, advertising, education, industrial simulation, and video editing.

Multimodal Multimodal Image input Audio input Video input
View model →
◎
ByteDance Seed

SeedRealtime

SeedRealtime

Real-time audio-visual assistants, scene-aware guidance, live explanation, interactive learning, accessibility, and proactive multimodal collaboration

Multimodal Multimodal Image input Audio input Video input
View model →
◎
ByteDance Seed

Seedream 5.0 Pro

Seedream 5.0

Professional image generation and editing, high-density infographics, advertising and e-commerce creative, multilingual visual content, precise spatial edits, layer separation, and multi-reference image workflows

Multimodal Multimodal Image input Tool use Web search
View model →
◎
SenseTime

SekoIDX

SenseNova Seko

Character-consistent image generation for multi-episode videos, motion comics, short dramas, storyboards, and cross-shot visual production

Multimodal Image Generation Image input
View model →
◎
SenseTime

SekoTalk-1.0

SekoTalk

Audio-driven digital humans, lip-sync video, singing avatars, multilingual character animation, multi-person dialogue, and long-duration talking-video generation

Multimodal Video Generation Image input Audio input
View model →
◎
SenseTime

SenseNova U1

SenseNova U1

Open-source visual understanding, image generation, image editing, infographic creation, visual reasoning, and continuous image-text workflows.

Multimodal Multimodal Image input
View model →
◎
SenseTime

SenseNova U1 Fast

SenseNova U1

Fast infographic generation, dense visual explanations, charts, diagrams, presentation graphics, and information-heavy layouts

Multimodal Lightweight
View model →
◎
SenseTime

SenseNova U1 Pro

SenseNova U

Professional image creation, infographics, advertising, e-commerce assets, presentations, educational diagrams, storyboards, and multi-step visual delivery workflows

Multimodal Multimodal Image input
View model →
◎
SenseNova

SenseNova U1.5 Lite

SenseNova U1.5

Text-to-image generation, reference-image creation, visual design, posters, infographics, product imagery, and iterative image editing

Multimodal Multimodal Image input
View model →
◎
SenseTime

SenseNova-U1.5-8B-MoT

SenseNova-U1.5

Native image generation, high-resolution visual creation, image editing, infographic and layout generation, visual understanding, and multimodal research or creative workflows

Multimodal Multimodal Image input
View model →
◎
SenseTime

SenseNova-Vision-7B-MoT

SenseNova-Vision

Unified computer-vision research, detection, OCR, segmentation, depth and normal estimation, visual grounding, and multi-view geometry

Multimodal Multimodal Image input Structured output
View model →
◎
OpenAI

Sora 2

Sora 2

Rapid video concepting, social clips, image-to-video experiments, prototypes, rough cuts, and audiovisual creative iteration

Multimodal Multimodal Image input
View model →
◎
OpenAI

Sora 2 Pro

Sora 2

Production-quality text-to-video and image-guided video generation, cinematic prototypes, marketing assets, and high-resolution short clips with synchronized audio.

Multimodal Video Generation Image input
View model →
◎
MiniMax

Speech-02-HD

MiniMax Speech

High-quality multilingual voiceovers, audiobooks, narration, digital characters, advertising, education, and zero-shot voice cloning

Multimodal Other 10,000 ctx Audio input Streaming
View model →
◎
MiniMax

Speech-02-Turbo

Speech-02

Low-latency multilingual text-to-speech, streaming voice agents, interactive applications, expressive narration, and voice cloning

Multimodal Other Audio input Streaming
View model →
◎
MiniMax

Speech-2.6-HD

Speech 2.6

High-quality voiceovers, audiobooks, narration, localization, e-learning, game dialogue, accessibility audio, and production speech

Multimodal Other
View model →
◎
MiniMax

Speech-2.8-HD

Speech 2.8

High-quality expressive narration, audiobooks, podcasts, advertising, character voices, multilingual speech, and applications prioritizing audio fidelity over the lowest latency

Multimodal Other Streaming
View model →
◎
MiniMax

Speech-2.8-Turbo

Speech 2.8

Real-time text-to-speech, voice assistants, conversational agents, interactive applications, multilingual narration, gaming characters and expressive voice experiences

Multimodal Other Streaming
View model →
◎
StepFun

Step Edge Gen

Step Edge

On-device text-to-image generation, local image editing, privacy-sensitive creative features, and low-latency edge-device applications.

Multimodal Other Image input
View model →
◎
StepFun

Step Edge GUI

Step Edge

Low-latency desktop and mobile GUI automation, visual grounding, local computer-use agents, and privacy-sensitive edge workflows

Multimodal Other Image input Tool use
View model →
◎
StepFun

StepAudio 3 Gen

StepAudio 3

Zero-shot text-to-speech, natural-language voice design, singing and vocal generation, music, sound effects, ambience, and complete multi-element audio scenes

Multimodal Other Audio input
View model →
◎
StepFun

StepAudio 3 Music

StepAudio 3

Song generation, instrumental music, lyric-to-song workflows, accompaniment, cover-style synthesis, and rapid music prototyping

Multimodal Other Audio input
View model →
◎
StepFun

StepAudio 3 Realtime

StepAudio 3

Natural realtime voice conversation, full-duplex interaction, interruption-aware assistants, emotional audio understanding and voice agents that use tools.

Multimodal Realtime Audio Audio input Tool use Streaming
View model →
◎
StepFun

StepAudio 3 TTS

StepAudio 3

Controllable multilingual text-to-speech, narration, voice interfaces, localization, and expressive spoken-audio generation

Multimodal Other Streaming
View model →
◎
NVIDIA

StreamPETR

StreamPETR

Camera-only multi-view 3D perception, autonomous-driving scene analysis, bird's-eye-view visualization, and object tracking

Multimodal Other Image input Video input Structured output
View model →
◎
MiniMax

T2V-01-Director

Video-01

Text-to-video generation with explicit cinematic camera-movement direction, short advertising concepts, storyboards, and controlled visual experiments

Multimodal Other
View model →
◎
Amazon

Titan Image Generator G1 v2

Titan Image Generator G1

Text-to-image generation, image editing, reference-guided composition, background removal, color-controlled visuals, image variations, and subject-consistent branded content.

Multimodal Other Image input
View model →
◎
OpenAI

TTS-1

TTS-1

Low-latency text-to-speech, realtime-oriented voice interfaces, narration, accessibility, and automated audio generation

Multimodal Other Streaming
View model →
◎
OpenAI

TTS-1 HD

TTS-1

High-quality text-to-speech generation, narration, accessibility audio, voice interfaces, and downloadable speech content

Multimodal Other Streaming
View model →
◎
Allen Institute for AI

Unified-IO

Unified-IO

Multimodal research, vision-language experiments, image generation, visual question answering, dense computer-vision tasks, and academic benchmarking

Multimodal Multimodal Image input
View model →
◎
Allen Institute for AI

Unified-IO 2

Unified-IO

Multimodal research, image understanding and generation, audio and video understanding, spatial prediction, embodied AI and robotic-manipulation experiments, and self-hosted academic prototyping.

Multimodal Multimodal Image input Audio input Video input
View model →
◎
Google DeepMind

Veo 3.1

Veo 3.1

Cinematic text-to-video and image-to-video generation, short-form storytelling, storyboarding, advertising concepts, visual effects exploration, and creative previsualization with synchronized audio

Multimodal Multimodal 1,024 ctx Image input Video input
View model →
◎
Google DeepMind

Veo 3.1 Fast

Veo 3.1

Fast, high-volume video generation, creative iteration, social content, advertising concepts, and automated production workflows

Multimodal Video Generation 1,024 ctx Image input Video input
View model →
◎
Google DeepMind

Veo 3.1 Lite

Veo 3.1

High-volume text-to-video and image-to-video generation, rapid creative iteration, social content, advertising variations, and cost-sensitive production workflows

Multimodal Other Image input
View model →
◎
NVIDIA

Video Super Resolution NIM

Video Super Resolution

Professional video upscaling, broadcast enhancement, streaming pipelines, pre-encoding optimization, denoising, and deblurring

Multimodal Other Video input Streaming
View model →
◎
Mistral AI

Voxtral TTS

Voxtral TTS

Multilingual voice generation, expressive voice agents, zero-shot voice cloning, custom voice adaptation, and low-latency speech output

Multimodal Text To Speech Audio input Streaming
View model →
◎
Tencent

WAND-Dubbing-Clone-V1

WAND-Dubbing-Clone

Multilingual video translation, voice-preserving dubbing, subtitle translation, online courses, films, and short-form video localization.

Multimodal Other Audio input Video input
View model →
◎
Tencent Cloud

WAND-Dubbing-Clone-v2

WAND-Dubbing-Clone

Multilingual video localization, translated online courses, short-form video dubbing, film and media localization, and long-video voice-preserving translation

Multimodal Other Video input
View model →
◎
Tencent

WAND-Vega-Image-1.0-Flash

WAND-Vega-Image

Fast everyday image creation, social-media graphics, marketing materials, short-video covers, and reference-guided image generation

Multimodal Multimodal Image input
View model →
◎
Tencent Cloud

WAND-Vega-Image-1.0-Lite

WAND-Vega-Image

Low-cost, high-volume text-to-image and reference-to-image generation, including e-commerce product imagery, batch visual assets, and marketing content

Multimodal Other Image input
View model →
◎
Tencent

WAND-Vega-Image-1.0-Pro

WAND-Vega-Image 1.0

Professional text-to-image and reference-to-image generation, brand visuals, refined product imagery, high-quality design assets, and 1K-to-4K creative production.

Multimodal Image Generation Image input
View model →
◎
Tencent Cloud

WAND-Vega-Video-1.0-Lite

WAND-Vega-Video

Fast, cost-conscious generation of short e-commerce, advertising, social media, and reference-guided video assets

Multimodal Lightweight Image input Audio input Video input
View model →
◎
xAI

grok-tts

Grok TTS

Expressive speech synthesis, voice agents, narration, podcasts, audiobooks, accessibility, and interactive audio applications

Multimodal Other Streaming
View model →
◎
DeepSeek

Janus-Pro-1B

Janus-Pro

Local multimodal research, image understanding, visual question answering, and compact text-to-image experimentation

Multimodal Multimodal 4,096 ctx Image input
View model →
◎
DeepSeek

Janus-Pro-7B

Janus-Pro

Local image understanding, text-to-image generation, and unified multimodal research

Multimodal Multimodal 4,096 ctx Image input
View model →
◎
Qwen

Qwen-Image-2.1

Qwen-Image

Local text-to-image generation, multi-reference composition, transparent asset creation, product visuals, and localized image editing

Multimodal Multimodal Image input
View model →
◎
Meta AI

SAM 3D Body

SAM 3D

Single-image 3D human mesh recovery, pose and shape estimation, AR/VR, robotics perception, and computer-vision research

Multimodal Other Image input
View model →
◎
Meta

SeamlessExpressive

Seamless

Noncommercial research on expressive multilingual speech-to-speech translation, prosody transfer and voice-style preservation

Multimodal Multimodal Audio input
View model →
◎
Meta

SeamlessStreaming

Seamless

Real-time multilingual speech recognition, simultaneous translation, speech-to-text translation, and speech-to-speech translation

Multimodal Multimodal Audio input Streaming
View model →
Learn more

About multimodal output models

What multimodal output means

Multimodal output is the ability of an AI model or model endpoint to directly produce one or more non-text types of content. Common examples include images, audio, video, and sometimes combinations of these with text.

A text-to-image model that returns an image, a text-to-speech model that returns an audio file, and a video-generation model that creates a video asset all produce multimodal output. A realtime voice model that returns synthesized audio also belongs in this category.

The term is not used consistently across the AI industry. Some providers use “multimodal” to describe a model that can understand several input types, even when it only returns text. For a capability catalogue, the important question is therefore not whether a provider calls a model multimodal, but what the endpoint actually returns.

Input and output are separate capabilities

One of the most common mistakes is assuming that a model which can process media can also generate it. These are different abilities.

  • An image-understanding model may accept a photograph and return a written description.
  • An audio-transcription model may accept a recording and return text.
  • A video-understanding model may answer questions about a clip without creating a new video.
  • An image-generation model may accept a text prompt and return a new image.
  • A speech-generation model may accept text and return spoken audio.

The first three examples use multimodal input but do not necessarily provide multimodal output. The latter examples directly generate a non-text result. A useful model record should list input modalities and output modalities independently.

What the model actually returns

Media output can be delivered in several technical forms. An API might return binary data, base64-encoded content, a downloadable file, a URL, a file identifier, a structured content part, or a stream of small audio chunks. The response may also include text such as a caption, metadata, safety information, or generation details.

Images

Image models commonly return PNG, JPEG, WebP, or another image representation. Important properties can include dimensions, aspect ratio, quality, background treatment, number of images, and whether the image is newly generated or edited from a supplied reference.

Audio

Audio output may be a complete file or a low-latency stream. Text-to-speech systems often expose voice, language, format, pronunciation, and style controls. Realtime systems may return audio incrementally while accepting audio input, which is different from waiting for a finished recording.

Video

Video generation is often asynchronous. The initial request may create a job or operation, and the application must wait for completion before downloading the resulting file. Video output can include synchronized or native audio, but this should be confirmed from the model documentation rather than assumed.

Native generation, tools, and applications

Native multimodal output means that the documented model endpoint itself generates the media. If an image endpoint creates image data or a speech endpoint returns synthesized audio, the output capability belongs directly to that model.

Tool-mediated output is different. A language model may return a function call containing instructions for an external image, audio, or video service. The complete application may then show the resulting media, but the language model's direct response was a tool call or structured data, not the media itself.

Applications can obscure this distinction by combining a text model, media-generation model, search system, storage service, and user interface behind one feature. When evaluating a capability, identify which component created the output. Function calling and JSON responses are not automatically multimodal output; they are usually structured text or control data.

Why this output is useful

Multimodal output is valuable when text alone is not the final result a person or system needs. The user supplies instructions, references, or source material, and the model returns media that can be reviewed, edited, played, published, or passed to another step in a workflow.

  • Visual creation: A designer can provide a written brief and reference images, then receive concept art, product variations, illustrations, or background edits.
  • Accessibility and narration: An application can convert written content into spoken audio for screen-reader alternatives, learning materials, or hands-free use.
  • Voice interaction: A realtime system can accept speech and return spoken responses for assistants, tutoring, customer support, or translation.
  • Video production: A creator can provide a script, prompt, or still image and receive a short animated clip for a storyboard, advertisement, lesson, or social post.
  • Media pipelines: One model can create an image, another can animate it, and a speech system can add narration. In this case, the workflow is multimodal even if no single model performs every step.

The output normally requires additional handling. An application may need to store the file, convert its format, moderate it, add metadata, place it in a content-management system, or send it to another model.

What to compare between models

The best model depends on the output you need and how that output will be used. Comparing only general intelligence or prompt quality is not enough.

Output type and delivery

First confirm whether the endpoint produces images, audio, video, or a combination. Then check whether the result arrives as bytes, a file, a URL, a content part, or a stream. The representation affects storage, playback, latency, and integration effort.

Control and fidelity

For images and video, examine prompt adherence, reference-image support, masks, editing controls, aspect ratios, resolution, keyframes, duration, and consistency across generations. For audio, consider voice selection, pronunciation, prosody, emotional direction, language coverage, and streaming support.

Consistency and reliability

Media generation can vary from one request to the next. Video models may show flicker, object changes, identity drift, or incorrect physical movement. Image models may struggle with exact layouts, small text, counting, hands, or logos. Audio models may mispronounce names or produce unnatural emphasis. Test representative examples rather than relying only on demonstrations.

Latency, limits, and cost

Some outputs are returned immediately, while others require a queued or asynchronous job. Compare generation time, streaming behavior, retry requirements, maximum duration or resolution, file-size limits, and the number of outputs allowed per request. Pricing may be calculated per image, request, audio duration, video duration, resolution, or token, so costs should be compared using the same unit.

Safety, privacy, and rights

Check content filters, rejection behavior, watermarking or provenance markers, retention policies, and restrictions on user-provided images, voices, or recordings. Commercial use may also depend on provider terms, training-data policies, likeness rights, music rights, and the status of generated or uploaded material.

Limitations and trade-offs

Non-text output usually requires more processing and storage than a text response. High-resolution images and longer videos can be slower and more expensive, while realtime audio requires a stable low-latency connection and suitable streaming infrastructure.

Generated media is not automatically accurate. An image may contain incorrect text or geometry, a video may change a character's appearance between frames, and a voice system may mispronounce specialized terms. Media should be reviewed when factual accuracy, identity, safety, or brand consistency matters.

Controllability is also model-specific. A model may accept a reference image but not support precise masks, fixed seeds, multiple keyframes, or reliable edits. A speech model may offer expressive instructions but limited control over individual phonemes. A video model may generate attractive clips but provide little control over camera motion or continuity.

Privacy is an important consideration when prompts include personal recordings, faces, documents, or proprietary designs. Applications should understand where inputs and outputs are processed, how long they are retained, and whether generated files are accessible through public or temporary URLs.

Do not confuse multimodal output with related categories

  • Multimodal input: the ability to accept images, audio, or video does not prove that the model can generate them.
  • Vision or media understanding: analyzing a photograph or video and returning text is not image or video generation.
  • Speech-to-text: processing audio to produce a transcript is audio input with text output, not audio generation.
  • Structured output: JSON, XML-like data, and schema-constrained responses are normally text or machine-readable data, not direct non-text media.
  • Function calling: a tool call instructs another service to act. It is not the same as the model returning the resulting image, audio, or video.
  • Product-level multimodality: an application may combine several specialized models, even when its central language model has text-only output.

Who needs multimodal output?

This capability is useful when the final user experience requires media rather than an explanation about media. Choose it when an application must create visuals, play generated speech, conduct spoken interaction, produce video, or pass generated media into a downstream production workflow.

You may not need a multimodal-output model if the task is limited to describing images, transcribing recordings, extracting information from video, generating JSON, or deciding which external tool should be called. In those cases, a text-output model with the appropriate input capability may be sufficient.

The practical decision is straightforward: define the final artifact first, then verify that the exact endpoint creates that artifact directly. This avoids confusing a product's overall feature set with the native output capabilities of one model.

How to evaluate a model in this category

  1. Identify the exact model, endpoint, and version.
  2. Record input and output modalities separately.
  3. Confirm whether media is generated natively or by an external tool.
  4. Inspect the actual response format and integration requirements.
  5. Test realistic prompts, references, languages, voices, resolutions, and durations.
  6. Measure quality, consistency, latency, failure rates, and retry behavior.
  7. Check controllability features and hard output limits.
  8. Review safety, privacy, provenance, licensing, and retention terms.
  9. Compare cost using the correct unit for the output type.

Provider documentation illustrates why these checks matter. Image, speech, realtime, and video endpoints often expose different interfaces and constraints even when they are offered within the same product family. Model names and availability can change, so the exact endpoint documentation and a small production-like test are more reliable than a broad marketing label.