AI model directory

AI Models

Explore AI models across providers and compare what they are designed to do. Browse language and reasoning models, coding models, multimodal systems, image and video models, audio and speech models, embeddings, realtime models and other specialized AI systems.

915 models tracked
915 Total models
33 Providers
21 Model types
823 Current / accessible
Model directory

Browse all models

01.AI General Purpose

Local bilingual text generation, research, fine-tuning, offline applications, and resource-conscious deployment

Context 4K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available open-weight model; older first-generation Yi base model

01.AI General Purpose

Long-document completion, local deployment, English-Chinese text generation, research, and downstream fine-tuning

Context 200K
Reasoning 4/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as downloadable open weights; legacy-generation model

Input No official hosted API price verified; downloadable weights
Output No official hosted API price verified; downloadable weights
01.AI General Purpose

Local bilingual English-Chinese chat, personal projects, academic experimentation, and fine-tuning on modest open-weight infrastructure

Context 4K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Open-weight and downloadable; legacy historical model

Input No official hosted API price verified; self-hosted weights have no per-token provider charge
Output No official hosted API price verified; self-hosted weights have no per-token provider charge
01.AI General Purpose

Local text generation, code completion, mathematics, bilingual English-Chinese applications, research, and downstream fine-tuning

Context 4K
Reasoning 6/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available as an open-weight downloadable model; older Yi-generation model with no verified first-party hosted API offering for this exact model

01.AI General Purpose

Long-context document processing, code generation, mathematics, bilingual English-Chinese text generation, local inference, and domain fine-tuning

Context 200K
Reasoning 5/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight base model; legacy generation but still publicly downloadable

Input No official hosted API price; downloadable weights
Output No official hosted API price; downloadable weights
01.AI General Purpose

Self-hosted bilingual text generation, research, custom fine-tuning, coding experiments, and English-Chinese applications

Context 4K
Reasoning 6/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight model; older-generation base checkpoint

Input No official hosted API price verified; self-hosted weights
Output No official hosted API price verified; self-hosted weights
01.AI General Purpose

Long-document analysis, English-Chinese generation, retrieval experiments, research, and self-hosted fine-tuning

Context 200K
Reasoning 7/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available as downloadable open-weight model; no official hosted API availability or retirement date verified

01.AI General Purpose

Self-hosted bilingual assistants, English-Chinese dialogue, open-weight LLM research, private inference, quantization, and fine-tuning

Context 4K
Reasoning 6/10
Speed 3/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Open-weight; downloadable; legacy-generation model with no verified provider-managed hosted API availability

01.AI General Purpose

Long-context chat, complex text analysis, multilingual generation, prediction, and general-purpose enterprise language applications.

Context 32K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Streaming
Status

Available through 01.AI API documentation; current lifecycle details are not explicitly stated by the provider.

Input $3 per 1 million input tokens
Output $3 per 1 million output tokens
01.AI General Purpose

Cost-sensitive hosted chat, Chinese-English generation, coding, mathematics, reasoning, summarization and high-volume API workloads

Context 16K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Streaming
Status

Proprietary hosted API model; current public availability and lifecycle status are not clearly documented

Input $0.14 per 1 million input tokens historically reported; verify current pricing with 01.AI
Output $0.14 per 1 million output tokens historically reported; verify current pricing with 01.AI
01.AI Lightweight

Local text generation, self-hosted applications, experimentation, domain adaptation, and fine-tuning

Context 4K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as an open-weight downloadable model for self-hosted and local inference; no hosted inference provider is currently listed on its Hugging Face model page.

01.AI General Purpose

Local conversational applications, instruction following, lightweight coding, bilingual English-Chinese text generation, experimentation, and self-hosted inference

Context 4K
Reasoning 4/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Open-weight; accessible for local deployment; no explicit retirement date found

01.AI General Purpose

Local text generation, research, fine-tuning, Chinese-English applications, and cost-sensitive self-hosted deployments

Context 4K
Reasoning 6/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Open-weight and downloadable; no official deprecation or shutdown date found

Input No official first-party hosted price; downloadable weights under Apache 2.0
Output No official first-party hosted price; downloadable weights under Apache 2.0
01.AI General Purpose

Self-hosted conversational assistants, local text generation, lightweight coding help, multilingual experimentation, and fine-tuning research.

Context 4K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight model; official weights remain accessible, with no exact provider-published deprecation or shutdown date verified.

Input No official hosted API price verified; self-hosted weights are available under Apache 2.0.
Output No official hosted API price verified; self-hosted weights are available under Apache 2.0.
01.AI General Purpose

Local bilingual text generation, research, domain adaptation, and fine-tuning

Context 4K
Reasoning 7/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning
Status

Open-weight and currently downloadable; standard 4K-context base checkpoint

01.AI General Purpose

Self-hosted bilingual assistants, English-Chinese text generation, coding support, mathematics, reasoning, fine-tuning, and privacy-sensitive deployments

Context 4K
Reasoning 7/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Open-weight model remains available for self-hosted deployment; 01.AI hosted model-platform API service ended on 2026-09-03

01.AI Coding

Local code completion, code generation, multilingual programming tasks, long-context source-code analysis, and downstream fine-tuning

Context 131K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as downloadable open-weight model; no exact first-party hosted API availability verified

Input No official first-party hosted API price identified; downloadable weights are available under Apache 2.0
Output No official first-party hosted API price identified; downloadable weights are available under Apache 2.0
01.AI Coding

Local code generation, completion, debugging, code explanation, lightweight IDE assistants, and model fine-tuning experiments

Context 131K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight model

Input No official hosted API pricing; model weights are openly available
Output No official hosted API pricing; model weights are openly available
01.AI Coding
01.AI
Yi-Coder

Yi-Coder-9B

Local code generation, completion, editing, repository-scale context, multilingual programming, and fine-tuning

Context 131K
Reasoning 6/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available open-weight base model

Input No official hosted API price; downloadable weights for self-hosting or third-party deployment
Output No official hosted API price; downloadable weights for self-hosting or third-party deployment
01.AI Coding

Self-hosted coding assistance, code generation, debugging, code explanation, code translation, and long-context repository analysis

Context 131K
Reasoning 6/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Open-weight and publicly available; no verified official hosted API listing for this exact model

Input No official hosted API price verified; downloadable weights are available under Apache 2.0
Output No official hosted API price verified; self-hosting and third-party inference costs vary
01.AI Multimodal

Local bilingual image understanding, visual question answering, OCR-oriented image analysis, and image-to-text applications

Context 4K
Reasoning 5/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Open-weight multimodal model family; current accessibility and active maintenance are not clearly documented by 01.AI

01.AI Multimodal

Local bilingual image understanding, visual question answering, OCR-assisted extraction, image summarization, and lightweight multimodal experimentation

Context 4K
Reasoning 4/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Available as an open-weight self-hosted model; no current official hosted inference provider is listed on its model page

01.AI Multimodal

Self-hosted bilingual image understanding, visual question answering, image text recognition, and research applications with substantial GPU capacity

Context 4K
Reasoning 6/10
Speed 3/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Available as open-weight downloadable model; no current first-party hosted API availability verified

01.AI Coding
01.AI
Yi Large

Yi Large FC

Function calling, tool selection, agent orchestration, and structured workflow automation

Context 33K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Documented by 01.AI; current live availability not independently verified

Input $3 per 1 million tokens
Output $3 per 1 million tokens
01.AI General Purpose

Low-cost short-context chat, text generation, summarization, classification, and general language applications

Context 4K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Streaming
Status

Documented in the 01.AI API catalog; current operational availability should be verified through the provider

Input $0.19 per 1 million tokens
Output $0.19 per 1 million tokens
AI21 Labs General Purpose

Long-context text generation, open-weight research, experimentation, and private or self-hosted deployment

Context 256K
Reasoning 6/10
Speed 8/10
Outputs
Text
Status

Legacy open-weight model; downloadable and accessible through AI21 Labs' official Hugging Face repository

AI21 Labs General Purpose

Long-document analysis, grounded generation, enterprise RAG, document summarization, information extraction, private deployment, and multilingual text workflows.

Context 256K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Current and available; the API endpoint jambalarge-1.7 points to the dated snapshot jambalarge-1.7-2025-07.

Input $2 per 1M input tokens
Output $8 per 1M output tokens
AI21 Labs General Purpose

Long-context document analysis, RAG, structured generation, function calling, multilingual enterprise assistants, and private deployment

Context 256K
Reasoning 5/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Legacy; superseded by newer Jamba releases, including AI21-Jamba-Mini-1.7. Public model weights remain available.

AI21 Labs General Purpose

Long-context document analysis, retrieval-augmented generation, enterprise assistants, structured text generation, multilingual workflows, and self-hosted deployments

Context 262K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Legacy and superseded by newer Jamba Large releases; downloadable open-weight checkpoint remains available

AI21 Labs General Purpose

Long-context retrieval-augmented generation, enterprise document analysis, grounded question answering, structured extraction, classification, and private deployment

Context 256K
Reasoning 6/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available; older Jamba 1.6 generation with Jamba 1.7 available as a newer successor

Input $2 per 1 million input tokens
Output $8 per 1 million output tokens
AI21 Labs General Purpose

Long-context RAG, grounded question answering, enterprise document processing, classification, structured text generation, function calling and privacy-sensitive private deployments.

Context 256K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Legacy/open-weight model; downloadable from AI21's official Hugging Face repository and available for self-managed deployment, but not featured among AI21's current primary Jamba models as of September 25, 2026.

AI21 Labs General Purpose

Long-document analysis, enterprise RAG, grounded question answering, structured text generation, private deployment, and cost-sensitive text workflows

Context 256K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available as an AI21 open-weight model; hosted API availability is not currently verified

Input $0.20 per 1 million input tokens when offered through AI21-hosted inference
Output $0.40 per 1 million output tokens when offered through AI21-hosted inference
AI21 Labs Lightweight

Long-context RAG, grounded enterprise question answering, document extraction, local inference, on-device assistants, and lightweight agent workflows

Context 262K
Reasoning 5/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current and available; open-weight release

AI21 Labs General Purpose

Long-context enterprise question answering, grounded generation, document analysis, instruction-heavy workflows, RAG systems and self-hosted deployments

Context 256K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Streaming
Status

Current open-weight model; available through Hugging Face and AI21 Studio

AI21 Labs Reasoning

Local reasoning, long-context document analysis, private RAG, extraction, coding assistance, and lightweight agent controllers

Context 256K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current open-weight model; available for download and local inference

Input No official AI21 hosted API price; self-hosted/open-weight model
Output No official AI21 hosted API price; self-hosted/open-weight model
Aleph Alpha General Purpose
Aleph Alpha
Llama-3.1-8B TFree HAT

Llama-3.1-8B TFree HAT Base

Research on tokenizer-free language modeling, English-German text generation, multilingual NLP, open-weight deployment, and custom model adaptation

Context 262K
Reasoning 5/10
Speed 4/10
Outputs
Text
Status

Available as open-weight research software through Hugging Face; not deployed by an inference provider

Aleph Alpha General Purpose

Steerable multilingual text generation, summarization, classification, question answering, and explainability-oriented enterprise workflows

Reasoning 4/10
Speed 6/10
Outputs
Text
Capabilities
Streaming
Status

Available; first-generation Luminous control model

Input $56.25 per 1 million tokens
Output $61.88 per 1 million tokens
Aleph Alpha General Purpose

Historical multilingual text completion, language understanding research, and compatibility work involving Aleph Alpha’s original Luminous API

Reasoning 5/10
Speed 3/10
Outputs
Text
Status

Legacy generation; historical API and Playground availability documented, current public availability unverified

Aleph Alpha General Purpose

Zero-shot multilingual text generation, instruction following, classification, conversational prototypes, and explainability-oriented enterprise workflows.

Context 2K
Reasoning 6/10
Speed 3/10
Outputs
Text
Capabilities
Streaming
Status

Legacy model; documented in Aleph Alpha SDK materials, with current endpoint availability dependent on provider access and deployment status.

Input $218.75 per 1 million tokens
Output $240.63 per 1 million tokens
Aleph Alpha Multimodal
Aleph Alpha
MAGMA

MAGMA

Vision-language research, image captioning, visual question answering, and experiments with adapter-based multimodal fine-tuning

Reasoning 3/10
Speed 3/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Legacy research/demo model; publicly released checkpoint and source code remain available, but it is not documented as a current hosted commercial model

Input No official hosted API pricing documented
Output No official hosted API pricing documented
Aleph Alpha General Purpose

Multilingual text generation, classification, summarization, question answering, engineering and automotive applications, and safety-conscious research deployments

Context 8K
Reasoning 4/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available open-weight release; safety-aligned variant

Input No public per-token price for the downloadable model; commercial hosted or on-premise access is available by agreement
Output No public per-token price for the downloadable model; commercial hosted or on-premise access is available by agreement
Aleph Alpha General Purpose

English and German instruction following, multilingual research, tokenizer-free language-model experimentation, and self-hosted deployment

Context 262K
Reasoning 6/10
Speed 5/10
Outputs
Text
Status

Available as downloadable open-weight research software

Input No official hosted API pricing; self-hosted deployment costs apply
Output No official hosted API pricing; self-hosted deployment costs apply

Local image understanding, visual question answering, captioning, document and chart analysis, counting, pointing, and multimodal research

Context 4K
Reasoning 6/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Available as an open-weight downloadable checkpoint

Input No official hosted API pricing; self-hosted weights
Output No official hosted API pricing; self-hosted weights

Local image-and-text understanding, visual question answering, document and chart analysis, image captioning, and multimodal research

Context 4K
Reasoning 6/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Available open-weight checkpoint; preview release

Input No official hosted API pricing; downloadable weights for self-hosted deployment
Output No official hosted API pricing; deployment cost depends on self-hosted or third-party infrastructure

Self-hosted image understanding, visual question answering, captioning, image grounding, pointing, counting, and multimodal research.

Context 4K
Reasoning 7/10
Speed 2/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning Streaming
Status

Available open-weight legacy research model; newer Molmo 2 models are Ai2's current successor family.

Allen Institute for Artificial Intelligence (Ai2)
Molmo 2

Molmo2-4B

Efficient local image and video understanding, visual grounding, pointing, captioning, counting, tracking, and multimodal research

Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input
Status

Current open-weight model

Input No official hosted API pricing; self-hosted model weights
Output No official hosted API pricing; self-hosted model weights

Open multimodal research, local image and video understanding, visual grounding, pointing, counting, captioning, and tracking

Context 66K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Current open-weight model

Input Not applicable; open-weight model with no official hosted API token pricing found
Output Not applicable; open-weight model with no official hosted API token pricing found

Open research and downstream fine-tuning for vision-guided robotic manipulation, spatial reasoning, trajectory planning, and robot action prediction

Context 4K
Reasoning 7/10
Speed 5/10
Outputs
Text Actions
Capabilities
Image input Multimodal input Fine-tuning
Status

Available open-weight research model; superseded by MolmoAct2 as Ai2's newer MolmoAct generation

Robotic manipulation research, action reasoning, downstream mid-training, and reproducing zero-shot SimplerEnv experiments

Reasoning 7/10
Speed 4/10
Outputs
Text Actions
Capabilities
Image input Multimodal input Fine-tuning
Status

Available open-weight preview checkpoint

Input No official hosted API pricing; self-hosted checkpoint
Output No official hosted API pricing; self-hosted checkpoint
Allen Institute for Artificial Intelligence (Ai2)
MolmoPoint

MolmoPoint-8B

Open research and applications requiring image or video grounding, visual pointing, object localization, counting, tracking, spatial reasoning, and multimodal analysis.

Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input
Status

Current open-weight model

Input No official hosted API pricing; downloadable checkpoint
Output No official hosted API pricing; downloadable checkpoint

Local inference, open-model research, benchmarking, continued pretraining, and task-specific fine-tuning.

Context 4K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available as a downloadable open-weight model; no official first-party hosted API identified.

Input No official Ai2 hosted API price; self-hosted model weights are downloadable.
Output No official Ai2 hosted API price; self-hosted model weights are downloadable.

Fully open language-model research, local inference, reproducible training experiments, evaluation, and downstream fine-tuning

Context 4K
Reasoning 6/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current downloadable open-weight model; no first-party hosted API deployment verified

Input No official hosted API price; downloadable weights are available
Output No official hosted API price; downloadable weights are available

Open-weight research, continued pretraining, fine-tuning, programming, mathematics, reading comprehension, and long-context language-model experiments

Context 66K
Reasoning 6/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current; open-weight and downloadable

Input No official hosted API price; downloadable weights
Output No official hosted API price; downloadable weights

Open-weight chat, tool-using assistants, multi-turn dialogue, synthetic data generation, self-hosting, and model research

Context 66K
Reasoning 7/10
Speed 5/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available open-weight checkpoint; superseded by Olmo 3.1 32B Instruct for newer instruction-tuned deployments

Input No official first-party hosted API price; downloadable open weights
Output No official first-party hosted API price; downloadable open weights

Open-model reasoning research, mathematics, coding, long-context analysis, and self-hosted deployment

Context 66K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as an open-weight model; superseded by the newer Olmo 3.1 32B Think model

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; deployment and inference costs are user-managed

Open English speech-to-text transcription, captioning, meetings, lectures, calls, podcasts, and local ASR research

Reasoning 1/10
Speed 6/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Available open-weight model

Input No official hosted API pricing; released for open-weight deployment
Output No official hosted API pricing; released for open-weight deployment

Open, self-hosted English transcription for meetings, lectures, calls, podcasts, accessibility, and speech-recognition research

Reasoning 1/10
Speed 5/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current open-weight model; downloadable and usable for local inference

Input No official hosted API pricing; self-hosted open weights
Output No official hosted API pricing; self-hosted open weights

English short- and long-form speech transcription, meeting and call transcription, lecture captioning, podcast processing, broadcast transcription, and speech analytics

Reasoning 1/10
Speed 6/10
Outputs
Text
Capabilities
Audio input
Status

Current open-weight model

Input No official hosted API price; self-hosted checkpoint
Output No official hosted API price; self-hosted checkpoint

Open English speech transcription, meeting and podcast transcription, captioning, timestamped audio indexing, and ASR research

Reasoning 1/10
Speed 7/10
Outputs
Text
Capabilities
Audio input
Status

Current open-weight model; available for download and self-managed inference

Input No official hosted API pricing; open weights available for self-hosted use
Output No official hosted API pricing; output is transcribed text generated during local inference

Efficient self-hosted English transcription, speech research, edge-oriented experiments, and applications where a small open ASR checkpoint is preferred.

Reasoning 1/10
Speed 8/10
Outputs
Text
Capabilities
Audio input
Status

Available open-weight checkpoint

Input No official hosted API price; self-hosted checkpoint
Output No official hosted API price; self-hosted checkpoint
Allen Institute for Artificial Intelligence (Ai2)
OlmoEarth v1.2

OlmoEarth-v1_2-Base

Satellite-image and Earth-observation embeddings, remote-sensing representation learning, geospatial classification, segmentation, and downstream fine-tuning

Reasoning 1/10
Speed 7/10
Outputs
Embeddings
Capabilities
Image input Multimodal input Fine-tuning
Status

Current open-weight model

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; produces feature representations and embeddings
Allen Institute for Artificial Intelligence (Ai2)
OlmoEarth v1.2

OlmoEarth-v1_2-Small

Satellite-image embeddings, remote-sensing representation learning, geospatial segmentation, land-cover analysis, and Earth observation research

Reasoning 1/10
Speed 7/10
Outputs
Embeddings
Capabilities
Image input Multimodal input Fine-tuning
Status

Current open-weight Earth observation foundation model

Input No official public hosted API price; downloadable weights are available
Output No official public hosted API price; model output is primarily embeddings and task representations

Self-hosted instruction following, reasoning, mathematics, coding, open-model research, and reproducible post-training experiments

Context 131K
Reasoning 8/10
Speed 3/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available open-weight model

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; downloadable weights

Large-scale research, instruction-following evaluation, open-weight post-training research, mathematical reasoning, coding benchmarks, and self-hosted experimentation.

Context 8K
Reasoning 8/10
Speed 2/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as an open-weight research and educational model; no hosted inference provider deployment is currently listed on its Hugging Face model page.

Input No official first-party hosted API price; self-hosted open weights
Output No official first-party hosted API price; self-hosted open weights
Allen Institute for Artificial Intelligence (Ai2)
Unified-IO

Unified-IO

Multimodal research, vision-language experiments, image generation, visual question answering, dense computer-vision tasks, and academic benchmarking

Reasoning 4/10
Speed 3/10
Outputs
Text Image
Capabilities
Image input Multimodal input
Status

Open-weight research release; publicly available inference code and checkpoints; legacy relative to Ai2's newer multimodal models

Allen Institute for Artificial Intelligence (Ai2)
Unified-IO

Unified-IO 2

Multimodal research, image understanding and generation, audio and video understanding, spatial prediction, embodied AI and robotic-manipulation experiments, and self-hosted academic prototyping.

Reasoning 5/10
Speed 2/10
Outputs
Text Image Speech Music Actions
Capabilities
Image input Audio input Video input Multimodal input Fine-tuning
Status

Open-weight research release; publicly accessible checkpoints and source code; no official hosted inference API or commercial token pricing identified.

Input No official hosted API pricing; self-hosted research checkpoints
Output No official hosted API pricing; self-hosted research checkpoints
Allen Institute for Artificial Intelligence (Ai2)
WildDet3D

WildDet3D

Open-vocabulary monocular 3D detection, spatial perception, robotics research, augmented reality, and lifting 2D prompts into metric 3D boxes

Reasoning 2/10
Speed 2/10
Capabilities
Image input Multimodal input Fine-tuning
Status

Current open-weight research model

Input No official hosted API pricing; downloadable checkpoint
Output No official hosted API pricing; downloadable checkpoint
Amazon Image Generation

Enterprise image generation and editing, product visualization, advertising and marketing assets, image variations, background removal, virtual try-on, and brand or subject-consistent visual content.

Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input Fine-tuning
Status

Legacy; currently accessible as of 2026-09-25; scheduled for end of life on 2026-09-30

Amazon Multimodal
Amazon
Amazon Nova

Amazon Nova Lite

Low-cost multimodal document analysis, image and video understanding, visual question answering, summarization, RAG, and tool-enabled agents.

Context 300K
Reasoning 5/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Fine-tuning
Status

Active

Input $0.06 per 1 million input tokens
Output $0.24 per 1 million output tokens
Amazon Lightweight

High-volume, low-latency text classification, summarization, translation, extraction, routing, FAQs and narrowly defined fine-tuned tasks

Context 128K
Reasoning 4/10
Speed 10/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

active

Input $0.035 per 1 million input tokens
Output $0.14 per 1 million output tokens
Amazon Multimodal

Long-context multimodal analysis, enterprise document workflows, complex tool calling, agentic orchestration, codebase analysis, and teacher-model distillation before retirement.

Context 1M
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

End-of-Life

Amazon Multimodal
Amazon
Amazon Nova

Amazon Nova Pro

Enterprise multimodal applications, document analysis, visual question answering, video understanding, long-context summarization, RAG, and tool-using assistants

Context 300K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Fine-tuning
Status

Active

Input $0.80 per 1 million input tokens
Output $3.20 per 1 million output tokens
Amazon Other
Amazon
Amazon Nova

Amazon Nova Reel

Short-form advertising, marketing concepts, product visualization, storyboards, social video drafts, and image-guided cinematic clips.

Speed 5/10
Outputs
Video
Capabilities
Image input Multimodal input
Status

Active through the current Nova Reel 1.1 workflow; the original Nova Reel v1.0 model is legacy and scheduled for end of life on 2026-09-30.

Input Not token-priced; video generation is priced per generated video second.
Output $0.08 per generated video second, subject to AWS region, pricing-tier, and current Bedrock pricing conditions.
Amazon Multimodal

Real-time voice assistants, customer-service automation, interactive education, language learning, and speech-enabled enterprise workflows

Context 300K
Reasoning 5/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Streaming
Status

Retired; legacy model with official end-of-life date of September 14, 2026

Amazon Multimodal
Amazon
Amazon Nova 2

Amazon Nova 2 Lite

High-volume multimodal applications, document and video analysis, customer service, business automation, software engineering, long-context workflows, and cost-sensitive AI agents.

Context 1M
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Active; generally available through Amazon Bedrock. AWS states EOL is no sooner than 2026-12-02.

Input $0.30 per 1 million input tokens on the global standard rate; regional and service-tier prices may vary.
Output $2.50 per 1 million output tokens on the global standard rate; regional and service-tier prices may vary.
Amazon Multimodal
Amazon
Amazon Nova 2

Amazon Nova 2 Sonic

Real-time voice assistants, customer-service automation, telephony, interactive learning, multilingual conversations, and tool-enabled speech agents.

Context 1M
Reasoning 6/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Streaming
Status

Active; EOL no sooner than 2026-12-02

Input $0.003 per 1,000 speech-input units; text-input pricing may apply separately and should be verified on the current Amazon Bedrock pricing page.
Output $0.012 per 1,000 speech-output units; text-output pricing may apply separately and should be verified on the current Amazon Bedrock pricing page.
Amazon Other
Amazon
Amazon Nova Act

Amazon Nova Act v1.0

Browser automation, visual UI navigation, repetitive web workflows, agentic QA, tool-oriented tasks, and human-supervised enterprise processes

Reasoning 7/10
Speed 7/10
Outputs
Actions
Capabilities
Image input Multimodal input Tool use
Status

Generally available

Input $4.75 per agent hour; token-level input pricing is not published for this model
Output $4.75 per agent hour; token-level output pricing is not published for this model
Amazon Other
Amazon
Amazon Nova Multimodal Embeddings

Amazon Nova Multimodal Embeddings

Cross-modal semantic search, multimodal RAG, digital asset discovery, recommendations, classification, and clustering

Context 8K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Capabilities
Image input Audio input Video input Multimodal input
Status

Active; generally available through Amazon Bedrock

Input Modality-dependent pricing; AWS charges based on processed input and processing mode. Current pricing should be checked on the Amazon Bedrock pricing page.
Output Not applicable as output-token pricing; the model returns embeddings and pricing is based primarily on input modality and processing mode.
Amazon Multimodal

Multimodal search, text-to-image retrieval, image similarity, visual recommendations, personalization, and image-text matching.

Context 256
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Capabilities
Image input Multimodal input Fine-tuning
Status

Active

Input Text: $0.0008 per 1,000 input tokens; images: $0.00006 per input image. Pricing may vary by AWS Region and current Bedrock pricing terms.
Output No separate output-token price; the model returns embedding vectors.
Amazon Other
Amazon
Titan Image Generator G1

Titan Image Generator G1 v2

Text-to-image generation, image editing, reference-guided composition, background removal, color-controlled visuals, image variations, and subject-consistent branded content.

Reasoning 1/10
Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input Fine-tuning
Status

Retired; AWS lists the model as legacy with an end-of-life date of 2026-06-30.

Input Not token-priced; image-generation pricing applies. AWS pricing examples list $0.01 per 1,024×1,024 standard-quality image.
Output $0.01 per 1,024×1,024 standard-quality image in the AWS pricing example; smaller images and premium quality use different pricing tiers.
Amazon Other

Semantic search, vector indexing, retrieval-augmented generation, personalization, clustering, classification, and recommendation pipelines.

Context 8K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Status

Available; older G1/V1 text-embedding generation

Input $0.10 per 1 million input tokens
Baidu Embedding
Baidu
Embedding-V1

Embedding-V1

Semantic search, vector retrieval, recommendation, semantic matching, knowledge bases, and retrieval-augmented generation

Context 384
Speed 7/10
Outputs
Embeddings
Status

Current and accessible through Baidu Qianfan and AI Studio embedding APIs

Input ¥0.0005 per 1,000 input tokens
Output ¥0.0005 per 1,000 completion tokens in current model-list metadata; the model practically returns vector embeddings rather than generated completion text
Baidu Multimodal

Multimodal understanding, long-context Chinese and English applications, complex reasoning, coding, tool-enabled agents, and enterprise workloads

Context 249K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current and accessible through Baidu Qianfan as of September 25, 2026

Input ¥0.006 per 1K input tokens for requests up to 32K tokens; ¥0.010 per 1K input tokens above 32K
Output ¥0.024 per 1K output tokens for requests up to 32K tokens; ¥0.040 per 1K output tokens above 32K
Baidu Image Generation
Baidu
ERNIE-iRAG

ERNIE-iRAG-1.0

Realistic text-to-image generation, reference-grounded visual creation, commercial-style imagery, and applications needing reduced generative-artificiality.

Reasoning 2/10
Speed 8/10
Outputs
Image
Status

Available; listed in Baidu AI Cloud's inference-service model updates and supported for offline batch inference in Qianfan ModelBuilder documentation.

Baidu Image Editing

Object removal, masked image repainting, image variation, and batch image-editing workflows.

Reasoning 1/10
Speed 6/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Available with access dependent on Baidu Qianfan service configuration; listed for Qianfan batch inference. Related AI作画-iRAG版 API sales stopped on 2026-04-30, but that product notice does not explicitly state that the Qianfan ERNIE-iRAG-Edit endpoi

Baidu Lightweight

Lightweight local text generation, Chinese and English language experimentation, compact conversational systems, and domain adaptation

Context 131K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current open-weight model

Baidu General Purpose

Fast, low-cost Chinese text generation, long-context applications, enterprise agents, content creation, reasoning, and code assistance.

Context 138K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Web search
Status

Current and accessible through Baidu Qianfan; canonical API endpoint is ernie-4.5-turbo-128k.

Input ¥0.0008 per 1,000 input tokens; web-search augmentation is priced at ¥0.004 per 1,000 tokens where applicable.
Output ¥0.0032 per 1,000 output tokens.
Baidu Multimodal
Baidu
ERNIE 4.5 Turbo

ERNIE 4.5 Turbo VL

Cost-efficient multimodal understanding, image and video analysis, OCR, document comprehension, translation, visual question answering, and code-related tasks

Context 128K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Current and available

Input ¥0.003 per 1,000 tokens
Output ¥0.009 per 1,000 tokens
Baidu Multimodal

Local or self-hosted multimodal applications, visual question answering, document and chart understanding, image analysis, video understanding, and efficient vision-language inference

Context 131K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Fine-tuning Streaming
Status

Current open-weight model; publicly released under the Apache 2.0 license

Baidu General Purpose
Baidu
ERNIE 5

ERNIE 5.1

Agentic workflows, web-search-assisted tasks, reasoning, Chinese-language knowledge work, creative writing, and general-purpose assistant applications.

Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Web search
Status

Current; officially released and accessible through Baidu's ERNIE website and AI Studio playground. Public Qianfan API availability for an exact ERNIE 5.1 endpoint was not verified.

Baidu Coding

Code completion, code generation, unit-test generation, code optimization, code explanation, and fine-tuning for software-development workflows.

Context 128K
Reasoning 4/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning
Status

Retired; Qianfan ModelBuilder listed August 14, 2025 as the retirement date.

Baidu Reasoning

Chinese-language reasoning, long-form analysis, complex calculations, literary and document writing, agent workflows, and function calling

Context 33K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Ready; currently accessible through Baidu Qianfan as ernie-x1-turbo-32k

Baidu Reasoning
Baidu
ERNIE X1

ERNIE X1.1

Deep reasoning, Chinese and English question answering, mathematics, coding, factual responses, tool calling, web-grounded applications, and agent workflows

Context 66K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Current preview model

Input ¥0.001 per 1K input tokens
Output ¥0.004 per 1K output tokens
Baidu Video Generation

Low-cost image-to-video generation, short social clips, product animation, marketing assets, and turning still images into dynamic scenes

Reasoning 1/10
Speed 8/10
Outputs
Video
Capabilities
Image input Multimodal input
Status

Current and available through Baidu Qianfan API

Input CNY 1.00 per 5-second video
Baidu Multimodal
Baidu
MuseSteamer 2.0

Baidu MuseSteamer 2.0

Chinese image-to-video generation, audiovisual storytelling, marketing videos, multi-person dialogue, synchronized speech, sound effects, and cinematic short-form content

Speed 7/10
Outputs
Video Speech
Capabilities
Image input Multimodal input
Status

Current model family; concrete Qianfan variants include Turbo, Lite, Pro, Turbo-I2V-Audio, and Turbo-I2V-Effect

Baidu Image Generation
Baidu
MuseSteamer Air

MuseSteamer-Air-Image

Low-cost text-to-image generation, marketing visuals, creative assets, illustrations, and rapid image prototyping

Reasoning 1/10
Speed 8/10
Outputs
Image
Status

Current and available

Input ¥0.05 per image at 1024x1024
Output ¥0.05 per generated image at 1024x1024
Baidu Multimodal
Baidu
PaddleOCR-VL

PaddleOCR-VL-1.5

Multilingual OCR and structured parsing of complex documents, including tables, formulas, charts, seals, scanned pages, warped documents, and screen photographs

Context 131K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Available; superseded by PaddleOCR-VL-1.6

Baidu Lightweight

Fast enterprise question answering, summarization, workflow nodes, agent response generation, and text processing with a 32K context window

Context 33K
Reasoning 5/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Available through Baidu’s documented Wenxin Workshop API; not found in the current Qianfan V2 model-list documentation, so new integrations should verify legacy endpoint compatibility.

Baidu Lightweight
Baidu
Qianfan-Agent-Lite

Qianfan-Agent-Lite-128K

Long-context Agent planning, task decomposition, component selection, and function-calling workflows.

Context 128K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Tool use
Status

Current platform-listed planning model; standalone lifecycle status is not separately published.

Baidu Lightweight

Fast enterprise question answering, lightweight agent planning, component selection, and streaming text applications

Context 8K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current and listed as available in Baidu Qianfan documentation; fast agent-oriented model

ByteDance Seed
Protenix

Protenix

Protein and biomolecular complex structure prediction, computational biology research, molecular design workflows, and self-hosted scientific inference

Reasoning 8/10
Speed 5/10
Capabilities
Fine-tuning
Status

Active open-source biomolecular structure prediction project with multiple model variants, including Protenix-v1 and Protenix-v2

ByteDance Seed Multimodal

Multimodal agent workflows, search and information retrieval, coding agents, GUI interaction, image and video understanding, complex instruction following, and long-context business tasks.

Context 256K
Reasoning 8/10
Speed 8/10
Outputs
Text Actions
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Beta; accessible through BytePlus ModelArk as seed-1-8-251228. The separate LAS multimodal deep-thinking operator using Seed1.8 ended service on 2026-09-20.

Input USD 0.25 per million tokens for prompts up to 128K tokens; USD 0.50 per million tokens for prompts over 128K and up to 256K; cached input USD 0.05 per million tokens.
Output USD 2.00 per million tokens for prompts up to 128K tokens; USD 4.00 per million tokens for prompts over 128K and up to 256K. Batch output pricing is USD 1.00 or USD 2.00 per million tokens by prompt-length tier.
ByteDance Seed General Purpose

Self-hosted general language modeling, long-context research, reasoning experiments, coding assistance, summarization, and foundation-model fine-tuning

Context 524K
Reasoning 7/10
Speed 5/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight model

Input No official hosted API price; downloadable weights are available under Apache-2.0
Output No official hosted API price; deployment cost depends on user infrastructure or third-party hosting
ByteDance Seed Multimodal

General-purpose Chinese and multilingual assistance, coding, reasoning, image and document understanding, and voice-interaction applications.

Context 33K
Reasoning 7/10
Speed 8/10
Outputs
Text Speech
Capabilities
Image input Audio input Multimodal input Tool use Web search
Status

Retired; the primary doubao-1-5-pro-32k-250115 deployment was scheduled to shut down on 2026-09-21 at 14:00 China Standard Time.

ByteDance Seed Multimodal
ByteDance Seed
Seed1.5

Seed1.5-VL

Visual reasoning, image and video understanding, OCR, visual grounding, GUI-agent research, gameplay analysis, and multimodal benchmark evaluation

Context 131K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Retired; Volcano Engine service ended on 2026-03-31

ByteDance Seed Multimodal
ByteDance Seed
Seed1.6

Seed1.6

Multimodal document analysis, visual question answering, coding, mathematics, general reasoning, long-context analysis and adaptive-thinking applications

Context 256K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Deprecated; new endpoint creation stopped on 2026-09-24; existing service scheduled for automatic migration or replacement on 2026-11-24

Deep reasoning, coding, mathematics, logical analysis, document understanding, and visual reasoning over images or videos

Context 256K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Deprecated; scheduled for retirement and migration to a Seed2.0 model

ByteDance Seed Multimodal

Cost-conscious production applications requiring long-context multimodal understanding, document and video analysis, coding assistance, tool use, GUI automation, and structured extraction.

Context 262K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current; latest documented release seed-2-0-lite-260428

Input $0.25 per 1M non-audio input tokens for prompts up to 128K; $0.50 per 1M for prompts above 128K and up to 256K. Audio input: $3.75 per 1M tokens up to 128K and $7.50 per 1M above 128K.
Output $2.00 per 1M output tokens for prompts up to 128K; $4.00 per 1M for prompts above 128K and up to 256K.
ByteDance Seed Lightweight

High-concurrency inference, batch generation, classification, extraction, summarization, and cost-sensitive multimodal workloads

Context 256K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input
Status

Current; available through ByteDance's Volcano Engine model API

Input $0.03 per 1M input tokens
Output $0.31 per 1M output tokens
ByteDance Seed Multimodal

Complex multimodal reasoning, long-chain agent workflows, visual and video analysis, document understanding, scientific research support, coding, and enterprise automation.

Context 200K
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Active and currently accessible through Volcano Engine Ark; canonical deployment version 260215

Input CNY 3.2 per million tokens for 0–32K input; CNY 4.8 per million tokens for 32–128K input; CNY 9.6 per million tokens for 128–256K input. Batch inference input pricing is CNY 1.6, 2.4, and 4.8 per million tokens for the same tiers.
Output CNY 16 per million tokens for 0–32K input; CNY 24 per million tokens for 32–128K input; CNY 48 per million tokens for 128–256K input. Batch inference output pricing is CNY 8, 12, and 24 per million tokens for the same tiers.
ByteDance Seed General Purpose

Complex agent workflows, high-value office and research tasks, long-horizon coding, document and visual analysis, video understanding, and tool-enabled productivity automation.

Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use
Status

Current and officially released; available through ByteDance Seed, Doubao, and Volcano Engine API channels.

ByteDance Seed Multimodal

Fast multimodal agents, coding assistants, document and video analysis, tool-calling workflows, structured data extraction, and cost-sensitive production applications

Context 256K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Current and available through Volcano Engine Ark; API model identifier doubao-seed-2-1-turbo-260628

Input CNY 6 per 1M input tokens for input lengths up to 256K tokens
Output CNY 30 per 1M output tokens
ByteDance Seed Multimodal

Long-form audiovisual storytelling, text-to-video, reference-based video generation, creative production, advertising, education, industrial simulation, and video editing.

Reasoning 2/10
Speed 7/10
Outputs
Video Audio
Capabilities
Image input Audio input Video input Multimodal input
Status

Current; available through ByteDance platforms including Jimeng AI, Doubao Pro, and the Seed platform. API access through BytePlus ModelArk was announced as forthcoming.

ByteDance Seed Multimodal
ByteDance Seed
Seedance 2.0

Seedance 2.0

Multimodal text-to-video and reference-based video creation, cinematic short clips, video editing and extension, multi-shot storytelling, and synchronized audio-video production.

Reasoning 1/10
Speed 6/10
Outputs
Video Speech
Capabilities
Image input Audio input Video input Multimodal input
Status

Current official model page remains available; older generation superseded by Seedance 2.5

ByteDance Seed
Seed GR

Seed GR-3

Embodied robotics research, long-horizon manipulation, bimanual control, dexterous object handling, and adapting robot policies to new objects and tasks

Reasoning 6/10
Speed 6/10
Outputs
Actions
Capabilities
Image input Multimodal input Fine-tuning
Status

Current officially documented research model; no public hosted API, commercial pricing, or downloadable weights verified

ByteDance Seed Multimodal
ByteDance Seed
SeedRealtime

SeedRealtime

Real-time audio-visual assistants, scene-aware guidance, live explanation, interactive learning, accessibility, and proactive multimodal collaboration

Reasoning 7/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current; fully rolled out for large-scale audio-visual full-duplex deployment

ByteDance Seed Multimodal
ByteDance Seed
Seedream 5.0

Seedream 5.0 Pro

Professional image generation and editing, high-density infographics, advertising and e-commerce creative, multilingual visual content, precise spatial edits, layer separation, and multi-reference image workflows

Reasoning 8/10
Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current and publicly documented; available through BytePlus ModelArk and ByteDance Seed services

Input First reference image free; additional reference images $0.003 per image
Output $0.045 per output image for images up to 2.36 million pixels; $0.09 per output image above 2.36 million pixels
ByteDance Seed Multimodal
ByteDance Seed
UI-TARS-1.5

UI-TARS-1.5-7B

Open-weight computer-use research, GUI grounding, browser automation prototypes, screenshot-based interface interaction, and visual action-model experimentation

Context 128K
Reasoning 8/10
Speed 5/10
Outputs
Text Actions
Capabilities
Image input Multimodal input
Status

Available open-weight research release; superseded by UI-TARS-2 for the provider's newer UI-TARS development line

Cerebras General Purpose
Cerebras
Cerebras-GPT

Cerebras-GPT-2.7B

Open-weight language-model research, local text generation, fine-tuning experiments, scaling-law studies, and educational or reference implementations.

Context 2K
Reasoning 3/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available as an open-weight research model

Cerebras General Purpose
Cerebras
Cerebras-GPT

Cerebras-GPT-6.7B

Open LLM research, scaling-law experiments, language-model evaluation, fine-tuning, and self-hosted English text generation

Context 2K
Reasoning 4/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Open-weight and publicly downloadable; research-oriented base model

Cerebras General Purpose
Cerebras
Cerebras-GPT

Cerebras-GPT-13B

Open LLM research, local text generation, reproducible training experiments, and downstream fine-tuning

Context 2K
Reasoning 2/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning
Status

Open-weight, publicly downloadable; no official retirement date identified

Cerebras Lightweight
Cerebras
Cerebras-GPT

Cerebras-GPT-111M

Research, small-scale language-model experimentation, causal text generation, fine-tuning studies, and local deployment

Context 2K
Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available open-weight model; legacy research release

Cerebras Lightweight
Cerebras
Cerebras-GPT

Cerebras-GPT-256M

Compact local language-model experiments, text completion, education, benchmarking, and fine-tuning research

Context 2K
Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available as downloadable open weights; research-oriented and not instruction-tuned

Cerebras Lightweight
Cerebras
Cerebras-GPT

Cerebras-GPT-590M

Local text generation, language-model research, benchmarking, fine-tuning experiments, and studying compute-efficient scaling at small model sizes

Context 2K
Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available open-weight research checkpoint

Cerebras General Purpose

Document-grounded conversational question answering, local retrieval-augmented generation, and open-weight research deployments

Context 8K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as a public open-weight checkpoint; not identified as a current Cerebras-hosted API model

Claude Lightweight

Low-latency assistants, customer support, high-volume text processing, coding assistance, image understanding, and parallel subagent workloads

Context 200K
Reasoning 8/10
Speed 10/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active (latest); retirement not sooner than October 15, 2026

Input $1 per million input tokens; cache write $1.25 per million tokens for 5 minutes or $2 per million tokens for 1 hour; cache read $0.10 per million tokens
Output $5 per million output tokens; batch API output pricing is 50% lower
Claude Reasoning
Claude
Claude Fable 5

Claude Fable 5

Complex reasoning, long-running autonomous agents, advanced coding, multi-stage research, document-heavy analysis, and high-value enterprise workflows

Context 1M
Reasoning 10/10
Speed 3/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active; tentative retirement not sooner than 2027-06-09

Input $10 per million input tokens; $12.50 per million tokens for 5-minute cache writes; $20 per million tokens for 1-hour cache writes; $1 per million tokens for cache reads
Output $50 per million output tokens; $25 per million output tokens for batch processing
Claude Reasoning
Claude
Claude Fable 5

Claude Fable 5.1

Long-running agentic coding, demanding reasoning, multistep research, complex document analysis, spreadsheets, presentations, and high-stakes knowledge work

Context 1M
Reasoning 10/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active (latest)

Input $10 per million input tokens; 5-minute cache writes $12.50 per million tokens; 1-hour cache writes $20 per million tokens; cache reads $0.25 per million tokens
Output $50 per million output tokens; batch output pricing $25 per million tokens
Claude Reasoning
Claude
Claude Mythos

Claude Mythos 5

Advanced cybersecurity, vulnerability research, biology, healthcare, life-sciences research, and long-running technical workflows requiring extensive context and reasoning

Context 1M
Reasoning 10/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active; invite-only; preview/beta on Amazon Bedrock

Input $10 per million input tokens; cached input: $1 per million tokens for cache reads; cache writes: $12.50 per million tokens for 5-minute TTL or $20 per million tokens for 1-hour TTL
Output $50 per million output tokens; Batch API input and output receive a 50% discount
Claude Reasoning
Claude
Claude Mythos

Claude Mythos 5.1

Vetted cybersecurity defense, advanced biology and life-sciences research, complex coding, long-running agents, technical investigation, and high-value research workflows

Context 1M
Reasoning 10/10
Speed 4/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Active; invitation-only through Project Glasswing and trusted-access programs

Input $10 per million input tokens; 5-minute cache write $12.50 per million tokens; 1-hour cache write $20 per million tokens; cache read $0.25 per million tokens
Output $50 per million output tokens; Batch API offers a 50% discount on input and output pricing
Claude Reasoning

Defensive cybersecurity research, vulnerability discovery, attack-surface auditing, autonomous coding, long-running agents, and large-context analysis.

Context 1M
Reasoning 9/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Deprecated; deprecated June 9, 2026. Anthropic recommends migrating to Claude Mythos 5.

Claude Reasoning
Claude
Claude Opus

Claude Opus 4.6

Complex reasoning, agentic coding, repository-scale software work, long-context analysis, research, document and spreadsheet workflows, and enterprise automation

Context 1M
Reasoning 10/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active legacy; retirement not sooner than 2027-02-05

Input $5 per million input tokens; cache write $6.25 per million tokens for 5 minutes or $10 per million tokens for 1 hour; cache read $0.50 per million tokens
Output $25 per million output tokens; Batch API input and output pricing discounted 50%
Claude General Purpose
Claude
Claude Opus

Claude Opus 5

Complex agentic coding, code review, enterprise analysis, long-context research, document workflows, financial and legal reasoning, and multi-step tool-using applications.

Context 1M
Reasoning 10/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active legacy; Anthropic recommends migration to Claude Opus 5.5. Retirement is scheduled no sooner than 2027-07-24.

Input $5 per million tokens; cache writes $6.25 per million tokens for 5 minutes or $10 per million tokens for 1 hour; cache reads $0.50 per million tokens; batch input receives a 50% discount.
Output $25 per million tokens; batch output receives a 50% discount.
Claude Reasoning
Claude
Claude Opus

Claude Opus 5.5

Long-running agentic coding, repository-scale software engineering, code review, knowledge work, multimodal document analysis, and tool-using applications

Context 1M
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active (latest)

Input $4 per million input tokens
Output $20 per million output tokens
Claude Reasoning
Claude
Claude Opus 4

Claude Opus 4.5

Complex software engineering, coding agents, multi-step research, computer-use workflows, enterprise analysis, visual document understanding, and high-value tool-using applications.

Context 200K
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active (legacy)

Input $5 per million input tokens
Output $25 per million output tokens
Claude Reasoning
Claude
Claude Opus 4

Claude Opus 4.7

Complex reasoning, agentic coding, repository-scale software engineering, research, document analysis, computer-use workflows, and high-stakes knowledge work

Context 1M
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active legacy; migration to Claude Opus 5.5 recommended

Input $5 per million input tokens; 5-minute cache write $6.25 per million; 1-hour cache write $10 per million; cache read $0.50 per million; Batch API $2.50 per million
Output $25 per million output tokens; Batch API $12.50 per million
Claude Reasoning
Claude
Claude Opus 4

Claude Opus 4.8

Complex reasoning, agentic coding, long-horizon workflows, enterprise knowledge work, document analysis, vision tasks, and computer-use agents

Context 1M
Reasoning 10/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active legacy model; migration to newer Opus models is recommended for new deployments

Input $5 per million input tokens
Output $25 per million output tokens
Claude Coding
Claude
Claude Sonnet

Claude Sonnet 4.5

Complex coding, software agents, computer-use workflows, visual document analysis, research, and long-running multi-step tasks

Context 200K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Legacy; still available as of 2026-09-24; migration to Claude Sonnet 5 recommended

Input $3 per million tokens; cache write $3.75 per million tokens for 5 minutes or $6 per million tokens for 1 hour; cache read $0.30 per million tokens
Output $15 per million tokens; Message Batches API offers a 50% discount on input and output pricing
Claude General Purpose
Claude
Claude Sonnet

Claude Sonnet 5

Agentic coding, software engineering, browser and computer-use workflows, long-context analysis, tool-driven automation, and high-volume assistants

Context 1M
Reasoning 9/10
Speed 9/10
Outputs
Text Actions
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current and generally available

Input $2 per million input tokens
Output $10 per million output tokens
Claude General Purpose
Claude
Claude Sonnet 4.6

Claude Sonnet 4.6

Agentic coding, computer use, long-context document analysis, enterprise knowledge work, structured extraction, and high-volume applications needing strong reasoning at Sonnet-tier pricing

Context 1M
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active legacy; Anthropic recommends migration to Claude Sonnet 5

Input $3 per million input tokens
Output $15 per million output tokens
Claude General Purpose
Claude
Claude Sonnet 5.5

Claude Sonnet 5.5

Fast coding agents, long-context analysis, image-aware workflows, document creation, and tool-enabled business applications

Context 1M
Reasoning 9/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Active (latest)

Input $2 per million tokens
Output $10 per million tokens
Cohere General Purpose
Cohere
Aya Expanse

Aya Expanse 32B

Multilingual generation, translation, summarization, customer support, global communication, and multilingual research

Context 128K
Reasoning 6/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning
Status

Live

Input $0.50 per 1 million tokens
Output $1.50 per 1 million tokens
Cohere Multimodal
Cohere
Aya Vision

Aya Vision 32B

Multilingual image understanding, OCR, image captioning, visual question answering, image-based translation, visual reasoning and research deployments using open weights.

Context 16K
Reasoning 7/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Streaming
Status

Live

Cohere Speech Recognition
Cohere
Cohere Transcribe

Cohere Transcribe

Multilingual audio transcription, enterprise speech archives, meeting and interview transcription, call-center audio, and low-latency ASR workflows

Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Audio input
Status

Live; open-source research release

Input No per-input token price published. API experimentation is free subject to rate limits; Model Vault uses hourly instance-based pricing.
Output No per-output token price published. API experimentation is free subject to rate limits; Model Vault uses hourly instance-based pricing.
Cohere Other
Cohere
Cohere Transcribe

Cohere Transcribe Arabic

Arabic speech transcription, multilingual Arabic-English audio, regional dialects, code-switched speech, call-center audio and high-throughput ASR workloads.

Reasoning 1/10
Speed 8/10
Outputs
Text
Capabilities
Audio input
Status

Live

Input Free through Cohere API for experimentation subject to rate limits; Model Vault deployment is priced per hour-instance and requires contacting Cohere.
Cohere General Purpose
Cohere
Command A

Command A

Enterprise RAG, long-context document analysis, multilingual applications, tool use, agentic workflows, financial text processing and structured text generation

Context 256K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Live

Input $2.50 per 1 million tokens
Output $10.00 per 1 million tokens
Cohere Multimodal
Cohere
Command A

Command A+

Enterprise agents, multimodal document and image analysis, multilingual workflows, reasoning-intensive automation, retrieval-augmented generation, and tool-using applications.

Context 128K
Reasoning 9/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Live

Input Free until applicable rate limits; production pricing is not publicly listed and may require contacting Cohere.
Output Free until applicable rate limits; production pricing is not publicly listed and may require contacting Cohere.
Cohere Reasoning

Complex enterprise agents, tool use, retrieval-augmented generation, multilingual reasoning, long-context analysis, and workflow automation.

Context 256K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Live

Input Free until applicable API rate limits; private Model Vault deployment is billed per instance-hour rather than per token. Standard Vault pricing lists $48/hour for L and $57.50/hour for XL performance tiers.
Output Free until applicable API rate limits on Cohere-hosted trial and production access; Model Vault pricing is instance-hour based rather than output-token based.
Cohere Specialized

High-quality multilingual text translation, enterprise document translation, and privacy-sensitive translation workflows.

Context 8K
Reasoning 4/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Live

Input Free until applicable rate limits are reached; production access requires contacting Cohere sales.
Output Free until applicable rate limits are reached; production access requires contacting Cohere sales.
Cohere Multimodal

Enterprise document intelligence, OCR, chart and table analysis, visual question answering, and multilingual image understanding

Context 128K
Reasoning 6/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Streaming
Status

Live

Input Free until applicable rate limits; production access requires contacting Cohere
Output Free until applicable rate limits; production access requires contacting Cohere
Cohere General Purpose

Low-cost enterprise RAG, multilingual document workflows, long-context chat, structured extraction, and tool-using agents

Context 128K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Live

Input $0.15 per 1M input tokens
Output $0.60 per 1M output tokens
Cohere Lightweight
Cohere
Command R

Command R7B

Cost-sensitive RAG, enterprise chat, tool use, coding assistance, and fast multi-step agents

Context 128K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Live

Input $0.0375 per 1 million tokens
Output $0.15 per 1 million tokens
Cohere General Purpose

Complex enterprise RAG, long-context document analysis, multilingual assistants, citations, structured data tasks and multi-step tool-use agents.

Context 128K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Live

Input $2.50 per 1 million tokens
Output $10 per 1 million tokens
Cohere Multimodal

Multilingual semantic search, multimodal RAG, PDF and document retrieval, image-to-text retrieval, classification, clustering, and enterprise vector indexing.

Context 128K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Capabilities
Image input Multimodal input
Status

Generally available

Input $0.12 per 1 million text input tokens; $0.47 per 1 million image tokens
Cohere Other

English semantic search, retrieval-augmented generation, vector indexing, classification, clustering, and similarity matching

Context 512
Speed 7/10
Outputs
Embeddings
Capabilities
Image input Multimodal input
Status

Live

Input $0.10 per 1 million input tokens
Output Not applicable; the model returns embeddings rather than billed generated text
Cohere Other

Multilingual semantic search, cross-lingual retrieval, RAG indexing, classification, clustering, and image-text similarity

Context 512
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Capabilities
Image input Multimodal input
Status

Active

Input $0.10 per 1M input tokens
Cohere Coding

Agentic software engineering, repository-level code changes, terminal-based coding agents, code review, local inference, and private deployment.

Context 256K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Live

Input Free until applicable Cohere API rate limits are reached; self-hosted deployment has infrastructure costs rather than Cohere per-token pricing.
Output Free until applicable Cohere API rate limits are reached; self-hosted deployment has infrastructure costs rather than Cohere per-token pricing.
Cohere Other

Enterprise machine translation, multilingual documentation, localization, internal communications, safety procedures, and private or self-hosted translation workflows

Context 16K
Reasoning 6/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Live

Input Cohere API: free until applicable rate limits are reached; Standard Model Vault: $17.50 per instance-hour for XL deployment
Output No separate token-based output price published; Cohere API is free until applicable rate limits are reached
Cohere Multimodal

High-volume enterprise document parsing, table and form extraction, search indexing, RAG ingestion, and document context for AI agents

Context 8K
Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Live; generally available

Input $1.50 per 1,000 pages through the Cohere API
Cohere Other

English semantic reranking for enterprise search, hybrid retrieval, RAG pipelines, FAQs, knowledge bases, documents, code retrieval, and semi-structured records

Context 4K
Reasoning 2/10
Speed 8/10
Status

Active and currently listed by Cohere; older Rerank 3.0 English model with newer Rerank alternatives available

Input $2.00 per 1,000 search units; one search unit covers one query with up to 100 documents, subject to chunking
Cohere Lightweight

Low-latency multilingual search reranking, high-throughput retrieval, enterprise search, hybrid search and RAG pipelines

Context 33K
Reasoning 1/10
Speed 9/10
Status

Current and available

Input Usage-based search-unit pricing; one search unit is one query with up to 100 documents to be ranked
Cohere Other
Cohere
Rerank 4.0

Rerank 4 Pro

High-quality multilingual reranking for enterprise search, RAG pipelines, semantic retrieval, and semi-structured document ranking.

Context 33K
Reasoning 4/10
Speed 7/10
Outputs
Text
Status

Generally available

Input $2.50 per 1,000 search units; one search unit is one query against up to 100 documents. Dedicated Model Vault pricing is separate.
Output Not applicable; the endpoint returns relevance scores and ranking results rather than generated tokens.
Cohere Lightweight

Multilingual translation, conversation, summarization, and text generation focused on African and West Asian languages; local and edge deployment

Context 8K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Live

Cohere Lightweight

Multilingual translation, cross-lingual text generation, localized assistants, education, research, and efficient local or edge deployment

Context 8K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Live

Cohere Lightweight

Efficient multilingual translation, text generation, localization, language learning, and local or edge deployment for European and Asia-Pacific languages

Context 8K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Live; available through the Cohere Chat API and as an open-weight model

Self-hosted general English-language instruction following, coding assistance, text generation, and domain-specific fine-tuning

Context 33K
Reasoning 7/10
Speed 5/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Open-weight model; Databricks Foundation Model API serving retired

Input 10.714 DBU per 1 million input tokens when historically available through Databricks pay-per-token serving; endpoint retired April 30, 2025
Output 32.143 DBU per 1 million output tokens when historically available through Databricks pay-per-token serving; endpoint retired April 30, 2025

Local experimentation, research, instruction-tuning studies, and organizations needing an openly downloadable model with commercial-use licensing

Context 2K
Reasoning 3/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning
Status

Legacy open-weight model; publicly downloadable; no official shutdown date found

Input No official hosted API price; model weights are available for self-hosted use
Output No official hosted API price; model weights are available for self-hosted use
DeepSeek Coding
DeepSeek
DeepSeek-Coder-V2

DeepSeek-Coder-V2

Self-hosted code generation, completion, repository analysis, debugging, code translation, and programming-focused research.

Context 128K
Reasoning 7/10
Speed 5/10
Outputs
Text
Capabilities
Streaming
Status

Legacy open-weight model series; downloadable checkpoints remain available, while the hosted Coder API line was superseded and merged into DeepSeek-V2.5.

DeepSeek Coding

Local code completion, fill-in-the-middle generation, IDE integrations, repository-level coding, and private or self-hosted inference

Context 131K
Reasoning 6/10
Speed 8/10
Outputs
Text
Status

Open-weight and downloadable; accessible through the official Hugging Face repository, with no verified provider-published shutdown date

DeepSeek Coding

Local code generation, code completion, debugging, code explanation, repository-scale prompts, and developers needing an open-weight coding model

Context 128K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available as an open-weight downloadable model; older but still accessible

DeepSeek General Purpose
DeepSeek
DeepSeek-LLM

DeepSeek-LLM 7B

Local bilingual text generation, research, experimentation, instruction tuning, and lightweight self-hosted assistants

Context 4K
Reasoning 5/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Legacy open-weight model; downloadable and usable through compatible local inference tools

Input No official hosted API price; downloadable weights for self-hosting
Output No official hosted API price; self-hosting and infrastructure costs apply
DeepSeek General Purpose
DeepSeek
DeepSeek-LLM

DeepSeek-LLM 67B

Local deployment, bilingual English-Chinese generation, language-model research, mathematics, coding experiments, and custom fine-tuning workflows

Context 4K
Reasoning 6/10
Speed 3/10
Outputs
Text
Status

Legacy open-weight model; downloadable checkpoints remain available, but it is superseded by newer DeepSeek model families

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; downloadable weights
DeepSeek Reasoning
DeepSeek
DeepSeek-Math

DeepSeek-Math-V2

Advanced mathematical reasoning, natural-language theorem proving, proof generation, proof verification, and research on self-correcting reasoning systems

Context 164K
Reasoning 10/10
Speed 2/10
Outputs
Text
Status

Available as an open-weight research model; no official DeepSeek-hosted API deployment or public model-specific API pricing verified

DeepSeek Multimodal
DeepSeek
DeepSeek-OCR

DeepSeek-OCR

Local OCR, document digitization, PDF and image parsing, layout-aware markdown conversion, table extraction, figure parsing, and visual-text compression research

Context 8K
Reasoning 2/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Streaming
Status

Current open-weight model; publicly available for local and self-hosted inference

DeepSeek Multimodal
DeepSeek
DeepSeek-OCR

DeepSeek-OCR 2

Local OCR, scanned-document transcription, layout-aware document parsing, table extraction, and document-to-Markdown workflows

Context 8K
Reasoning 2/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Streaming
Status

Current open-weight model; self-hosted checkpoint

DeepSeek Reasoning
DeepSeek
DeepSeek-Prover

DeepSeek-Prover-V1

Lean 4 theorem proving, formal mathematics, automated proof generation, and theorem-proving research

Reasoning 8/10
Speed 5/10
Outputs
Text
Status

Open-weight research release; legacy predecessor to DeepSeek-Prover-V1.5 and DeepSeek-Prover-V2

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; downloadable weights
DeepSeek Reasoning

Lean 4 theorem proving, formal mathematics, automated proof synthesis, proof-search research, and verifier-guided reasoning

Context 164K
Reasoning 9/10
Speed 2/10
Outputs
Text
Status

Open-weight and downloadable; currently accessible from the official model repository; no first-party hosted API identified for this exact checkpoint

DeepSeek Reasoning
DeepSeek
DeepSeek-Prover V1.5

DeepSeek-Prover-V1.5-Base

Lean 4 theorem proving, formal mathematics research, proof completion, proof-search experiments, and open-weight model fine-tuning

Context 4K
Reasoning 8/10
Speed 4/10
Outputs
Text
Status

Legacy open-weight model; downloadable and usable, but superseded by newer DeepSeek-Prover releases

DeepSeek Reasoning
DeepSeek
DeepSeek-R1

DeepSeek-R1

Mathematical reasoning, coding, technical analysis, research, and self-hosted reasoning applications

Context 128K
Reasoning 9/10
Speed 4/10
Outputs
Text
Capabilities
Streaming
Status

Open-weight checkpoint available; original hosted API identity superseded and scheduled for discontinuation

DeepSeek Reasoning
DeepSeek
DeepSeek-R1

DeepSeek-R1-Zero

Reasoning research, mathematics, coding experiments, reinforcement-learning studies, open-weight evaluation, and model distillation

Context 131K
Reasoning 9/10
Speed 3/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Open-weight and downloadable; experimental/legacy research model; not a current first-party hosted API model

DeepSeek Reasoning

Local mathematical reasoning, compact reasoning experiments, educational applications, lightweight coding assistance, and self-hosted inference on limited hardware

Context 131K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Active open-weight model; downloadable for self-hosted inference

DeepSeek General Purpose
DeepSeek
DeepSeek-V2

DeepSeek-V2

Local or third-party deployment for general text generation, translation, mathematics, research, and code generation

Context 128K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Streaming
Status

Legacy open-weight model; downloadable and usable through local or third-party inference, but not listed in DeepSeek's current first-party hosted API catalog

DeepSeek Lightweight
DeepSeek
DeepSeek-V2

DeepSeek-V2-Lite

Local text generation, Chinese and English language tasks, MoE research, fine-tuning, and efficient self-hosted inference

Context 33K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Open-weight and downloadable; legacy self-hosting model

Input No official DeepSeek-hosted API price documented for this exact model; self-hosted weights are available under the DeepSeek Model License.
Output No official DeepSeek-hosted API price documented for this exact model; self-hosted inference costs depend on hardware and serving infrastructure.
DeepSeek General Purpose
DeepSeek
DeepSeek-V2

DeepSeek-V2.5

Open-weight general language generation, coding assistance, code completion, and self-hosted experimentation

Context 128K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Legacy open-weight model; the V2.5 series was superseded by newer DeepSeek model families, while the model weights remain available

DeepSeek General Purpose
DeepSeek
DeepSeek-V3

DeepSeek-V3

Open-weight general language generation, coding, long-context text tasks, research, and cost-sensitive third-party inference.

Context 128K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Superseded and no longer current as a first-party hosted API model; open-weight checkpoint remains available for self-hosted and third-party deployment.

Input $0.27 per 1M tokens for cache misses; $0.07 per 1M tokens for cache hits, historical launch pricing
Output $1.10 per 1M tokens, historical launch pricing
DeepSeek General Purpose
DeepSeek
DeepSeek-V3

DeepSeek-V3.1

Open-weight reasoning, coding, tool-calling, long-context analysis, and self-hosted agent systems

Context 128K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Legacy hosted API generation; official open-weight release remains available

DeepSeek General Purpose

Open-weight deployment, coding assistance, long-context text processing, reasoning workflows, search agents, and terminal-oriented automation

Context 128K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Open-weight checkpoint available; dedicated DeepSeek API endpoint retired on 2025-10-15

Input $0.56 per 1M input tokens cache miss; $0.07 per 1M cached input tokens during historical API availability
Output $1.68 per 1M output tokens during historical API availability
DeepSeek Reasoning

Difficult mathematics, advanced coding, scientific reasoning, long-form analysis, benchmark evaluation, and research deployment

Context 164K
Reasoning 10/10
Speed 4/10
Outputs
Text
Capabilities
Streaming
Status

Retired hosted API; open-weight model remains available for self-hosting and third-party deployment

Input Historical temporary API pricing was the same as DeepSeek-V3.2; no current hosted price because the endpoint expired on 2025-12-15
Output Historical temporary API pricing was the same as DeepSeek-V3.2; no current hosted price because the endpoint expired on 2025-12-15
DeepSeek Multimodal

Local image understanding, visual question answering, diagram and document analysis, multimodal research, and compact deployments

Context 4K
Reasoning 3/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Available open-weight checkpoint; legacy first-generation model

DeepSeek Reasoning
DeepSeek
DeepSeek V3

DeepSeek-V3.2

Open-weight reasoning, coding, long-context analysis, tool-using agents, research workflows, and cost-sensitive deployments

Context 131K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Legacy open-weight model; former DeepSeek API aliases deepseek-chat and deepseek-reasoner were scheduled for discontinuation on 2026-07-24

Input $0.028 per 1M tokens cached; $0.28 per 1M tokens cache miss during official API availability
Output $0.42 per 1M tokens during official API availability
DeepSeek Lightweight

Long-context reasoning, coding, agent workflows, and cost-sensitive API applications

Context 1M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Retired; legacy API identifier temporarily routed to DeepSeek-V4.1-Flash from September 10, 2026

Input $0.14 per 1 million input tokens
Output $0.28 per 1 million output tokens
DeepSeek General Purpose

Self-hosted language-model research, custom post-training, domain adaptation, and large-context text generation

Context 1.05M
Reasoning 7/10
Speed 6/10
Outputs
Text
Status

Current downloadable open-weight base checkpoint; no Hugging Face Inference Provider deployment listed

DeepSeek Multimodal

Image understanding, screenshot and chart analysis, multimodal coding agents, visual tool-use workflows, and text-plus-image reasoning

Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use
Status

Retired as an independent model on 2026-09-10; legacy API identifier temporarily routes requests to DeepSeek-V4.1-Flash

Input $0.15 per 1M cache-miss input tokens off-peak or $0.30 peak; $0.003 off-peak or $0.006 peak for cache-hit input when using the current routed Flash pricing
Output $0.60 per 1M output tokens off-peak or $1.20 peak when using the current routed Flash pricing
DeepSeek General Purpose
DeepSeek
DeepSeek V4

DeepSeek-V4-Pro

Complex reasoning, coding agents, long-context analysis, tool-using workflows, and large document or codebase processing

Context 1M
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Deprecated for independent serving; deepseek-v4-pro API requests are routed to DeepSeek-V4.1-Flash until V4.1-Pro launches

Input Official V4-Pro list price: US$0.66 per 1M cache-miss input tokens off-peak and US$1.32 peak; US$0.022 per 1M cache-hit input tokens off-peak and US$0.044 peak. DeepSeek announced that routed requests use V4.1-Flash rates.
Output Official V4-Pro list price: US$1.98 per 1M output tokens off-peak and US$3.96 peak. DeepSeek announced that routed requests use V4.1-Flash rates.
DeepSeek Lightweight
DeepSeek
DeepSeek V4.1

DeepSeek-V4.1-Flash

Low-cost, high-throughput reasoning and coding, long-context analysis, agentic workflows, tool calling, and text-plus-image understanding.

Context 1.05M
Reasoning 9/10
Speed 10/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Current and available through the DeepSeek API; the canonical API identifier is deepseek-flash.

Input $0.15 per 1M input tokens off-peak or $0.30 peak for cache misses; $0.003 off-peak or $0.006 peak for cache hits.
Output $0.60 per 1M output tokens off-peak or $1.20 peak.
DeepSeek Multimodal

Local research, image understanding, visual question answering, multimodal prototyping, and lightweight text-to-image experimentation

Context 4K
Reasoning 4/10
Speed 7/10
Outputs
Text Image
Capabilities
Image input Multimodal input
Status

Available as an open-weight research model; superseded in the Janus series by newer Janus-Pro variants but still downloadable and usable.

Input No official hosted API pricing; self-hosted model weights
Output No official hosted API pricing; self-hosted model weights
DeepSeek Multimodal

Local research, visual question answering, image interpretation, and compact text-to-image experimentation

Context 4K
Reasoning 5/10
Speed 7/10
Outputs
Text Image
Capabilities
Image input Multimodal input
Status

Available as an open-weight downloadable checkpoint; no official hosted inference API identified

Input No official hosted API pricing
Output No official hosted API pricing

Browser automation, visual UI interaction, repetitive web workflows, form filling, and user-interface testing

Context 128K
Reasoning 7/10
Speed 6/10
Outputs
Text Actions
Capabilities
Image input Multimodal input Tool use Streaming
Status

Legacy preview; currently documented and accessible, with newer Gemini 3.x models recommended for new computer-use applications

Input $1.25 per 1M tokens for prompts up to 200K tokens; $2.50 per 1M tokens for prompts over 200K tokens
Output $10.00 per 1M tokens for prompts up to 200K tokens; $15.00 per 1M tokens for prompts over 200K tokens
Google DeepMind General Purpose

Large-scale multimodal processing, low-latency reasoning, coding, data extraction, tool-using agents, and applications requiring a very large context window.

Context 1.05M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Stable and currently served through the Gemini API with restricted access for users who have actively used Gemini 2.5 models; not deprecated; no shutdown date announced.

Input $0.30 per 1M text/image/video tokens; $1.00 per 1M audio tokens; batch input $0.15 per 1M text/image/video tokens.
Output $2.50 per 1M tokens, including thinking tokens.
Google DeepMind Lightweight

High-volume classification, simple extraction, lightweight multimodal analysis, routing, tagging, summarization, and extremely latency-sensitive applications.

Context 1.05M
Reasoning 6/10
Speed 10/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Stable; currently served through the Gemini API with access limited to users who have actively used Gemini 2.5 models. Google recommends newer models for new projects.

Input $0.10 per 1M text/image/video input tokens; $0.30 per 1M audio input tokens. Batch/Flex: $0.05 per 1M text/image/video tokens and $0.15 per 1M audio tokens.
Output $0.40 per 1M output tokens. Batch/Flex: $0.20 per 1M output tokens.

Real-time voice and video agents, speech-to-speech assistants, interactive customer support, tutoring, coaching, and multimodal Live API applications

Context 131K
Reasoning 7/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Preview; currently listed by Google with limited access for users who have actively used Gemini 2.5 models; no shutdown date announced

Input $0.50 per 1M text tokens; $3.00 per 1M audio or video tokens
Output $2.00 per 1M text tokens; $12.00 per 1M audio tokens

Low-latency controllable text-to-speech, voice assistants, narration, read-aloud features, and multi-speaker audio generation

Context 8K
Reasoning 2/10
Speed 8/10
Outputs
Speech
Status

Preview; currently available with limited access conditions

Input $0.50 per 1M input text tokens; batch: $0.25 per 1M input text tokens
Output $10.00 per 1M output audio tokens; batch: $5.00 per 1M output audio tokens
Google DeepMind
Gemini 2.5

Gemini 2.5 Pro

Advanced coding, complex reasoning, mathematics, STEM analysis, long documents, large codebases, multimodal analysis, and tool-using agents.

Context 1.05M
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Stable and generally available through the Gemini API; access is currently limited to users who have actively used Gemini 2.5 models. Google states that the model is not deprecated and will continue to be served until further notice.

Input $1.25 per 1M tokens for prompts up to 200K tokens; $2.50 per 1M tokens for prompts over 200K tokens. Batch/Flex rates are $0.625/$1.25 respectively.
Output $10.00 per 1M tokens for prompts up to 200K tokens; $15.00 per 1M tokens for prompts over 200K tokens. Batch/Flex rates are $5.00/$7.50 respectively. Prices include thinking tokens.

High-fidelity single-speaker and multi-speaker narration, audiobooks, podcasts, professional voiceovers, and scripted creative audio

Context 8K
Reasoning 1/10
Speed 6/10
Outputs
Speech
Status

Preview; currently listed and accessible through the Gemini API, with restricted access to Gemini 2.5 models

Input $1.00 per 1M text tokens standard; $0.50 per 1M text tokens batch
Output $20.00 per 1M audio tokens standard; $10.00 per 1M audio tokens batch
Google DeepMind
Gemini 2.5 Flash Image

Nano Banana

Fast image generation, conversational image editing, image transformation, and high-volume visual workflows

Context 66K
Reasoning 3/10
Speed 8/10
Outputs
Text Image
Capabilities
Image input Multimodal input
Status

Deprecated; scheduled to shut down on 2026-10-02

Input $0.30 per 1 million text or image input tokens; Batch/Flex $0.15; Priority $0.54
Output $0.039 per generated image; Batch/Flex $0.0195; Priority $0.0702

Agentic workflows, everyday coding, reasoning and planning, multimodal analysis, long-context document work, and cost-sensitive tool-using applications

Context 1.05M
Reasoning 9/10
Speed 9/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Preview

Input $0.50 per 1 million input tokens
Output $3 per 1 million output tokens

Complex reasoning, advanced software engineering, long-context multimodal analysis, and agentic workflows requiring reliable tool use

Context 1.05M
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Preview

Input $2.00 per 1M tokens for prompts up to 200K tokens; $4.00 per 1M tokens for prompts over 200K tokens
Output $12.00 per 1M tokens for prompts up to 200K tokens; $18.00 per 1M tokens for prompts over 200K tokens, including thinking tokens
Google DeepMind General Purpose

Fast multimodal applications, coding assistants, long-context document and video analysis, tool-using agents, enterprise workflows, and rapid agentic execution

Context 1.05M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Stable; previous-generation Flash model; currently accessible; no shutdown date announced

Input $1.50 per 1 million input tokens
Output $7.50 per 1 million output tokens
Google DeepMind General Purpose

Coding, software engineering, tool-using agents, long-context multimodal analysis, structured extraction and high-throughput enterprise workflows

Context 1.05M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Generally available; previous-generation Flash model, currently supported

Input $0.75 per 1M tokens through 2026-12-31; $1.50 per 1M tokens starting 2027-01-01
Output $3.75 per 1M tokens, including thinking tokens, through 2026-12-31; $7.50 per 1M tokens starting 2027-01-01

Long-horizon software engineering, autonomous agents, multimodal document workflows, enterprise knowledge work, and tool-using applications.

Context 1.05M
Reasoning 9/10
Speed 9/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Generally available

Input $0.75 per 1 million tokens through December 31, 2026; $1.50 per 1 million tokens from January 1, 2027. Batch/Flex introductory input price: $0.375 per 1 million tokens.
Output $3.75 per 1 million tokens, including thinking tokens, through December 31, 2026; $7.50 per 1 million tokens from January 1, 2027. Batch/Flex introductory output price: $1.875 per 1 million tokens.
Google DeepMind Lightweight

High-volume translation, classification, extraction, summarization, document processing, and lightweight tool-using agent workflows

Context 1.05M
Reasoning 7/10
Speed 10/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Deprecated; generally available and accessible until scheduled shutdown on May 7, 2027

Input $0.25 per 1M text/image/video tokens; $0.50 per 1M audio tokens; batch/flex: $0.125 per 1M text/image/video tokens and $0.25 per 1M audio tokens
Output $1.50 per 1M tokens standard; $0.75 per 1M tokens for batch/flex inference

Low-latency voice agents, real-time dialogue, multimodal live sessions, and interactive audio applications

Context 131K
Reasoning 7/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Legacy preview; currently accessible; Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or $0.005 per audio minute; $1.00 per 1M image/video tokens or $0.002 per image/video minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or $0.018 per audio minute
Google DeepMind
Gemini 3.1 Flash Audio

Gemini 3.1 Flash TTS

Controllable expressive speech, narration, accessibility, scripted audio, and multi-speaker TTS prototypes

Context 8K
Reasoning 2/10
Speed 8/10
Outputs
Speech
Capabilities
Streaming
Status

Legacy preview; currently accessible; no shutdown date announced

Input $1.00 per 1 million text tokens standard; $0.50 per 1 million text tokens batch
Output $20.00 per 1 million audio tokens standard; $10.00 per 1 million audio tokens batch
Google DeepMind
Gemini 3.1 Flash Lite Image

Nano Banana 2 Lite

Fast, low-cost 1K image generation and editing, rapid visual prototyping, interactive applications, high-volume image variations, storyboarding, and lightweight creative workflows

Context 66K
Reasoning 3/10
Speed 10/10
Outputs
Text Image
Capabilities
Image input Video input Multimodal input
Status

Generally available; scheduled retirement June 28, 2027 or later

Input $0.25 per 1 million input tokens for text, image, and video input; batch pricing $0.125 per 1 million input tokens
Output $1.50 per 1 million text/thought output tokens and $30 per 1 million image output tokens; approximately $0.0336 per 1K image; batch image output pricing $15 per 1 million image tokens
Google DeepMind General Purpose

Agentic workflows, coding agents, long-context multimodal analysis, tool use, and scaled production applications

Context 1.05M
Reasoning 9/10
Speed 8/10
Outputs
Text Actions
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Generally available; stable

Input $1.50 per 1M tokens standard; $0.75 per 1M tokens Batch and Flex; $2.70 per 1M tokens Priority
Output $9.00 per 1M tokens standard; $4.50 per 1M tokens Batch and Flex; $16.20 per 1M tokens Priority
Google DeepMind Lightweight

High-volume, latency-sensitive agentic workflows, document parsing, translation, classification, data extraction, search-backed applications, and multimodal sub-agents.

Context 1.05M
Reasoning 8/10
Speed 10/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

General availability

Input $0.30 per 1 million tokens for standard text, image, video, and audio input; $0.15 per 1 million input tokens for Batch and Flex; $0.54 per 1 million input tokens for Priority.
Output $2.50 per 1 million output tokens for standard usage; $1.25 per 1 million output tokens for Batch and Flex; $4.50 per 1 million output tokens for Priority.

Low-latency, real-time speech-to-speech translation for calls, meetings, travel, customer support, and multilingual voice applications

Context 131K
Reasoning 2/10
Speed 10/10
Outputs
Text Speech
Capabilities
Audio input Streaming
Status

Preview

Input $3.50 per 1 million input audio tokens, approximately $0.0053 per minute
Output $21.00 per 1 million output audio tokens, approximately $0.0315 per minute
Google DeepMind
Gemini 3.5 Audio

Gemini 3.5 Transcribe

Pre-recorded audio transcription, multilingual speech recognition, speaker-labeled transcripts, timestamped transcripts, smart dictation, and domain-specific vocabulary

Context 96K
Reasoning 2/10
Speed 9/10
Outputs
Text
Capabilities
Audio input Multimodal input
Status

Generally available (GA)

Input $2.00 per 1M audio input tokens, approximately $0.003 per minute; free tier available
Output $12.00 per 1M text output tokens, approximately $0.002 per minute; free tier available

High-volume text-to-speech production, low-latency voice-agent cascades, read-aloud applications, voice replication, and everyday single-speaker speech

Context 8K
Reasoning 1/10
Speed 9/10
Outputs
Speech
Capabilities
Streaming
Status

Generally available

Input $0.50 per 1 million text tokens through December 31, 2026; $1.00 per 1 million text tokens starting January 1, 2027. Batch and Flex input: $0.25 through December 31, 2026; $0.50 starting January 1, 2027.
Output $6.00 per 1 million audio tokens through December 31, 2026; $12.00 per 1 million audio tokens starting January 1, 2027. Standard equivalent: approximately $0.0015 per 10 seconds of audio through December 31, 2026.

Studio-quality narration, audiobooks, expressive voice acting, complex multi-speaker dialogue, regional accents, difficult pronunciations, long-form narration, voice design, and voice replication.

Context 8K
Reasoning 1/10
Speed 8/10
Outputs
Speech
Capabilities
Streaming
Status

Generally available (GA); no shutdown date announced

Input $0.50 per 1 million text input tokens through December 31, 2026; $1.00 per 1 million text input tokens from January 1, 2027
Output $9.00 per 1 million audio output tokens through December 31, 2026; $18.00 per 1 million audio output tokens from January 1, 2027. Equivalent to $0.00225 per 10 seconds of audio through December 31, 2026 and $0.0045 per 10 seconds thereafter.

Low-latency voice agents, real-time audio-to-audio dialogue, multimodal assistants, and interactive tool-using applications

Context 131K
Reasoning 7/10
Speed 10/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Stable; generally available

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or approximately $0.005 per minute; $1.00 per 1M image/video tokens or approximately $0.002 per minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or approximately $0.018 per minute

Complex real-time voice agents, multi-step problem solving, asynchronous tool workflows, technical support, travel coordination, and spoken STEM or coding tutoring

Context 131K
Reasoning 9/10
Speed 7/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Stable; generally available

Input $0.75 per 1M text tokens; $3.00 per 1M audio tokens or $0.005 per audio minute; $1.00 per 1M image/video tokens or $0.002 per minute
Output $4.50 per 1M text tokens; $12.00 per 1M audio tokens or $0.018 per audio minute
Google DeepMind
Gemini Deep Research

Gemini Deep Research

Autonomous market research, due diligence, literature reviews, competitive analysis, source-heavy investigations, and cited research reports

Context 1.05M
Reasoning 8/10
Speed 8/10
Outputs
Text Image
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Preview

Input Variable pay-as-you-go pricing based on underlying model usage and tools; no fixed per-input-token price is specified for the agent.
Output Variable pay-as-you-go pricing based on research depth and generated output; typical total cost is estimated at approximately $1–$3 per standard task.
Google DeepMind
Gemini Deep Research

Gemini Deep Research Max

Comprehensive market research, competitive analysis, due diligence, literature reviews, and source-rich investigative reports

Context 1.05M
Reasoning 10/10
Speed 3/10
Outputs
Text Image
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Preview; currently available through the Interactions API in the Gemini API and Google AI Studio

Input Pay-as-you-go based on underlying Gemini model inference and tool usage; no fixed model-specific input rate published
Output Pay-as-you-go based on underlying Gemini model inference and tool usage; no fixed model-specific output rate published
Google DeepMind
Gemini Embedding

Gemini Embedding

Text semantic search, RAG retrieval, document matching, classification, clustering, and recommendation systems

Context 2K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Status

Current, scheduled for shutdown on 2028-05-14

Input $0.00015 per 1,000 input tokens for online Vertex AI requests; $0.00012 per 1,000 input tokens for batch requests
Output No charge for embedding output
Google DeepMind
Gemini Embedding

Gemini Embedding 2

Cross-modal semantic search, multimodal RAG, vector retrieval, recommendations, classification, clustering, and indexing mixed text and media collections.

Context 8K
Reasoning 2/10
Speed 7/10
Outputs
Embeddings
Capabilities
Image input Audio input Video input Multimodal input
Status

Generally available

Input Standard paid tier: text $0.20 per 1M tokens; images $0.45 per 1M tokens or $0.00012 per image; audio $6.50 per 1M tokens or $0.00016 per second; video $12.00 per 1M tokens or $0.00079 per frame. Batch pricing is 50% lower.
Output No separate output-token price; the model returns embeddings. Output is included in the input-modality pricing structure.
Google DeepMind
Gemini Image

Nano Banana 2

Fast, high-volume image generation and editing, visual iteration, marketing assets, diagrams, infographics, localization, and applications requiring image-search grounding

Context 131K
Reasoning 7/10
Speed 9/10
Outputs
Text Image
Capabilities
Image input Video input Multimodal input Web search
Status

Generally available; the stable Gemini API model is gemini-3.1-flash-image

Input $0.50 per 1 million tokens for text/image input; batch input $0.25 per 1 million tokens
Output $3 per 1 million text/thinking tokens and $60 per 1 million image-output tokens; approximately $0.067 per 1K image, $0.101 per 2K image, and $0.151 per 4K image; batch image output $30 per 1 million image-output tokens
Google DeepMind
Gemini Image

Nano Banana Pro

Professional image generation and editing, complex compositions, product mockups, infographics, branded creative, multilingual localization, and high-fidelity visual prototyping

Context 66K
Reasoning 8/10
Speed 6/10
Outputs
Text Image
Capabilities
Image input Multimodal input Web search
Status

Current stable model

Input $2.00 per 1M text/image input tokens; approximately $0.0011 per input image
Output $12.00 per 1M text and thinking tokens; $120.00 per 1M image tokens, equivalent to approximately $0.134 per 1K/2K image and $0.24 per 4K image

Fast text-to-video, image-to-video, conversational video editing, video extension, interpolation, marketing content, and short-form cinematic production

Context 1.05M
Reasoning 5/10
Speed 9/10
Outputs
Video Audio
Capabilities
Image input Video input Multimodal input Streaming
Status

Current; stable model available as gemini-omni-1.1-flash, with gemini-omni-flash-preview also documented

Input $1.50 per 1M input tokens for text, image, video, and audio inputs under standard pricing
Output $17.50 per 1M video output tokens; $9.00 per 1M text output tokens where applicable
Google DeepMind
Gemini Robotics ER

Gemini Robotics ER 2

High-level robot planning, spatial reasoning, video progress tracking, tool orchestration, and multi-robot collaboration

Context 131K
Reasoning 8/10
Speed 4/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Public preview

Input $1.00 per 1 million tokens for the standard preview endpoint through December 31, 2026; $2.00 per 1 million tokens starting January 1, 2027
Output $5.00 per 1 million tokens, including reasoning tokens, through December 31, 2026; $10.00 per 1 million tokens starting January 1, 2027

Low-latency robotic agents, continuous audio/video monitoring, function-based robot orchestration, warehouse workflows, and multi-robot coordination

Context 131K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Public preview

Input $1.00 per 1M tokens through December 31, 2026; $2.00 per 1M tokens from January 1, 2027
Output $5.00 per 1M tokens through December 31, 2026; $10.00 per 1M tokens from January 1, 2027

Full-length AI-generated songs, vocal music, instrumental arrangements, songwriting experiments, soundtracks, and image-inspired music creation

Context 131K
Reasoning 1/10
Speed 7/10
Outputs
Text Music
Capabilities
Image input Multimodal input
Status

Generally available

Input $0.08 per full song; no free tier
Output $0.08 per full song

Interactive instrumental music generation, live musical improvisation, prompt-driven DJ tools, MIDI-controlled experiences, and real-time creative audio applications

Reasoning 1/10
Speed 9/10
Outputs
Music
Capabilities
Streaming
Status

Experimental and currently documented through the Gemini API as lyria-realtime-exp

Input Not publicly listed for this model
Output Not publicly listed for this model
Google DeepMind
Veo 3.1

Veo 3.1

Cinematic text-to-video and image-to-video generation, short-form storytelling, storyboarding, advertising concepts, visual effects exploration, and creative previsualization with synchronized audio

Context 1K
Reasoning 1/10
Speed 5/10
Outputs
Video Audio
Capabilities
Image input Video input Multimodal input
Status

Preview; currently accessible through Google APIs and Google products

Input $0.40 per second for standard Veo 3.1 video with audio at 720p or 1080p; $0.60 per second at 4K
Output Generated video with native audio; pricing is charged per generated video second
Google DeepMind Video Generation

Fast, high-volume video generation, creative iteration, social content, advertising concepts, and automated production workflows

Context 1K
Speed 9/10
Outputs
Video Audio
Capabilities
Image input Video input Multimodal input
Status

Generally available on Vertex AI; preview model on the Gemini API

Input $0.10 per video second at 720p; $0.12 per video second at 1080p; $0.30 per video second at 4K where supported
Output Video with natively generated audio, priced per generated video second

High-volume text-to-video and image-to-video generation, rapid creative iteration, social content, advertising variations, and cost-sensitive production workflows

Reasoning 1/10
Speed 8/10
Outputs
Video Audio
Capabilities
Image input Multimodal input
Status

Preview; currently available through the Gemini API and Google Cloud Vertex AI

Input $0.05 per second for 720p video with audio; $0.08 per second for 1080p video with audio
Output Video with native audio, priced per generated-video second
IBM watsonx General Purpose

Japanese text generation, summarization, classification, extraction, question answering, and Japanese-English translation in legacy IBM enterprise deployments

Context 4K
Reasoning 4/10
Speed 7/10
Outputs
Text
Status

Deprecated; withdrawn from IBM Software Hub 5.2.0 and removed from watsonx.ai in 2.2.0

IBM watsonx General Purpose

Multilingual enterprise question answering, retrieval-augmented generation, summarization, extraction, classification, and text generation in English, German, Spanish, French, and Portuguese

Context 8K
Reasoning 4/10
Speed 3/10
Outputs
Text
Status

Retired; deprecated January 15, 2025 and withdrawn from standard watsonx.ai availability April 16, 2025

IBM watsonx General Purpose

Fine-tuning, domain adaptation, text generation, summarization, extraction, classification, and question answering

Context 4K
Reasoning 4/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available for deployment on demand in IBM watsonx.ai; also available as an Apache 2.0 open-weight model through IBM's Hugging Face organization.

Input $0.0006 per 1,000 tokens
Output $0.0006 per 1,000 tokens
IBM watsonx General Purpose

Fine-tuning, long-context document processing, enterprise text classification, extraction, summarization, question answering, and self-managed deployment

Context 131K
Reasoning 5/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available; base model intended for fine-tuning and dedicated deployment

IBM watsonx General Purpose

Self-hosted enterprise assistants, long-context document analysis, RAG, summarization, extraction, multilingual text workflows, and function calling

Context 131K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available as an open-weight model; superseded by Granite-3.3-8B-Instruct for newer deployments

IBM watsonx Reasoning

Long-context enterprise RAG, document and meeting summarization, information extraction, multilingual dialogue, code-related tasks, function calling, and self-hosted deployments

Context 131K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available; legacy in some IBM watsonx catalogs

IBM watsonx General Purpose

Local or dedicated enterprise assistants, RAG, long-document summarization, multilingual text workflows, coding assistance, reasoning, and resource-conscious inference

Context 131K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available; open-weight model with local deployment and dedicated IBM watsonx deployment options

IBM watsonx General Purpose

Self-hosted enterprise assistants, long-context RAG, summarization, multilingual text tasks, coding assistance, function calling, and cost-sensitive deployments.

Context 131K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Retired from IBM watsonx.ai deploy-on-demand service on 2026-02-22; open-weight checkpoint remains available for self-hosted and compatible third-party deployments.

Input $0.0002 per 1,000 tokens, historical IBM watsonx Developer Hub listing; no current IBM-hosted price verified after withdrawal.
Output $0.0002 per 1,000 tokens, historical IBM watsonx Developer Hub listing; no current IBM-hosted price verified after withdrawal.
IBM watsonx General Purpose

Enterprise RAG, multi-tool agents, function calling, customer-support automation, multilingual instruction following, and long-context workloads

Context 131K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning
Status

Available; open-weight instruct model

Input $0.0000636 per 1,000 tokens on IBM watsonx.ai multitenant hardware
Output $0.000265 per 1,000 tokens on IBM watsonx.ai multitenant hardware
IBM watsonx Lightweight

Efficient enterprise assistants, multilingual text generation, RAG, classification, extraction, summarization, coding assistance, fill-in-the-middle completion, structured JSON, and tool-calling workflows

Context 128K
Reasoning 5/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current; open-weight instruct model; available for download and listed for deploy-on-demand use in IBM watsonx.ai

IBM watsonx Lightweight
IBM watsonx
Granite 4.1

Granite 4.1 3B

Efficient local or private deployment, multilingual enterprise text processing, RAG, summarization, extraction, coding assistance, function calling, and lightweight AI assistants.

Context 131K
Reasoning 5/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Currently available open-weight model; earlier-generation Granite model superseded by the newer Granite 4.2 family, with no verified deprecation or shutdown date.

Input No official hosted API token price published; downloadable weights are available under Apache 2.0.
Output No official hosted API token price published; downloadable weights are available under Apache 2.0.
IBM watsonx General Purpose
IBM watsonx
Granite 4.1

Granite 4.1 8B

Self-hosted enterprise assistants, multilingual text generation, RAG, coding assistance, structured extraction, and tool-calling agents

Context 131K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight instruction model

Input Not available for official per-token hosted pricing
Output Not available for official per-token hosted pricing
IBM watsonx General Purpose
IBM watsonx
Granite 4.1

Granite 4.1 30B

Self-hosted enterprise assistants, long-context RAG, multilingual applications, coding, structured extraction, and tool-calling agents

Context 131K
Reasoning 7/10
Speed 5/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight instruct model; publicly available

IBM watsonx Reasoning
IBM watsonx
Granite 4.2

Granite 4.2 3B

Efficient reasoning, coding assistance, tool calling, multilingual dialogue, local deployment, and lightweight enterprise agents

Context 131K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning
Status

Current; publicly available open-weight model

IBM watsonx Reasoning
IBM watsonx
Granite 4.2

Granite 4.2 8B

Local or self-hosted reasoning, coding assistants, tool calling, multilingual dialogue, retrieval-augmented generation, and agentic workflows

Context 131K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current; open-weight and downloadable

Input No official hosted API price; downloadable Apache 2.0 weights
Output No official hosted API price; inference infrastructure costs depend on deployment
IBM watsonx Reasoning
IBM watsonx
Granite 4.2

Granite 4.2 30B

Self-hosted enterprise reasoning, coding agents, multilingual applications, tool calling, long-context workflows, and organizations requiring Apache 2.0 licensing

Context 131K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Fine-tuning
Status

Current; open-weight and available for download

Input No official IBM hosted API price; model weights are available for self-hosted deployment
Output No official IBM hosted API price; model weights are available for self-hosted deployment
IBM watsonx General Purpose
IBM watsonx
Granite 7B

Granite-7B-Lab

Self-hosted English text generation, summarization, extraction, classification, experimentation with LAB-aligned open-weight models, and resource-conscious deployments

Context 8K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Withdrawn from IBM watsonx.ai on 2025-01-07; open-weight checkpoint remains available for self-hosted deployment

IBM watsonx General Purpose
IBM watsonx
Granite 13B V2

granite-13b-chat-v2

English enterprise chat, retrieval-augmented generation, question answering, summarization, extraction, and classification

Context 8K
Reasoning 4/10
Speed 6/10
Outputs
Text
Status

Withdrawn; access ended January 19, 2025

Input $0.0006 per 1,000 input tokens, historical watsonx.ai rate
Output $0.0006 per 1,000 output tokens, historical watsonx.ai rate

Self-hosted coding assistants, code generation, code explanation, code conversion, repository-scale prompts, and historical or reproducible research

Context 128K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Retired; withdrawn from IBM watsonx.ai on 2025-07-17

Input $0.0006 per 1,000 input tokens historically on IBM watsonx.ai
Output $0.0006 per 1,000 output tokens historically on IBM watsonx.ai

Schema linking and relevant-table or relevant-column selection in text-to-SQL pipelines

Context 8K
Reasoning 3/10
Speed 3/10
Outputs
Text
Status

Available; deploy on demand only

Input Not available as pay-as-you-go token pricing; deploy-on-demand infrastructure pricing applies
Output Not available as pay-as-you-go token pricing; deploy-on-demand infrastructure pricing applies

Natural-language-to-SQL generation over structured databases, analytics assistants, and the SQL-generation stage of text-to-SQL pipelines.

Context 8K
Reasoning 5/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available; deploy-on-demand model in IBM watsonx.ai

Input Not publicly listed as a standalone token price; IBM lists the model as deploy-on-demand.
Output Not publicly listed as a standalone token price; IBM lists the model as deploy-on-demand.

Self-hosted coding assistants, code generation, code conversion, code explanation, and programming experiments where an Apache 2.0 open-weight model is preferred.

Context 8K
Reasoning 5/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as an open-weight model and documented for IBM watsonx.ai deploy-on-demand use; deprecated and withdrawn from the watsonx.ai multitenant offering according to IBM's 2025 lifecycle notice.

Self-hosted code generation, code explanation, code repair, code conversion, coding assistants, and legacy Granite Code compatibility

Context 8K
Reasoning 6/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Legacy/deprecated; available in some IBM watsonx environments and retained for historical or scientific use

IBM watsonx Embedding

Low-cost multilingual semantic search, cross-lingual retrieval, vector databases, similarity matching, and RAG pipelines

Context 512
Reasoning 1/10
Speed 9/10
Outputs
Embeddings
Capabilities
Fine-tuning
Status

Retired from IBM watsonx.ai; open-weight checkpoint remains available through IBM's Hugging Face organization

Input No official hosted API price verified; open-weight model intended for local or third-party deployment
Output Not applicable to token generation; produces 384-dimensional embeddings
IBM watsonx Embedding

English semantic search, vector retrieval, retrieval-augmented generation, document similarity, clustering, and enterprise information retrieval

Context 8K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Status

Current open-weight model

Input No official hosted API price; downloadable weights are available under the Apache 2.0 license
Output No official hosted API price; downloadable weights are available under the Apache 2.0 license

Canopy-height mapping, forest monitoring, vegetation analysis, carbon-cycle research, ecological assessment, and remote-sensing experimentation.

Reasoning 1/10
Speed 5/10
Capabilities
Image input Fine-tuning
Status

Available as downloadable open-weight model; not deployed by an inference provider; IBM repository disclosure states that the project is not maintained as an IBM product.

IBM watsonx
Granite Geospatial

granite-geospatial-uki

Remote-sensing feature extraction, Earth-observation transfer learning, and regional flood-segmentation workflows using multispectral and SAR satellite imagery

Reasoning 1/10
Speed 5/10
Capabilities
Image input Multimodal input Fine-tuning
Status

Available as an open-weight research model; no hosted inference provider is currently listed

Input No official hosted API pricing; open-weight model
Output No official hosted API pricing; open-weight model
IBM watsonx
Granite Guardian

Granite Guardian 3.8B

Prompt and response safety classification, jailbreak detection, RAG groundedness and relevance checks, hallucination evaluation, and enterprise AI guardrails

Context 131K
Reasoning 5/10
Speed 7/10
Outputs
Text
Status

Deprecated in IBM watsonx.ai documentation; open-weight model and official model materials remain available

Input $0.0002 per 1,000 input tokens in IBM watsonx.ai documentation
Output $0.0002 per 1,000 output tokens in IBM watsonx.ai documentation
IBM watsonx
Granite Guardian

Granite Guardian 4.1 8B

AI safety guardrails, jailbreak detection, RAG groundedness and relevance checks, function-call hallucination detection, custom criteria evaluation, and best-of-N response ranking

Context 8K
Reasoning 6/10
Speed 7/10
Outputs
Text
Status

Current and available as an open-weight model

Input Not applicable; no official hosted token price found for the downloadable model
Output Not applicable; no official hosted token price found for the downloadable model
IBM watsonx Speech Recognition
IBM watsonx
Granite Speech 4.1

Granite Speech 4.1 2B NAR

High-throughput multilingual speech transcription where low inference latency is more important than maximum recognition accuracy.

Context 4K
Reasoning 2/10
Speed 10/10
Outputs
Text
Capabilities
Audio input
Status

Current; open-weight; Apache 2.0

Input No official hosted API pricing; model weights are available for self-hosted use.
Output No official hosted API pricing; model weights are available for self-hosted use.

Fast local English speech-to-text transcription, edge-device ASR, research, browser demos, and high-throughput noncommercial batch processing

Reasoning 1/10
Speed 10/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current; research and noncommercial use only

Input No official hosted API price; downloadable model weights
Output No official hosted API price; downloadable model weights
IBM watsonx
Granite Time Series

granite-ttm-512-96-r2

Lightweight multivariate forecasting at minute- and hour-level resolutions, especially when 512 historical observations and a 96-point forecast horizon are appropriate.

Context 512
Reasoning 0/10
Speed 9/10
Capabilities
Fine-tuning
Status

Available

Input $0.13 per 1,000 input data points
Output $0.38 per 1,000 output data points
IBM watsonx
Granite Time Series TTM

Granite-TTM-R3

High-throughput multivariate time-series forecasting, zero-shot or few-shot forecasting, probabilistic prediction, exogenous-variable forecasting, and CPU-friendly production deployment.

Reasoning 0/10
Speed 9/10
Capabilities
Fine-tuning
Status

Current open-weight model family; available through the IBM Granite Hugging Face repository. No hosted inference provider deployment was listed on the reviewed model page.

Efficient multivariate forecasting for regularly sampled energy, traffic, manufacturing, network, sales, and sensor data

Context 1K
Reasoning 1/10
Speed 9/10
Capabilities
Fine-tuning
Status

Available through IBM watsonx.ai; downloadable model branch

Input watsonx.ai API pricing class 14
Output watsonx.ai API pricing class 15

Multivariate forecasting when at least 1,536 historical observations per channel are available, especially demand, traffic, electricity, manufacturing, finance, and other minute- or hour-level forecasting tasks.

Context 2K
Reasoning 0/10
Speed 9/10
Capabilities
Fine-tuning
Status

Available

Input $0.00013 per 1,000 data points
Output $0.00038 per 1,000 data points
IBM watsonx Multimodal

Chart extraction, table parsing, semantic key-value extraction, visual document processing, and local enterprise RAG pipelines

Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Current and downloadable; newer Granite Vision 4.1 4B version available

Input No official hosted API price found; open-weight model intended for self-hosted deployment
Output No official hosted API price found; open-weight model intended for self-hosted deployment
IBM watsonx Multimodal

Enterprise document understanding, chart and table extraction, OCR-oriented image analysis, visual question answering, and multimodal RAG

Context 131K
Reasoning 5/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Available; legacy relative to Granite Vision 4.0 3B Vision

IBM watsonx Multimodal
IBM watsonx
Granite Vision 4.1

Granite Vision 4.1 4B

Structured extraction from charts, tables, invoices, forms, and enterprise document images

Reasoning 3/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning Streaming
Status

Current open-weight model

Input No official hosted API price specified; downloadable open-weight model
Output No official hosted API price specified; downloadable open-weight model

English semantic search, dense retrieval, vector indexing, duplicate-question matching, and lightweight retrieval-augmented generation

Context 512
Speed 8/10
Outputs
Embeddings
Status

Deprecated; withdrawn in Dallas on 2026-09-08 and scheduled for withdrawal in other listed regions on 2027-01-12.

Input USD 0.0001 per 1,000 input tokens on watsonx.ai
Output Not separately priced; the model returns embedding vectors rather than generated output tokens
LG AI Research General Purpose

Bilingual English-Korean chat, instruction following, Korean-language applications, local inference, and open-weight LLM research

Context 4K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Streaming
Status

Available open-weight model; older EXAONE generation and superseded by newer EXAONE releases

Input No official hosted API pricing published; self-hosted weights
Output No official hosted API pricing published; self-hosted weights
LG AI Research Lightweight

Lightweight local text generation, English-Korean assistants, summarization, rewriting, classification, and deployment on resource-constrained devices

Context 33K
Reasoning 5/10
Speed 8/10
Outputs
Text
Status

Current downloadable open-weight model; research use permitted, commercial use requires a separate license

Input No official hosted API pricing; downloadable weights
Output No official hosted API pricing; downloadable weights
LG AI Research
EXAONE 4.0

EXAONE-4.0-32B

Self-hosted Korean, English, and Spanish language tasks requiring a combination of general-purpose generation, reasoning, coding, long-context processing, or agentic tool use

Context 131K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Available open-weight model; no official deprecation or shutdown date found

Input No official hosted API pricing found; intended primarily for self-hosted deployment
Output No official hosted API pricing found; intended primarily for self-hosted deployment
LG AI Research Multimodal
LG AI Research
EXAONE 4.5

EXAONE-4.5-33B

Open-weight multimodal reasoning, Korean-language tasks, document understanding, OCR, visual question answering, and self-hosted deployments

Context 262K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Current open-weight model

LG AI Research
EXAONE Deep

EXAONE-Deep-7.8B

Local research, mathematical reasoning, science problem solving, coding evaluation, Korean and English text generation, and self-hosted experimentation

Context 33K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current and accessible open-weight research release

LG AI Research
EXAONE Deep

EXAONE-Deep-32B

Mathematical reasoning, scientific problem solving, coding evaluation, and local research deployments

Context 33K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Streaming
Status

Current open-weight research model; downloadable and locally deployable

Input No official hosted API price; downloadable weights for local or third-party deployment
Output No official hosted API price; downloadable weights for local or third-party deployment

Research on EGFR mutation prediction, computational pathology, molecular subtyping, and whole-slide image biomarker analysis

Reasoning 0/10
Speed 0/10
Capabilities
Image input
Status

Available as a gated open-source research release; not deployed by a hosted inference provider

Input No official hosted API pricing published
Output No official hosted API pricing published
Meta AI Multimodal
Meta AI
DINOv3

DINOv3

Image embeddings, dense feature extraction, image retrieval, classification, segmentation, depth estimation, object discovery, video tracking pipelines, and geospatial computer vision

Reasoning 0/10
Speed 7/10
Outputs
Embeddings
Capabilities
Image input Fine-tuning
Status

Current; downloadable open-weight research model suite

Meta AI General Purpose
Meta AI
Llama 3.1

Llama 3.1 8B

Local inference, fine-tuning, private deployment, text generation, retrieval-augmented generation, research, and cost-sensitive applications

Context 131K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current open-weight static model; downloadable and usable through compatible local or hosted inference deployments

Meta AI General Purpose

High-quality open-weight research, multilingual applications, coding, reasoning, synthetic-data generation, model distillation and self-hosted deployments with substantial infrastructure

Context 131K
Reasoning 9/10
Speed 3/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available as an open-weight model; older generation superseded by newer Llama releases

Meta AI Lightweight
Meta AI
Llama 3.2

Llama 3.2 1B

Private local inference, mobile and edge assistants, summarization, rewriting, retrieval-supported generation, and lightweight multilingual applications

Context 128K
Reasoning 3/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available downloadable open-weight model; static checkpoint

Input No official Meta-hosted API price; downloadable weights are available under the Llama 3.2 Community License
Output No official Meta-hosted API price; deployment and inference costs depend on hardware or third-party hosting
Meta AI Lightweight
Meta AI
Llama 3.2

Llama 3.2 3B

Private local inference, multilingual text generation, edge applications, model adaptation, and fine-tuning

Context 128K
Reasoning 5/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning
Status

Available; static pretrained open-weight model

Input No official Meta-hosted API price; downloadable weights
Output No official Meta-hosted API price; downloadable weights
Meta AI Multimodal

Visual question answering, image reasoning, chart and document understanding, image captioning, multimodal research, and self-hosted or partner-hosted AI applications.

Context 128K
Reasoning 8/10
Speed 4/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Fine-tuning
Status

Available open-weight model; static model trained on an offline dataset

Meta AI Multimodal

Self-hosted visual question answering, image captioning, document analysis, visual reasoning, and multimodal assistants

Context 128K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Fine-tuning
Status

Available open-weight static model

Meta AI General Purpose

Self-hosted or hosted multilingual chat, coding assistance, long-context text generation, tool calling, synthetic data, and applications requiring open model weights

Context 128K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Fine-tuning
Status

Available open-weight model; static model trained on an offline dataset

Input No universal Meta-hosted per-token price; downloadable weights and third-party hosting options are available
Output No universal Meta-hosted per-token price; downloadable weights and third-party hosting options are available
Meta AI Multimodal

Open-weight multimodal assistants, image understanding, visual question answering, coding, multilingual applications, creative writing and long-context text processing.

Context 1M
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Fine-tuning Streaming
Status

Available; open-weight static checkpoint

Meta AI Multimodal

Long-context document and code analysis, visual question answering, multimodal assistants, multilingual applications, self-hosted inference, and customized deployments

Context 10M
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Fine-tuning Streaming
Status

Available; open-weight model

Meta AI Other
Meta AI
Llama Guard

Llama Guard 4

Text and image moderation for prompts and generated responses in generative AI systems

Context 8K
Reasoning 2/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Current and available as an open-weight model

Meta AI Other
Meta AI
Llama Guard 3

Llama Guard 3-1B

Low-cost prompt and response safety classification, local moderation, mobile and edge deployments, and customizable LLM guardrails

Context 131K
Reasoning 3/10
Speed 9/10
Outputs
Text
Capabilities
Fine-tuning
Status

Current open-weight model; downloadable subject to access approval

Meta AI Other

Safety classification of mixed text-and-image prompts and text responses in multimodal LLM systems

Context 128K
Reasoning 4/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Available open-weight multimodal safety model; Meta's current model repositories continue to list it.

Meta AI Other
Meta AI
Llama Prompt Guard 2

Llama Prompt Guard 2 86M

Multilingual prompt-injection detection, jailbreak screening, agent security, and filtering untrusted text before it reaches an LLM

Context 512
Reasoning 1/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Current open-weight model; gated download access

Meta AI Reasoning
Meta AI
Muse Glimmer

Muse Glimmer

Local agents, long-running tool workflows, coding assistants, multimodal document and screenshot understanding, private on-device inference, and model customization

Context 131K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Fine-tuning Streaming
Status

Current open-weight model; self-hosted and available through selected third-party hosted inference providers

Input No official Meta-hosted API token price; open weights can be downloaded and self-hosted without per-token charges
Output No official Meta-hosted API token price; third-party hosted inference pricing varies by provider
Meta AI Multimodal
Meta AI
Muse Image

Muse Image 1.0

Text-to-image generation, precise image editing, multi-image composition, anchored visual series, product imagery, creative assets, and grounded visual content.

Reasoning 8/10
Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input Tool use Web search
Status

Current and available through Meta Model API; also available in selected Meta AI consumer experiences

Output $0.01 per successfully generated image
Meta AI Reasoning
Meta AI
Muse Spark

Muse Spark 1.1

Agentic workflows, coding agents, computer-use automation, multimodal document and media analysis, long-context reasoning, tool orchestration, and web-grounded applications.

Context 1.05M
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current but superseded by Muse Spark 1.2 and Muse Spark 1.3; available on the Meta Model API Standard tier in public preview for US developers.

Input $1.25 per 1 million input tokens; cached input $0.15 per 1 million tokens
Output $4.25 per 1 million output tokens
Meta AI Coding
Meta AI
Muse Spark

Muse Spark 1.2

Long-horizon coding agents, repository-scale software engineering, multimodal code generation, debugging, refactoring, and tool-driven workflows

Context 1.05M
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Available; previous version, with Muse Spark 1.3 recommended for new work

Input $1.25 per 1 million input tokens; $0.15 per 1 million cached input tokens
Output $4.25 per 1 million output tokens
Meta AI Multimodal
Meta AI
Muse Spark

Muse Spark 1.3

Long-horizon coding agents, software engineering, browser and computer-use workflows, tool orchestration, large repositories, document analysis, and multimodal reasoning

Context 1.05M
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current; available through Meta Model API and Muse Code

Input $1.25 per 1M input tokens; $0.15 per 1M cached input tokens
Output $4.25 per 1M output tokens
Meta AI Other

Real-time speech-to-text, live captions, voice agents, meeting transcription, call intelligence, dictation, and speaker-aware transcription

Reasoning 1/10
Speed 9/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current and available through Meta Model API

Input $3.00 per 1,000 minutes ($0.18 per hour)
Meta AI Other
Meta AI
Omnilingual wav2vec 2.0

Omnilingual wav2vec 2.0

Multilingual speech representation learning, audio embeddings, low-resource language research, and custom downstream speech systems

Reasoning 1/10
Speed 5/10
Outputs
Embeddings
Capabilities
Audio input Fine-tuning
Status

Current; open-source research model family

Input Free to download and self-host; no official hosted API price published
Output Free to download and self-host; no official hosted API price published
Meta AI Multimodal

Cross-modal audio-video-text retrieval, audiovisual embeddings, sound-event understanding, media indexing, and multimodal perception systems.

Reasoning 2/10
Speed 7/10
Outputs
Embeddings
Capabilities
Audio input Video input Multimodal input
Status

Current and openly available

Input No official hosted API pricing; open checkpoints are available for self-managed use.
Output No official hosted API pricing; the model returns embeddings rather than generated media.
Meta AI Other

Single-image reconstruction of textured 3D objects from natural scenes, 3D computer-vision research, Gaussian-splat workflows, and rapid asset prototyping

Reasoning 2/10
Speed 3/10
Capabilities
Image input
Status

Available research release; gated model checkpoints

Meta AI Multimodal
Meta AI
SAM Audio

SAM Audio

Prompted audio separation, speech and noise isolation, instrument and vocal extraction, audiovisual sound segmentation, and audio-editing research

Reasoning 2/10
Speed 7/10
Outputs
Audio
Capabilities
Audio input Video input Multimodal input
Status

Current open research release; downloadable checkpoints and public demo available

Input No hosted API pricing published; downloadable checkpoints available under the SAM License
Output No hosted API pricing published; model outputs separated target and residual audio
Meta AI Multimodal

Multilingual automatic speech recognition, speech-to-text translation, text translation, text-to-speech translation, and speech-to-speech translation

Reasoning 1/10
Speed 5/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Fine-tuning
Status

Available open-weight research model; noncommercial research use

Input Not applicable; no official hosted API pricing published
Output Not applicable; no official hosted API pricing published
Meta AI Other
Meta AI
Segment Anything

SAM 3.1

Open-vocabulary object detection, pixel-level image segmentation, and multi-object video tracking

Reasoning 2/10
Speed 9/10
Outputs
Image
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Current and available; hosted through Meta Model API and available as released research checkpoints

Input $2.50 per 1,000 images; $0.20 per 1,000 video frames
Output Included in the image and video segmentation pricing; no separate output-token price
Meta AI Multimodal

Computational neuroscience, fMRI response prediction, brain encoding, multisensory research, and in-silico experiment design

Reasoning 0/10
Speed 0/10
Capabilities
Image input Audio input Video input Multimodal input Fine-tuning
Status

Open-weight research release

Microsoft Copilot
BiomedParse

MedImageParse

Text-guided biomedical image segmentation, annotation assistance, organ and tumor delineation, pathology-cell analysis, and research-oriented medical imaging pipelines

Reasoning 1/10
Speed 5/10
Outputs
Text Image
Capabilities
Image input Multimodal input Fine-tuning
Status

Current; open-weight research model and available for managed deployment through Microsoft Foundry

Input Not publicly listed as a per-inference model price; Microsoft Foundry deployments incur managed Azure compute charges
Output Not publicly listed as a per-inference model price; Microsoft Foundry deployments incur managed Azure compute charges

First-pass chest X-ray report drafting, structured findings extraction, radiologist workflow assistance, research evaluation, and institution-specific fine-tuning

Reasoning 2/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Limited preview; registration and eligibility approval required

Input $0.00218 per image for standard inference
Output Included in the per-image inference price; output is generated text findings

Local image captioning, object detection, phrase grounding, region description, OCR and lightweight computer-vision pipelines

Context 1K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Current open-weight model; downloadable from Hugging Face and usable for local or self-hosted inference

Local image captioning, object detection, visual grounding, OCR, region annotation and multi-task computer-vision pipelines.

Context 1K
Reasoning 2/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Current open-weight model; publicly accessible through Hugging Face for local or self-hosted inference.

Open-weight reasoning, mathematics, coding, research, and general text-generation applications requiring DeepSeek-R1-style reasoning with Microsoft post-training

Context 164K
Reasoning 8/10
Speed 3/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available as an open-weights model and through Microsoft Foundry hosted API

Microsoft Copilot
MAI-Image

MAI-Image-2.5

High-quality text-to-image generation, photorealistic imagery, product and marketing visuals, presentation graphics, and precise image-to-image editing.

Context 131K
Reasoning 6/10
Speed 6/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Preview

Input $5 per 1 million text input tokens; $8 per 1 million image input tokens
Output $47 per 1 million image output tokens
Microsoft Copilot
MAI-Image-2.5

MAI-Image-2.5-Flash

Fast, cost-conscious text-to-image generation, image editing, creative production, concept visualization, and high-volume image workflows

Context 32K
Reasoning 5/10
Speed 9/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Public preview; scheduled for retirement on 2026-10-01

Input $1.75 per 1M text input tokens; $1.75 per 1M image input tokens
Output $19.50 per 1M image output tokens
Microsoft Copilot
MAI-Image-2.5

MAI-Image-2.5-Pro

High-fidelity text-to-image generation, precise image editing, hero imagery, commercial and photorealistic creative work, accurate in-image typography, and visually dense scenes requiring consistent objects, characters, materials, and spatial relationship

Context 131K
Reasoning 7/10
Speed 4/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Public preview

Input $5 per 1 million text input tokens; $8 per 1 million image input tokens
Output $106 per 1 million image output tokens
Microsoft Copilot
MAI-Transcribe

MAI-Transcribe-1.5

Multilingual speech-to-text, captions, meeting transcription, accessibility, call analysis, content workflows, voice-agent audio understanding, and domain-specific terminology

Reasoning 1/10
Speed 9/10
Outputs
Text
Capabilities
Audio input
Status

Preview; currently accessible through Microsoft Foundry and Azure Speech

Input $0.36 per hour of audio
Microsoft Copilot
MAI-Transcribe

MAI-Transcribe-2

Multilingual audio transcription, meeting and contact-center records, captions, clinical notes, accessibility, media search, voice-agent evaluation, and domain-specific transcription with speaker labels and timestamps

Speed 9/10
Outputs
Text
Capabilities
Audio input
Status

Public preview

Input $0.10 per hour of audio; limited-time launch pricing through December 31, 2026
Output Included; no separate text-output charge documented
Microsoft Copilot
MAI-Voice

MAI-Voice-2

Expressive long-form narration, audiobooks, podcasts, educational content, voice-over, accessibility, and high-fidelity branded audio

Reasoning 1/10
Speed 7/10
Outputs
Speech
Capabilities
Audio input Multimodal input
Status

Public preview

Input $22 per 1 million characters
Output $22 per 1 million characters
Microsoft Copilot
MedImageParse

MedImageParse 3D

Text-prompted segmentation of complete CT or MRI volumes, organ and lesion delineation, volumetry, annotation assistance, and medical-imaging research

Reasoning 1/10
Speed 5/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Available through Microsoft Foundry classic managed compute; research and model-development use only

Microsoft Copilot General Purpose
Microsoft Copilot
Microsoft Foundry model-router

model-router

Enterprise applications with mixed-complexity workloads, model selection automation, agent workflows, cost optimization, and configurable quality-versus-latency trade-offs.

Context 200K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Current; latest documented version is 2025-11-18

Input Not publicly listed as a numeric rate; Microsoft’s pricing page currently shows $- and directs users to Azure pricing for applicable deployment details.
Output Not separately listed; billing depends on model-router and the underlying routed model pricing configuration.
Microsoft Copilot General Purpose

Local and private text generation, coding assistance, mathematics, reasoning, summarization, and latency-sensitive applications

Context 4K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Retired from Microsoft Foundry; open-weight repository remains available for independent deployment

Input $0.00017 per 1,000 tokens historically on Azure AI; hosted price no longer applicable after retirement
Output $0.00068 per 1,000 tokens historically on Azure AI; hosted price no longer applicable after retirement
Microsoft Copilot General Purpose

Long-context chat, retrieval-augmented generation, document summarization, coding, mathematics, reasoning, and self-hosted text-generation applications

Context 131K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Generally available in Microsoft Foundry; open-weight checkpoint available for download and self-hosting

Input $0.17 per 1 million input tokens in Microsoft Foundry
Output $0.68 per 1 million output tokens in Microsoft Foundry

Local or low-latency text generation, lightweight assistants, mathematics, coding, summarization, and resource-constrained deployments

Context 4K
Reasoning 6/10
Speed 9/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Retired from Microsoft Foundry on August 30, 2025

Input $0.00013 per 1,000 input tokens (historical Azure pricing)
Output $0.00052 per 1,000 output tokens (historical Azure pricing)

Long-document analysis, local assistants, code and math tasks, retrieval-augmented generation, private or offline inference, and resource-constrained deployments.

Context 131K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Retired from Microsoft Foundry on August 30, 2025; open-weight checkpoint remains downloadable and usable for self-hosted inference.

Input $0.00013 per 1,000 input tokens, historical Azure Models-as-a-Service pricing; hosted offering retired
Output $0.00052 per 1,000 output tokens, historical Azure Models-as-a-Service pricing; hosted offering retired

Local and private text generation, instruction following, code and mathematics assistance, summarization, extraction, and resource-constrained deployments

Context 8K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Retired from Microsoft Foundry on 2025-08-30; downloadable open-weight checkpoint remains available

Input $0.00015 per 1,000 input tokens (historical Azure Models as a Service pricing; hosted service retired)
Output $0.0006 per 1,000 output tokens (historical Azure Models as a Service pricing; hosted service retired)

Long-context text generation, document analysis, summarization, local assistants, coding support, mathematics, and cost-sensitive self-hosted applications

Context 131K
Reasoning 6/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Retired from Microsoft Foundry on 2025-08-30; open-weight checkpoint remains available for self-hosted deployment

Input USD 0.00015 per 1,000 input tokens historically on Azure Foundry; hosted Azure access is retired
Output USD 0.0006 per 1,000 output tokens historically on Azure Foundry; hosted Azure access is retired

Local multilingual chat, summarization, document analysis, coding assistance, long-context retrieval, and resource-constrained deployments

Context 131K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight model; also supported in Microsoft and third-party local or hosted inference environments

Microsoft Copilot General Purpose

Long-context multilingual assistants, coding, mathematics, reasoning, retrieval-augmented generation, and controlled local deployment.

Context 131K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Retired from Azure Foundry on August 30, 2025; downloadable open-weight checkpoint remains available for self-hosted deployment.

Input $0.16 per 1M input tokens historically on Azure; current hosted pricing unavailable after retirement.
Output $0.64 per 1M output tokens historically on Azure; current hosted pricing unavailable after retirement.
Microsoft Copilot General Purpose
Microsoft Copilot
Phi-4

Phi-4

Efficient local or hosted text generation, STEM reasoning, mathematics, coding assistance, technical question answering, and research on small language models

Context 16K
Reasoning 8/10
Speed 8/10
Outputs
Text
Status

Preview in Microsoft Foundry; open-weight model available under the MIT license

Efficient mathematical reasoning, math tutoring, automated assessment, lightweight reasoning agents, edge deployment, mobile applications, and latency-sensitive local inference

Context 66K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current open-weight model; available through Microsoft’s Hugging Face repository and documented Azure AI Foundry availability

Efficient local or cloud text generation, multilingual applications, mathematics, coding, reasoning, retrieval-augmented generation, and edge deployment

Context 131K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Generally available

Input $0.00075 per 1,000 tokens in Microsoft's February 2025 Azure announcement; current Foundry pricing may vary
Output $0.0003 per 1,000 tokens in Microsoft's February 2025 Azure announcement; current Foundry pricing may vary

Compact multimodal assistants, OCR, document and chart analysis, image question answering, speech recognition, speech translation, audio summarization, and private or local deployment

Context 131K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Audio input Multimodal input Fine-tuning Streaming
Status

Available; open-weight; listed in Microsoft Foundry and Hugging Face

Input Not publicly listed as a current numeric price; Azure pricing page displays $-
Output Not publicly listed as a current numeric price; Azure pricing page displays $-

Mathematical reasoning, coding, science, logic, algorithmic problem solving, research, and resource-constrained local deployments

Context 33K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Streaming
Status

Preview; open-weight and currently listed in Microsoft Foundry

Input No official per-token price published; open-weight model for self-hosted or separately priced managed deployment
Output No official per-token price published; open-weight model for self-hosted or separately priced managed deployment
MiniMax Multimodal

Short text-to-video and image-to-video clips, cinematic experiments, advertising concepts, social content, and scenes with complex motion.

Reasoning 1/10
Speed 5/10
Outputs
Video
Capabilities
Image input Multimodal input
Status

Legacy or superseded; no longer prominent in MiniMax's current first-party model catalog

Output Historical launch configurations were priced per generated video rather than by tokens; documented configurations included 512p/6s, 512p/10s, 768p/6s, 768p/10s, and 1080p/6s.
MiniMax Video Generation

Short cinematic videos, image animation, realistic human motion, stylized scenes, visual effects, advertising concepts, and social-media content.

Reasoning 2/10
Speed 6/10
Outputs
Video
Capabilities
Image input Multimodal input
Status

Legacy or superseded video model; still referenced in MiniMax consumer and platform offerings, while MiniMax H3 is the current primary video model in developer documentation.

Output $0.28 per 768P/6s clip; $0.56 per 768P/10s clip; $0.49 per 1080P/6s clip
MiniMax Other

Fast image-to-video generation, high-volume short-form content, social media clips, advertisements, and rapid creative iteration

Reasoning 1/10
Speed 9/10
Outputs
Video
Capabilities
Image input Multimodal input
Status

Legacy; current availability should be verified

Input $0.19 per 768p 6-second clip; $0.32 per 768p 10-second clip; $0.33 per 1080p 6-second clip
MiniMax Other
MiniMax
Hailuo Video

MiniMax T2V-01

Short text-driven video concepts, storyboards, and early Hailuo-style cinematic experiments

Reasoning 2/10
Speed 6/10
Outputs
Video
Status

Legacy and superseded; no longer listed in MiniMax's current primary video-generation catalog

MiniMax Other
MiniMax
Image-01

image-01

Prompt-based image generation, reference-guided variations, commercial visuals, and batch creative production

Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and accessible through MiniMax's image-generation console and API platform

MiniMax Other
MiniMax
MiniMax ASR

ASR 1.0

Multilingual audio transcription, meeting transcription, speaker-labeled transcripts, live captions, call analysis, and subtitle generation

Reasoning 1/10
Speed 7/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current public model

Input $0.38 per hour of processed audio
Output No separate output charge; billing is based on input audio duration
MiniMax Multimodal
MiniMax
MiniMax H3

H3 Max

Fast short-form text-to-video, image-to-video, reference-guided generation, and synchronized-audio production

Reasoning 2/10
Speed 10/10
Outputs
Video Audio
Capabilities
Image input Audio input Video input Multimodal input Streaming
Status

Current; commercially hosted by fal

Output $0.05 per second at 480p; $0.08 per second at 768p; $0.16 per second at 1080p
MiniMax Multimodal
MiniMax
MiniMax H3

MiniMax H3

Multimodal commercial video generation, reference-based editing, product and advertising content, short cinematic clips, and locally deployed 768p workflows

Reasoning 2/10
Speed 5/10
Outputs
Video Music
Capabilities
Image input Audio input Video input Multimodal input Fine-tuning
Status

Current; open-weight release and hosted API available

Input $0.08 per second for 768p video output; $0.13 per second for 2K video output; H3-Context-IR: $0.90 per million input tokens
Output $0.08 per second for 768p video output; $0.13 per second for 2K video output; H3-Regenerate-2K: $0.05 per second of regenerated output; H3-Context-IR: $3.60 per million output tokens
MiniMax Coding
MiniMax
MiniMax M2

MiniMax M2

Coding agents, multi-step tool workflows, long-context codebase analysis, research automation, and self-hosted experimentation

Context 197K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Legacy/open-weight model; current hosted availability and pricing are not confirmed in MiniMax's latest public model catalog

Input $0.30 per 1 million input tokens at launch; current pricing unverified
Output $1.20 per 1 million output tokens at launch; current pricing unverified
MiniMax Coding
MiniMax
MiniMax M2

MiniMax M2.1

Multilingual software engineering, coding agents, tool-using workflows, application development, and cost-sensitive automation

Context 205K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Legacy but currently available

Input $0.30 per 1 million tokens; prompt-cache read $0.03 per 1 million tokens; prompt-cache write $0.375 per 1 million tokens
Output $1.20 per 1 million tokens
MiniMax Reasoning
MiniMax
MiniMax M2

MiniMax M2.5

Coding agents, software engineering, search and browser agents, tool-calling workflows, office automation, and long-context technical work

Context 205K
Reasoning 9/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Web search Fine-tuning Streaming
Status

Current and accessible; open-weight model; also available through the MiniMax API and MiniMax Agent

Input $0.15 per 1 million input tokens for standard-speed M2.5 at launch; current prices may vary by platform and endpoint
Output $1.20 per 1 million output tokens for standard-speed M2.5
MiniMax Coding
MiniMax
MiniMax M2

MiniMax M2.7

Agentic software engineering, repository-level coding, production debugging, complex tool workflows, office document automation, and long-context professional tasks

Context 197K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use
Status

Current; available through MiniMax API, MiniMax Agent, and downloadable open weights

Input $0.30 per 1 million tokens; cache read $0.06 per 1 million tokens; cache write $0.375 per 1 million tokens
Output $1.20 per 1 million tokens
MiniMax Coding

Low-latency coding assistants, multilingual software development, tool-using agents, long-horizon workflows, and interactive office automation

Context 205K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current and available through the MiniMax Open Platform API; historical model variant

Input ¥4.20 per 1M input tokens
Output ¥16.80 per 1M output tokens
MiniMax Coding

Low-latency coding assistants, software-engineering agents, tool-using workflows, search tasks, and long-context productivity automation

Context 205K
Reasoning 8/10
Speed 10/10
Outputs
Text
Capabilities
Tool use Web search Fine-tuning Streaming
Status

Current and available; highspeed variant of MiniMax M2.5

Input $0.30 per 1 million tokens
Output $2.40 per 1 million tokens
MiniMax Coding

Low-latency coding assistants, software-engineering agents, tool-calling workflows, and interactive developer applications

Context 205K
Reasoning 8/10
Speed 10/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current and available through the MiniMax API Platform; highspeed variant of MiniMax M2.7

Input $0.60 per million tokens; cache read $0.06 per million tokens; cache write $0.375 per million tokens
Output $2.40 per million tokens
MiniMax Multimodal
MiniMax
MiniMax M3

MiniMax M3

Long-context coding agents, autonomous tool-using workflows, multimodal document and video analysis, and private deployment

Context 1M
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Active; open-weight and available through the MiniMax API

Input $0.30 per million tokens for context up to 512K; $0.60 per million tokens for context from 512K to 1M. Listed rates are subject to MiniMax's current pricing terms.
Output $1.20 per million tokens for context up to 512K; $2.40 per million tokens for context from 512K to 1M. Listed rates are subject to MiniMax's current pricing terms.
MiniMax Other
MiniMax
MiniMax Music

MiniMax Music 2.0

Generating complete songs with expressive vocals, lyrics, melodies, instrumental arrangements, duets, a cappella passages, and cinematic musical soundscapes.

Reasoning 1/10
Speed 6/10
Outputs
Music
Status

Legacy or limited availability; MiniMax states that music models were no longer available through Token Plan from 2026-08-20, but a full model retirement date was not verified.

MiniMax Other
MiniMax
MiniMax Music

MiniMax Music 2.6

Text-to-music, instrumental generation, game and video scoring, detailed musical direction, and genre reinterpretation with Cover mode

Speed 7/10
Outputs
Music
Capabilities
Audio input Multimodal input
Status

Legacy or superseded; current public MiniMax navigation highlights Music 3.0, and Music 2.6 availability should be verified before use

MiniMax Other
MiniMax
MiniMax Music

MiniMax Music Cover

Reinterpreting existing songs in new genres, vocal styles, arrangements, and production directions while preserving the source melody

Reasoning 1/10
Speed 6/10
Outputs
Music
Capabilities
Audio input Multimodal input
Status

Current; specialized music-cover model introduced with MiniMax Music 2.6

Input Not applicable to token pricing; audio-generation pricing varies by MiniMax platform or deployment
Output Not applicable to token pricing; charged per generated audio result where applicable
MiniMax Other
MiniMax
MiniMax Speech

Speech-02-HD

High-quality multilingual voiceovers, audiobooks, narration, digital characters, advertising, education, and zero-shot voice cloning

Context 10K
Reasoning 1/10
Speed 6/10
Outputs
Speech
Capabilities
Audio input Multimodal input Streaming
Status

Legacy or older generation; still accessible through some partner platforms, but not listed as a core model on MiniMax's current global pricing page

Input CNY 3.5 per 10,000 characters on Alibaba Cloud Model Studio in China; current direct MiniMax pricing for this exact model is not verified
MiniMax Other

Low-latency multilingual text-to-speech, streaming voice agents, interactive applications, expressive narration, and voice cloning

Reasoning 1/10
Speed 9/10
Outputs
Speech
Capabilities
Audio input Multimodal input Streaming
Status

Legacy or superseded; exact current first-party availability is unclear

MiniMax Other

Real-time voice agents, conversational assistants, customer-service automation, interactive characters, multilingual speech, and low-latency text-to-speech

Reasoning 1/10
Speed 9/10
Outputs
Speech
Capabilities
Streaming
Status

Legacy or superseded; current first-party availability is unverified

MiniMax Other
MiniMax
Speech 2.6

Speech-2.6-HD

High-quality voiceovers, audiobooks, narration, localization, e-learning, game dialogue, accessibility audio, and production speech

Speed 7/10
Outputs
Speech
Status

Legacy or transition-era model; not prominently listed in MiniMax's current first-party speech catalog as of September 25, 2026

MiniMax Other
MiniMax
Speech 2.8

Speech-2.8-HD

High-quality expressive narration, audiobooks, podcasts, advertising, character voices, multilingual speech, and applications prioritizing audio fidelity over the lowest latency

Speed 8/10
Outputs
Speech
Capabilities
Streaming
Status

Current and available through the MiniMax API

Input $100 per 1 million characters for text-to-audio usage
MiniMax Other

Real-time text-to-speech, voice assistants, conversational agents, interactive applications, multilingual narration, gaming characters and expressive voice experiences

Reasoning 1/10
Speed 9/10
Outputs
Speech
Capabilities
Streaming
Status

Current and accessible through the MiniMax Open Platform API

Input $60 per 1 million characters
MiniMax Other

Text-to-video generation with explicit cinematic camera-movement direction, short advertising concepts, storyboards, and controlled visual experiments

Reasoning 1/10
Speed 6/10
Outputs
Video
Status

Legacy or limited availability; current official catalog status and continued first-party access are not clearly verified

Low-latency IDE autocomplete, fill-in-the-middle completion, code generation, code editing, test generation and developer assistants

Context 128K
Reasoning 6/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Active

Input $0.30 per million tokens; cached input $0.03 per million tokens
Output $0.90 per million tokens
Mistral AI
Leanstral

Leanstral 1.5

Lean 4 theorem proving, formal verification, autoformalization, proof debugging, and agentic proof engineering

Context 256K
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Tool use
Status

Public Preview; scheduled for retirement on 2026-09-30

Input Free
Output Free
Mistral AI Lightweight
Mistral AI
Ministral 3

Ministral 3 3B

Low-cost edge and local inference, image-aware assistants, document analysis, structured extraction, lightweight agents, task routing, and privacy-sensitive deployments.

Context 256K
Reasoning 6/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Active; generally available

Input $0.10 per 1 million tokens; cached input $0.01 per 1 million tokens
Output $0.10 per 1 million tokens
Mistral AI Lightweight
Mistral AI
Ministral 3

Ministral 3 8B

Efficient edge and local inference, image understanding, document workflows, structured extraction, lightweight agents, and high-volume text generation.

Context 256K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Active; generally available

Input $0.15 per million tokens; cached input $0.015 per million tokens
Output $0.15 per million tokens
Mistral AI Multimodal
Mistral AI
Ministral 3

Ministral 3 14B

Private assistants, local vision-language applications, multilingual workloads, document and image analysis, and cost-efficient agentic systems

Context 262K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Fine-tuning Streaming
Status

Active; generally available

Input $0.20 per 1 million tokens; cached input $0.02 per 1 million tokens
Output $0.20 per 1 million tokens
Mistral AI
Mistral Embed

Mistral Embed

Semantic search, retrieval-augmented generation, vector databases, document classification, clustering, duplicate detection and general text retrieval

Context 8K
Speed 8/10
Outputs
Embeddings
Status

Generally available

Input $0.10 per 1 million tokens
Output $0.10 per 1 million tokens
Mistral AI Multimodal
Mistral AI
Mistral Large

Mistral Large 3

Long-context enterprise assistants, multilingual applications, image-aware document analysis, agentic workflows, coding, RAG and self-hosted sovereign deployments

Context 256K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Fine-tuning
Status

Active; generally available; open-weight

Input $0.50 per 1M tokens; cached input $0.05 per 1M tokens
Output $1.50 per 1M tokens
Mistral AI Multimodal
Mistral AI
Mistral Medium

Mistral Medium 3.5

Agentic coding, software engineering, long-context analysis, multimodal document workflows, structured outputs and multi-step tool use

Context 256K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

GA; currently available; open weights

Input $1.50 per million input tokens; $0.15 per million cached input tokens
Output $7.50 per million output tokens
Mistral AI
Mistral OCR

OCR 4.0

High-volume OCR, structured document extraction, enterprise search, RAG ingestion, invoice processing, compliance workflows, and document automation

Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Generally available; superseded by OCR 4.1 as Mistral's latest OCR model

Input $4 per 1,000 pages; $0.40 per 1,000 cached pages; Batch API pricing reported at $2 per 1,000 pages
Output Included in page-based processing price; no separate output-token price
Mistral AI Multimodal
Mistral AI
Mistral Small

Mistral Small 4

Cost-efficient general chat, multimodal document analysis, coding, agentic workflows, and configurable reasoning

Context 256K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Fine-tuning
Status

Active; generally available

Input $0.15 per 1M tokens; cached input $0.015 per 1M tokens
Output $0.60 per 1M tokens

High-volume document extraction, scanned forms, handwriting, invoices, complex tables, archival digitization, and document-to-knowledge pipelines.

Reasoning 2/10
Speed 8/10
Outputs
Text Image
Capabilities
Image input Multimodal input
Status

Legacy; available for existing integrations and production workloads. OCR 4 is the newer model.

Input $2 per 1,000 pages
Output $3 per 1,000 annotated pages

OCR, document parsing, structured extraction, enterprise search, RAG ingestion, invoice processing, and document AI workflows

Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Generally Available

Input $4 per 1,000 pages; cached input $0.40 per 1,000 pages
Mistral AI Multimodal

Production-scale audio understanding, multilingual transcription, audio Q&A, meeting and call summarization, speech translation, and voice-driven function calling

Context 32K
Reasoning 6/10
Speed 6/10
Outputs
Text
Capabilities
Audio input Multimodal input Tool use Fine-tuning
Status

Active; generally available

Input $0.004 per minute of audio; $0.10 per 1 million input tokens
Output $0.40 per 1 million output tokens
Mistral AI Text To Speech
Mistral AI
Voxtral TTS

Voxtral TTS

Multilingual voice generation, expressive voice agents, zero-shot voice cloning, custom voice adaptation, and low-latency speech output

Reasoning 1/10
Speed 9/10
Outputs
Speech
Capabilities
Audio input Multimodal input Streaming
Status

GA; currently available through the Mistral API and Mistral Studio

Input $0 per 1 million input characters
Output $16 per 1 million output characters
Moonshot AI Multimodal
Moonshot AI
Kimi-Audio

Kimi-Audio-7B

Fine-tuning and research on speech recognition, audio understanding, audio classification, audio question answering, and speech-audio generation

Context 8K
Reasoning 4/10
Speed 5/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Fine-tuning Streaming
Status

Current open-weight base model; not instruction-tuned

Input No official hosted API pricing; downloadable open weights
Output No official hosted API pricing; downloadable open weights
Moonshot AI Multimodal

Self-hosted speech recognition, audio understanding, audio question answering, audio captioning, and spoken conversational agents

Reasoning 6/10
Speed 5/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Streaming
Status

Available as an open-weight self-hosted model

Input No official hosted API pricing
Output No official hosted API pricing
Moonshot AI Multimodal

Self-hosted image and video understanding, OCR, long-document analysis, visual question answering, screenshot perception, and efficient multimodal applications

Context 131K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Fine-tuning Streaming
Status

Current open-weight/downloadable model; a newer Kimi-VL-A3B-Thinking-2506 variant is recommended for stronger multimodal reasoning

Input No official Moonshot-hosted API price published; self-hosted/open-weight model
Output No official Moonshot-hosted API price published; self-hosted/open-weight model
Moonshot AI Reasoning

Local multimodal reasoning, mathematical visual question answering, OCR, document understanding, image and video analysis, and research on open-weight vision-language models

Context 131K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Streaming
Status

Available as downloadable open weights; superseded by Kimi-VL-A3B-Thinking-2506

Moonshot AI Reasoning

Open-weight image and video reasoning, OCR, chart interpretation, visual mathematics, long PDFs, high-resolution screenshots and GUI-agent grounding

Context 131K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Streaming
Status

Current open-weight model; publicly available for self-hosted and third-party inference

Input No official hosted API price; open-weight model
Output No official hosted API price; open-weight model
Moonshot AI General Purpose

Fine-tuning, foundation-model research, custom language systems, coding experiments, and self-hosted inference

Context 131K
Reasoning 8/10
Speed 3/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Open-weight and downloadable; legacy relative to newer Kimi K2.x models but still accessible from the official Hugging Face repository

Input No official hosted API price for the Base checkpoint; self-hosted weights
Output No official hosted API price for the Base checkpoint; self-hosted weights
Moonshot AI General Purpose

Open-weight coding assistants, tool-using agents, general-purpose chat, and self-hosted research deployments

Context 131K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Open-weight checkpoint available; legacy relative to Moonshot AI's current hosted API catalog

Input USD 0.60 per 1 million input tokens historically for the Kimi K2 preview API; current first-party pricing for this exact model is unverified
Output USD 2.50 per 1 million output tokens historically for the Kimi K2 preview API; current first-party pricing for this exact model is unverified
Moonshot AI Reasoning

Self-hosted reasoning agents, autonomous research, long-horizon tool workflows, coding, and complex multi-step analysis

Context 262K
Reasoning 9/10
Speed 5/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Retired from Moonshot's direct API; open-weight checkpoint remains available for self-hosted or third-party deployment

Moonshot AI Multimodal
Moonshot AI
Kimi K2

Kimi K2.6

Long-horizon software engineering, agentic coding, visual document understanding, tool-using workflows, and multi-agent orchestration

Context 262K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search
Status

Current; open-source model, available through Kimi, Kimi API, Kimi Code, and downloadable model weights

Long-horizon software engineering, repository-level coding, multi-file refactoring, debugging, coding agents, and tool-driven development workflows

Context 262K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Current and accessible through Kimi API; Kimi Code default service has moved to Kimi K2.8 Preview, while Kimi K2.7 Code remains available through API and Kimi K2.7 Code HighSpeed service.

Input ¥6.50 per 1M uncached input tokens; ¥1.30 per 1M cache-hit input tokens
Output ¥27.00 per 1M output tokens

Long-context software development, repository analysis, code completion, multi-file refactoring, and agentic coding workflows

Context 1.05M
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Preview; fully rolled out in Kimi Code

Input Not publicly listed; accessed through Kimi Code membership and quota
Output Not publicly listed; accessed through Kimi Code membership and quota
Moonshot AI Multimodal
Moonshot AI
Kimi K2.5

Kimi K2.5

Multimodal coding, visual debugging, long-context analysis, tool-using agents, and complex research or office workflows

Context 256K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

current

Input $0.60 per 1 million tokens; cached input approximately $0.10 per 1 million tokens
Output $3.00 per 1 million tokens
Moonshot AI Multimodal
Moonshot AI
Kimi K3

Kimi K3

Long-context coding, software engineering, multimodal document and video understanding, agentic workflows, technical research, and complex reasoning

Context 1M
Reasoning 9/10
Speed 5/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current and available; open-weight model

Input $3.00 per MTok cache-miss input; $0.30 per MTok cache-hit input
Output $15.00 per MTok
Moonshot AI Lightweight

Long-context text generation, local inference, continued pretraining, research, and fine-tuning

Context 1.05M
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current open-weight model; publicly downloadable

Input Not applicable; open-weight checkpoint with no official hosted API price
Output Not applicable; open-weight checkpoint with no official hosted API price

Native-resolution image feature extraction, vision-language model backbones, high-resolution document and image understanding, and multimodal research

Outputs
Embeddings
Capabilities
Image input
Status

Current open-weight model; available for local use through Hugging Face Transformers; not deployed by a Hugging Face Inference Provider

NAVER AI Lightweight

Fast, cost-sensitive Korean text generation, classification, summarization, report drafting, data expansion and customized enterprise chatbots

Context 4K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current and available through CLOVA Studio and related NAVER Cloud services

Input Current standard inference price not clearly exposed in the retrieved official pricing table; NAVER's launch announcement described HCX-DASH-001 as costing approximately one-fifth of HCX-003.
Output Current standard inference price not clearly exposed in the retrieved official pricing table; NAVER's launch announcement described HCX-DASH-001 as costing approximately one-fifth of HCX-003.
NAVER AI General Purpose
NAVER AI
HyperCLOVA X

HCX-003

Korean-focused text generation, summarization, extraction, classification, tuning, and batch text workflows

Context 8K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current and available through CLOVA Studio

NAVER AI Multimodal
NAVER AI
HyperCLOVA X

HCX-005

Korean-language business applications, image understanding, document and visual analysis, instruction following, and API-based assistants.

Context 128K
Reasoning 6/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Fine-tuning Streaming
Status

current

Input Priced per 1,000 input tokens; current amount not displayed in the accessible official pricing table.
Output Priced per 1,000 output tokens; current amount not displayed in the accessible official pricing table.
NAVER AI Reasoning
NAVER AI
HyperCLOVA X

HCX-007

Complex reasoning, mathematics, science, language reasoning, writing, long-context text generation, and Korean-language enterprise applications

Context 128K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current and accessible through CLOVA Studio

NAVER AI Lightweight
NAVER AI
HyperCLOVA X

HCX-DASH-002

Fast, high-throughput text generation, classification, summarization, simple extraction, and function-calling workflows

Context 32K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available

NAVER AI Lightweight
NAVER AI
HyperCLOVA X SEED

HyperCLOVA X SEED 0.5B

Lightweight Korean conversational interfaces, mobile and edge applications, smart-home devices, wearables, and customer-support chatbots

Context 4K
Reasoning 3/10
Speed 9/10
Outputs
Text
Capabilities
Fine-tuning
Status

Current open-weight model; available for download

NAVER AI Lightweight
NAVER AI
HyperCLOVA X SEED

HyperCLOVA X SEED 1.5B

Korean-language applications, lightweight local inference, basic translation, education, business communication, specialized chatbots, and domain fine-tuning.

Context 16K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Current downloadable open-weight model; Hugging Face repository is gated and requires acceptance of access conditions.

NAVER AI Multimodal
NAVER AI
HyperCLOVA X SEED

HyperCLOVA X SEED 3B

Korean image and video understanding, visual question answering, chart and diagram interpretation, OCR-assisted analysis, tourism and cultural applications, and locally deployable fine-tuned systems.

Context 16K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Fine-tuning Streaming
Status

Available open-weight model

Input No official hosted API price identified; model weights are available for download under the hyperclovax-seed license.
Output No official hosted API price identified; deployment costs depend on infrastructure or hosting provider.
NAVER AI Multimodal
NAVER AI
HyperCLOVA X SEED

HyperCLOVA X SEED 4B

Korean visual and document understanding, video-and-audio analysis, edge AI, public-sector systems, defense intelligence, and air-gapped deployments

Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input
Status

Publicly announced; current model identity, with detailed commercial/API availability not publicly specified

NAVER AI Multimodal
NAVER AI
HyperCLOVA X SEED

HyperCLOVA X SEED 8B Omni

Korean-first any-to-any multimodal assistants, speech and vision applications, multimodal research, and self-hosted deployments

Context 33K
Reasoning 7/10
Speed 6/10
Outputs
Text Image Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current open-weight model

NAVER AI Reasoning

Korean-language reasoning, mathematics, coding, instruction following, tool-connected agents, and self-hosted commercial applications

Context 33K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current open-weight model; free for commercial use under the HyperCLOVA X SEED license

Input Free model weights; no official hosted API input price published for this exact checkpoint
Output Free model weights; no official hosted API output price published for this exact checkpoint
NAVER AI Reasoning
NAVER AI
HyperCLOVA X SEED Think

HyperCLOVA X SEED 32B Think

Korean-language reasoning, visual question answering, long-context multimodal analysis, document and chart understanding, coding assistance, and tool-using AI agents

Context 128K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Current; open-weight/open-source release

Input No official hosted API price published; self-hosting costs depend on infrastructure
Output No official hosted API price published; self-hosting costs depend on infrastructure
NVIDIA AI Multimodal
NVIDIA AI
Active Speaker Detection

Active Speaker Detection

Real-time or batch speaker identification and active-speaker tagging in broadcast, video communication, dubbing, localization, conferencing, and media analytics workflows.

Reasoning 1/10
Speed 8/10
Capabilities
Audio input Video input Multimodal input Streaming
Status

Current; downloadable NVIDIA NIM endpoint

NVIDIA AI
AI4M Relighting

Relighting

Real-time video relighting, virtual production, media effects, HDR-based lighting changes, and foreground/background compositing.

Reasoning 1/10
Speed 8/10
Outputs
Video
Capabilities
Image input Video input Multimodal input Streaming
Status

Current and accessible through NVIDIA AI for Media and NVIDIA NIM documentation; version 1.1.0 is documented.

NVIDIA AI Reasoning

Autonomous-driving research, trajectory prediction, interpretable motion planning, navigation-conditioned driving, visual question answering and safety-oriented model evaluation

Reasoning 8/10
Speed 5/10
Outputs
Text Actions
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Current; open weights available; also available through NVIDIA Alpamayo 1.5 NIM

Multilingual offline speech transcription, speech translation, and NeMo-based ASR research or deployment

Reasoning 1/10
Speed 6/10
Outputs
Text
Capabilities
Audio input
Status

Active; open-weight checkpoint and NVIDIA deployment options available

NVIDIA AI Multimodal
NVIDIA AI
Cosmos-Transfer2.5

Cosmos-Transfer2.5-2B

Controllable video world generation, robotics sim-to-real augmentation, autonomous-vehicle simulation, and Physical AI synthetic-data generation

Reasoning 1/10
Speed 3/10
Outputs
Video
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Available; legacy relative to Cosmos 3; original repository under limited maintenance

Input No official token-based hosted price; NVIDIA lists a free NIM endpoint. Downloadable model is intended for self-hosted deployment.
Output No official token-based hosted price; video-generation costs depend on deployment infrastructure and hardware.
NVIDIA AI Multimodal

Edge physical AI, robotics, visual reasoning, world simulation, video generation, and action-policy prototyping

Context 131K
Reasoning 7/10
Speed 9/10
Outputs
Text Image Video Actions
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Current; open model; gated Hugging Face access

NVIDIA AI Multimodal

Physical AI, robotics, autonomous-vehicle simulation, multimodal world generation, future-state prediction, action reasoning, and synthetic training data

Reasoning 7/10
Speed 6/10
Outputs
Text Image Video Audio Actions
Capabilities
Image input Audio input Video input Multimodal input Fine-tuning
Status

Current; downloadable open-weight model and available through an NVIDIA NIM endpoint

NVIDIA AI Multimodal

High-quality Physical AI simulation, synthetic-data generation, robotics and autonomous-vehicle research, multimodal world modeling, and teacher-model distillation

Context 262K
Reasoning 8/10
Speed 2/10
Outputs
Text Image Video Audio Actions
Capabilities
Image input Audio input Video input Multimodal input Fine-tuning
Status

Current; open-weight model; commercially and non-commercially usable

NVIDIA AI Reasoning

Physical-world video and image understanding, robotic perception, embodied-agent planning, spatial-temporal reasoning, and Physical AI research

Context 256K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Current; downloadable and hosted NIM endpoint

Input Free downloadable endpoint; hosted pricing not specified in the reviewed NVIDIA sources
Output Free downloadable endpoint; hosted pricing not specified in the reviewed NVIDIA sources
NVIDIA AI
FourCastNet

FourCastNet

Rapid global weather forecasting, ensemble simulation, climate-risk analysis, renewable-energy forecasting, and scientific weather-model research

Reasoning 0/10
Speed 9/10
Status

Current and downloadable; available through NVIDIA Earth-2 FourCastNet NIM and related deployment tooling

Input Not publicly listed as a token-priced model; downloadable model and NIM access are subject to NVIDIA access and licensing terms
Output Not publicly listed as a token-priced model; deployment and hosted inference pricing may depend on the selected NVIDIA service or infrastructure
NVIDIA AI
GenMol

GenMol

De novo molecular design, fragment-constrained generation, linker design, scaffold decoration, hit generation, and lead optimization

Context 512
Reasoning 2/10
Speed 7/10
Outputs
Text
Status

Current; downloadable model and available as an NVIDIA NIM

Input No public per-token model price found; downloadable weights are available under the NVIDIA Open Model License. NIM access is subject to NVIDIA API or deployment terms.
Output No public per-token model price found; NIM output pricing is not specified in the reviewed official documentation.
NVIDIA AI Multimodal

Humanoid robot manipulation, cross-embodiment policy learning, robot demonstration fine-tuning, physical AI research, and action-sequence deployment

Reasoning 4/10
Speed 6/10
Outputs
Actions
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Current; general availability

NVIDIA AI Multimodal

Quantum-computing calibration plot interpretation, QPU bring-up and retuning workflows, experiment diagnosis, parameter extraction, fit-quality assessment, and domain-specific calibration agents.

Context 262K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning Streaming
Status

Current; available through NVIDIA Build, NVIDIA NIM, and downloadable checkpoints

NVIDIA AI Multimodal
NVIDIA AI
LipSync

LipSync

Generative lip dubbing, multilingual video localization, broadcasting, conferencing, and digital-human facial animation

Reasoning 1/10
Speed 8/10
Outputs
Image
Capabilities
Image input Audio input Video input Multimodal input Streaming
Status

Current; downloadable model and NVIDIA LipSync NIM; private access may be required for some workflows

NVIDIA AI Multimodal
NVIDIA AI
Llama Nemotron Rerank VL

Llama Nemotron Rerank VL 1B v2

Reranking text, document images, and image-text candidates in visual search, multimodal RAG, and question-answering retrieval pipelines

Context 8K
Reasoning 2/10
Speed 7/10
Capabilities
Image input Multimodal input
Status

Current; downloadable and available through NVIDIA NIM and retrieval APIs

NVIDIA AI
MolMIM

MolMIM

Small-molecule generation, molecular embeddings, chemical-space exploration, lead optimization, and oracle-guided drug-design workflows

Context 128
Outputs
Text Embeddings
Status

Available through NVIDIA BioNeMo Framework and NVIDIA NIM; research and development model

NVIDIA AI Reasoning

Agentic reasoning, long-context analysis, coding, tool-calling workflows, retrieval-augmented generation, collaborative agents, and high-volume enterprise inference.

Context 1M
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current; open-weight; available through NVIDIA NIM, NVIDIA's hosted API trial endpoint, downloadable checkpoints, and self-hosted deployments.

NVIDIA AI Multimodal

Multimodal document intelligence, OCR, long-video and audio understanding, cross-modal reasoning, voice agents, and self-hosted enterprise inference

Context 262K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current; generally available open-weight checkpoint with hosted NVIDIA NIM access

NVIDIA AI Multimodal
NVIDIA AI
Nemotron 3 VoiceChat

NVIDIA Nemotron 3 VoiceChat

Real-time full-duplex voice agents, interruptible conversational interfaces, speech-to-speech research, and NVIDIA GPU-based enterprise voice applications

Reasoning 5/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Fine-tuning Streaming
Status

Early access; available for evaluation through NVIDIA NIM and qualified access programs

Input Free endpoint for NVIDIA NIM trial access; no general production price published
Output Free endpoint for NVIDIA NIM trial access; no general production price published
NVIDIA AI
Nemotron Content Safety

Nemotron 3.5 Content Safety

Multilingual text-and-image moderation, LLM and VLM guardrails, response safety evaluation, and custom enterprise safety policies

Context 128K
Reasoning 5/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Streaming
Status

Current; open-weight model and downloadable NVIDIA NIM; latest NGC container version bf16-v1.1 as of September 17, 2026

NVIDIA AI
Nemotron OCR

Nemotron OCR v1

English OCR, document ingestion, layout-aware text extraction, multimodal retrieval, RAG preprocessing, and enterprise document intelligence

Reasoning 1/10
Speed 8/10
Outputs
Text
Capabilities
Image input
Status

Available; English-only OCR model; newer Nemotron OCR v2 is available for updated English and multilingual OCR deployments

Multilingual OCR, scanned documents, forms, reports, charts, tables, image-based search, document ingestion, and retrieval-augmented generation preprocessing

Reasoning 1/10
Speed 9/10
Outputs
Text
Capabilities
Image input
Status

Current; downloadable model and available through NVIDIA NIM and NVIDIA-hosted services

NVIDIA AI
NVIDIA Maxine Eye Contact

Eye Contact

Gaze correction in video conferencing, telepresence, digital-human applications, and video-processing pipelines

Reasoning 1/10
Speed 8/10
Outputs
Image
Capabilities
Image input Video input Multimodal input
Status

Current; downloadable NVIDIA NIM model

Mandarin-English automatic speech recognition, code-switched transcription, streaming transcription, and self-hosted enterprise speech-to-text

Reasoning 1/10
Speed 8/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current and downloadable; available through NVIDIA NIM and Riva interfaces

Input No public per-token price; downloadable model and NVIDIA service terms apply
Output No public per-token price; downloadable model and NVIDIA service terms apply

High-speed English speech transcription, captions, meeting transcription, audio search, and timestamped media workflows

Reasoning 1/10
Speed 10/10
Outputs
Text
Capabilities
Audio input Fine-tuning Streaming
Status

Current open-weight model; publicly available through Hugging Face and NVIDIA NeMo

Multilingual sentence and document translation, localization, marketing content, and developer translation workflows

Context 8K
Reasoning 2/10
Speed 7/10
Outputs
Text
Status

Current and accessible through NVIDIA NIM and downloadable model weights

Input Free hosted endpoint listed on Build.NVIDIA.com; no self-hosted per-token price specified
Output Free hosted endpoint listed on Build.NVIDIA.com; no self-hosted per-token price specified
NVIDIA AI
StreamPETR

StreamPETR

Camera-only multi-view 3D perception, autonomous-driving scene analysis, bird's-eye-view visualization, and object tracking

Speed 8/10
Outputs
Video
Capabilities
Image input Video input Multimodal input Streaming
Status

Current NVIDIA NIM endpoint; free endpoint access requires an NVIDIA API key

Input Free endpoint
NVIDIA AI
Studio Voice

Studio Voice

Real-time enhancement of speech captured with low-quality microphones in noisy or reverberant environments, including broadcast, conferencing, telecommunications, and media production.

Reasoning 1/10
Speed 9/10
Outputs
Speech
Capabilities
Audio input Streaming
Status

Current and available through NVIDIA NIM, hosted preview services, and NVIDIA audio software; downloadable deployment may require applicable NVIDIA licensing or subscription access.

NVIDIA AI
Synthetic Video Detector

NVIDIA Synthetic Video Detector

AI-generated video detection, media authentication, digital forensics, content moderation, and media-integrity monitoring

Reasoning 1/10
Speed 8/10
Outputs
Text
Capabilities
Video input Streaming
Status

Current; downloadable NIM and managed trial endpoint, with self-hosted deployment requiring AI for Media Private Access

Input Free managed trial endpoint; self-hosted and enterprise licensing terms may apply
OpenAI General Purpose
OpenAI
Chat Latest

Chat Latest

ChatGPT-style instant responses, general-purpose writing and analysis, image-aware conversations, and tool-assisted workflows

Context 400K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current rolling alias; underlying model snapshot is regularly updated

Input $5.00 per 1 million input tokens; cached input $0.50 per 1 million tokens
Output $30.00 per 1 million output tokens
OpenAI Multimodal
OpenAI
CLIP

CLIP

Zero-shot image classification, image-text similarity, semantic image retrieval, multimodal indexing, and computer-vision research

Context 77
Reasoning 2/10
Speed 7/10
Outputs
Embeddings
Capabilities
Image input Multimodal input
Status

Public research release with downloadable weights; not verified as a current OpenAI hosted API model

OpenAI Coding

Codex CLI coding workflows, code question answering, code editing, repository tasks, and low-latency software-engineering assistance

Context 200K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; API access ended on 2026-02-12

Input $1.50 per 1M input tokens; $0.375 per 1M cached input tokens
Output $6.00 per 1M output tokens
OpenAI Other
OpenAI
Computer-Using Agent

computer-use-preview

Controlled browser automation, computer-use research, UI testing, and repetitive interface workflows

Context 8K
Reasoning 6/10
Speed 7/10
Outputs
Text Actions
Capabilities
Image input Multimodal input Tool use
Status

Deprecated

Input $3.00 per 1 million input tokens; tool-specific computer-use calls may incur separate fees
Output $12.00 per 1 million output tokens
OpenAI Other
OpenAI
DALL·E

DALL·E 2

Historical research on text-to-image generation, legacy image workflows, and comparisons with newer OpenAI image models.

Outputs
Image
Capabilities
Image input Multimodal input
Status

Retired; deprecated and removed from the OpenAI API on May 12, 2026.

OpenAI Other
OpenAI
DALL·E

DALL·E 3

Historical text-to-image generation, concept art, illustration, visual ideation, marketing imagery, and prompt-following research

Reasoning 1/10
Speed 6/10
Outputs
Image
Status

Retired; deprecated and removed from the OpenAI API on May 12, 2026

Input Not applicable to current use; historical pricing was charged per generated image rather than per input token
Output Historical API pricing started at $0.04 per 1024×1024 standard-quality image; higher prices applied to HD and larger formats
OpenAI Lightweight

Maintaining legacy text-completion applications, historical GPT-3 base-model behavior, and existing compatible fine-tuned workflows before shutdown

Reasoning 2/10
Speed 7/10
Outputs
Text
Status

Deprecated; currently accessible but scheduled to shut down on 2026-09-28

Input $0.40 per 1 million input tokens
Output $0.40 per 1 million output tokens
OpenAI General Purpose

Legacy text completion, code continuation, and inference from existing davinci-002 fine-tuned models before shutdown

Reasoning 3/10
Speed 6/10
Outputs
Text
Status

Deprecated; API access scheduled to shut down on September 28, 2026

Input $2.00 per 1M tokens
Output $2.00 per 1M tokens
OpenAI General Purpose

Low-cost, high-volume text generation, summarization, classification, extraction, simple chatbots, and legacy API integrations

Context 16K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Deprecated; still available through the OpenAI API

Input $0.50 per 1 million input tokens
Output $1.50 per 1 million output tokens
OpenAI General Purpose
OpenAI
GPT-4

GPT-4

Maintaining established GPT-4 integrations, general-purpose text generation, analysis, writing, and coding workloads

Context 8K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Legacy; older high-intelligence GPT model

Input $30 per 1 million prompt tokens
Output $60 per 1 million completion tokens
OpenAI Multimodal

Legacy high-context text and image analysis, function calling, JSON-mode workflows, and existing GPT-4 Turbo integrations

Context 128K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Deprecated but currently accessible; scheduled for shutdown on October 23, 2026

Input $10 per 1 million input tokens
Output $30 per 1 million output tokens
OpenAI General Purpose

Historical long-context text generation, document analysis, structured text generation, and general-purpose assistant applications.

Context 128K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Retired; the gpt-4-turbo-preview alias pointed to gpt-4-0125-preview, which was shut down on 2026-03-26.

Input $10 per 1 million tokens
Output $30 per 1 million tokens
OpenAI General Purpose
OpenAI
GPT-4.1

GPT-4.1

Software engineering, long-context document analysis, precise instruction following, structured extraction, tool-enabled agents, and image understanding

Context 1.05M
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Fine-tuning
Status

Current; default GPT-4.1 alias with gpt-4.1-2025-04-14 snapshot

Input $2.00 per 1M input tokens; $0.50 per 1M cached input tokens
Output $8.00 per 1M output tokens
OpenAI Lightweight

Fast, cost-efficient instruction following, coding assistance, image understanding, structured extraction, tool calling, and long-context API applications

Context 1.05M
Reasoning 6/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Fine-tuning
Status

current

Input $0.40 per 1 million tokens; cached input $0.10 per 1 million tokens
Output $1.60 per 1 million tokens
OpenAI Lightweight

High-volume, latency-sensitive classification, extraction, routing, summarization, lightweight assistants, image-assisted analysis, and simple tool-calling workflows

Context 1.05M
Reasoning 4/10
Speed 10/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Fine-tuning Streaming
Status

Deprecated; currently accessible as of September 23, 2026; scheduled for shutdown on October 23, 2026

Input $0.10 per 1M input tokens; $0.025 per 1M cached input tokens
Output $0.40 per 1M output tokens
OpenAI General Purpose

Historical general-purpose writing, creative work, nuanced communication, image understanding, programming assistance, and applications needing function calling or structured outputs.

Context 128K
Reasoning 7/10
Speed 3/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired from the API on 2025-07-14; retired from ChatGPT in June 2026

Input $75.00 per 1 million input tokens; $37.50 per 1 million cached input tokens
Output $150.00 per 1 million output tokens
OpenAI Multimodal

Fast general-purpose conversations, vision, voice interactions, coding, and everyday productivity

Context 128K
Reasoning 8/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Active

OpenAI Multimodal
OpenAI
GPT-4o

GPT-4o

General-purpose assistants, image understanding, coding help, structured extraction, multilingual generation, and latency-sensitive API workflows

Context 128K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Fine-tuning Streaming
Status

Current in the OpenAI API; retired from ChatGPT on 2026-02-13. The gpt-4o-2024-05-13 snapshot is scheduled for API shutdown on 2026-10-23.

Input $2.50 per 1M input tokens; $1.25 per 1M cached input tokens
Output $10.00 per 1M output tokens
OpenAI Multimodal

Voice assistants, spoken conversational agents, audio-enabled customer service, and applications requiring direct audio understanding and speech generation

Context 128K
Reasoning 7/10
Speed 6/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Streaming
Status

Retired; API access ended May 7, 2026

Input $2.50 per 1M text input tokens; $40 per 1M audio input tokens
Output $10.00 per 1M text output tokens; $80 per 1M audio output tokens
OpenAI Lightweight

Low-cost, high-volume text and image understanding, classification, extraction, translation, tagging, customer support, routing, and structured data generation

Context 128K
Reasoning 5/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Fine-tuning Streaming
Status

Current canonical model alias; dated snapshot gpt-4o-mini-2024-07-18 is available

Input $0.15 per 1M input tokens
Output $0.60 per 1M output tokens
OpenAI Multimodal

Lower-cost audio understanding, conversational voice interfaces, and applications requiring text and spoken-audio input/output

Context 128K
Reasoning 3/10
Speed 8/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Streaming
Status

Deprecated; scheduled for API shutdown on 2027-01-20

Input Text: $0.15 per 1M tokens; audio: $10.00 per 1M tokens
Output Text: $0.60 per 1M tokens; audio: $20.00 per 1M tokens
OpenAI Lightweight

Low-cost realtime voice assistants, speech-to-speech interfaces, interactive audio applications, and conversational prototypes

Context 16K
Reasoning 4/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use
Status

Deprecated; scheduled for API removal on 2027-01-20

Input Text: $0.60 per 1M input tokens; audio: $10.00 per 1M audio tokens; cached input: $0.30 per 1M tokens
Output Text: $2.40 per 1M output tokens; audio: $20.00 per 1M audio tokens
OpenAI Multimodal

Low-latency voice assistants, speech-to-speech applications, live translation, language learning, and interactive customer support

Context 32K
Reasoning 6/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use
Status

Retired; API access ended 2026-05-07

Input $5 per 1M text tokens; $40 per 1M audio tokens; cached text and audio input $2.50 per 1M tokens
Output $20 per 1M text tokens; $80 per 1M audio tokens
OpenAI Other

Historical web-search applications built around OpenAI Chat Completions

Context 128K
Reasoning 5/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Retired; shut down on 2026-07-23

Input $2.50 per 1 million input tokens; historical web-search tool fees applied separately per search call
Output $10.00 per 1 million output tokens
OpenAI Other

Accurate speech-to-text conversion, meeting transcription, call transcription, voice-agent input, and prompted domain-specific transcription

Context 16K
Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Audio input Multimodal input Streaming
Status

Deprecated; currently accessible; scheduled for API shutdown on 2027-02-26

Input $2.50 per 1M audio input tokens
Output $10.00 per 1M audio output tokens
OpenAI Other

Lower-cost multilingual speech transcription, meeting notes, call-center transcripts, voice-note conversion, and audio-to-text pipelines

Context 16K
Reasoning 1/10
Speed 8/10
Outputs
Text
Capabilities
Audio input Multimodal input Streaming
Status

Current and available

Input $1.25 per 1 million audio tokens
Output $5.00 per 1 million audio tokens
OpenAI Other
OpenAI
GPT-4o Mini

GPT-4o Mini TTS

Fast, controllable text-to-speech for narration, voice interfaces, customer service, accessibility, and realtime audio applications.

Context 2K
Reasoning 1/10
Speed 9/10
Outputs
Speech
Capabilities
Streaming
Status

Current; the canonical alias currently points to the gpt-4o-mini-tts-2025-12-15 snapshot.

Input $0.60 per 1M text input tokens
Output $12.00 per 1M audio output tokens
OpenAI Lightweight
OpenAI
GPT-4o Mini Search Preview

GPT-4o Mini Search Preview

Legacy Chat Completions applications requiring low-cost, search-grounded text responses

Context 128K
Reasoning 5/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Retired; access shut down on July 23, 2026

Input $0.15 per 1 million input tokens, plus a separate fee per web-search tool call
Output $0.60 per 1 million output tokens
OpenAI Other
OpenAI
GPT-4o Transcribe

GPT-4o Transcribe Diarize

Multi-speaker meeting, interview, call, podcast, and research transcription with speaker labels.

Context 16K
Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Deprecated; currently accessible and scheduled for removal from the API on February 26, 2027.

Input $2.50 per 1M audio tokens
Output $10.00 per 1M audio tokens
OpenAI Reasoning
OpenAI
GPT-5

GPT-5

Complex coding, reasoning, research, long-context analysis, visual understanding, tool-using agents, and structured professional workflows

Context 400K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current canonical alias, but previous-generation model; dated snapshot gpt-5-2025-08-07 is deprecated and scheduled for API shutdown on 2026-12-11

Input $1.25 per 1 million input tokens; cached input $0.125 per 1 million tokens
Output $10.00 per 1 million output tokens
OpenAI General Purpose

ChatGPT-aligned conversational applications, text generation, image-aware question answering, structured outputs, and tool-enabled workflows requiring GPT-5 compatibility

Context 128K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Deprecated

Input $1.25 per 1M input tokens; $0.125 per 1M cached input tokens
Output $10.00 per 1M output tokens
OpenAI Lightweight

Cost-sensitive reasoning, coding assistance, structured extraction, document processing, high-volume automation, and tool-enabled workflows

Context 400K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current API alias; dated snapshot gpt-5-mini-2025-08-07 is deprecated

Input US$0.25 per 1 million input tokens; cached input US$0.025 per 1 million tokens
Output US$2.00 per 1 million output tokens
OpenAI Lightweight

High-volume classification, summarization, extraction, ranking, routing, image-assisted analysis, and lightweight coding subagents

Context 400K
Reasoning 7/10
Speed 10/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Deprecated dated snapshot; currently accessible until scheduled shutdown on 2026-12-11

Input $0.05 per 1 million tokens; cached input $0.005 per 1 million tokens
Output $0.40 per 1 million tokens
OpenAI Reasoning

Difficult research, mathematics, science, complex coding, high-stakes analysis, and tool-using workflows where maximum answer quality matters more than latency or cost.

Context 400K
Reasoning 10/10
Speed 3/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current canonical alias; dated snapshot gpt-5-pro-2025-10-06 is deprecated and scheduled for shutdown on 2026-12-11.

Input $15 per 1M input tokens
Output $120 per 1M output tokens
OpenAI Coding

Agentic software engineering, repository-level coding, code review, debugging, refactoring, test generation, and frontend work using screenshots

Context 400K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; API access shut down on 2026-07-23

Input $1.25 per 1M input tokens; $0.125 per 1M cached input tokens
Output $10.00 per 1M output tokens
OpenAI General Purpose
OpenAI
GPT-5

GPT-5.1

Coding, long-context analysis, tool-using agents, structured outputs, and multi-step workflows

Context 400K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current

Input $1.25 per 1 million input tokens; $0.125 per 1 million cached input tokens
Output $10.00 per 1 million output tokens
OpenAI General Purpose

Conversational assistants, instruction following, image-grounded chat, structured extraction, streaming responses, and tool-using API workflows.

Context 128K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; the API alias gpt-5.1-chat-latest was shut down on 2026-07-23. GPT-5.1 models were retired from ChatGPT on 2026-03-11.

Input $1.25 per 1 million input tokens; cached input $0.125 per 1 million tokens
Output $10.00 per 1 million output tokens
OpenAI Coding

Agentic software engineering, code generation, debugging, refactoring, testing, code review, and long-running Codex workflows

Context 400K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; API access shut down on July 23, 2026

Input $1.25 per 1M input tokens; $0.125 per 1M cached input tokens
Output $10.00 per 1M output tokens
OpenAI Coding
OpenAI
GPT-5.1-Codex

GPT-5.1-Codex-Max

Long-running agentic coding, repository-scale refactoring, multi-file implementation, debugging, code review, pull-request creation, and extended Codex workflows.

Context 400K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; API access ended 2026-07-23

Input $1.25 per 1M tokens; cached input $0.125 per 1M tokens
Output $10.00 per 1M tokens
OpenAI Coding
OpenAI
GPT-5.1-Codex

GPT-5.1-Codex Mini

Cost-sensitive agentic coding, code editing, repository maintenance, and Codex-style workflows

Context 400K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; API access shut down on 2026-07-23

Input $0.25 per 1 million tokens; cached input $0.025 per 1 million tokens
Output $2.00 per 1 million tokens
OpenAI Reasoning
OpenAI
GPT-5.2

GPT-5.2

Complex professional work, long-context analysis, coding, document and spreadsheet workflows, visual understanding, and multi-step agents

Context 400K
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Currently available; previous flagship model

Input $1.75 per 1M input tokens; $0.175 per 1M cached input tokens
Output $14.00 per 1M output tokens
OpenAI General Purpose

ChatGPT-aligned conversational applications, general writing, summarization, translation, vision-enabled assistants, and tool-calling workflows

Context 128K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; API access ended on 2026-08-10

Input $1.75 per 1 million input tokens; $0.175 per 1 million cached input tokens
Output $14.00 per 1 million output tokens
OpenAI Reasoning

Complex professional reasoning, advanced analysis, scientific and mathematical work, high-quality coding, long-context document analysis, and tool-using workflows.

Context 400K
Reasoning 10/10
Speed 4/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Previous Pro model; currently available through the Responses API

Input $21 per 1M tokens
Output $168 per 1M tokens
OpenAI Coding

Long-horizon agentic coding, large refactors, code migrations, repository-scale changes, terminal workflows, Windows development and defensive cybersecurity

Context 400K
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; API access shut down on 2026-07-23

Input $1.75 per 1 million input tokens
Output $14.00 per 1 million output tokens
OpenAI General Purpose

Fast general-purpose conversation, writing, summarization, text-and-image understanding, streaming responses, and function-calling applications

Context 128K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Retired; API access ended 2026-08-10

Input $1.75 per 1M tokens; cached input $0.175 per 1M tokens
Output $14.00 per 1M tokens
OpenAI Coding

Long-running agentic software engineering, codebase maintenance, debugging, testing, web development, tool-driven development, and technical computer workflows

Context 400K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current and available through OpenAI API and Codex surfaces

Input $1.75 per 1 million input tokens
Output $14.00 per 1 million output tokens
OpenAI Reasoning
OpenAI
GPT-5.4

GPT-5.4

Complex professional work, advanced reasoning, software engineering, long-horizon agents, visual document analysis, computer use, research, and tool-heavy workflows

Context 1.05M
Reasoning 10/10
Speed 8/10
Outputs
Text Actions
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current

Input $2.50 per 1 million input tokens; $0.25 per 1 million cached input tokens
Output $15.00 per 1 million output tokens
OpenAI Lightweight

High-volume coding assistants, computer-use agents, subagents, tool calling, image reasoning, document workflows, and latency-sensitive applications

Context 400K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current

Input $0.75 per 1 million input tokens; $0.075 per 1 million cached input tokens
Output $4.50 per 1 million output tokens
OpenAI Lightweight

High-volume classification, data extraction, ranking, image understanding, routing, and lightweight coding subagents

Context 400K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current; API-only model

Input $0.20 per 1M input tokens; $0.02 per 1M cached input tokens
Output $1.25 per 1M output tokens
OpenAI Reasoning

High-stakes reasoning, professional knowledge work, long-context analysis, complex coding, web research and agentic workflows requiring maximum answer quality

Context 1.05M
Reasoning 10/10
Speed 4/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current; available in ChatGPT for Pro and Enterprise users and in the Responses API for developers

Input $30 per 1 million input tokens
Output $180 per 1 million output tokens
OpenAI Coding

Authorized vulnerability research, defensive cybersecurity operations, malware analysis, security testing, and binary reverse engineering.

Reasoning 9/10
Outputs
Text
Status

Deprecated; currently accessible through restricted Trusted Access for Cyber channels; scheduled for API shutdown on October 1, 2026.

OpenAI General Purpose
OpenAI
GPT-5.5

GPT-5.5

Complex coding, long-context research, professional analysis, tool-heavy agents, computer use, and multi-step workflow execution

Context 1.05M
Reasoning 10/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current; available through the OpenAI API, ChatGPT, and Codex

Input $5.00 per 1 million input tokens; cached input $0.50 per 1 million tokens
Output $30.00 per 1 million output tokens
OpenAI Reasoning

High-accuracy reasoning, complex coding, long-context research, data analysis, and multi-step professional workflows

Context 1.05M
Reasoning 10/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search
Status

Current

Input USD 30 per 1 million input tokens; no cached-input discount
Output USD 180 per 1 million output tokens
OpenAI Coding

Authorized vulnerability research, exploit validation, exploit-chain development, advanced security testing, vulnerability triage, and defensive cybersecurity agents

Context 400K
Reasoning 10/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current; restricted access through OpenAI Daybreak Red with separate approval and provisioning

Input $12.50 per 1 million input tokens; cached input $1.25 per 1 million tokens
Output $75.00 per 1 million output tokens
OpenAI Lightweight

High-volume classification, summarization, routing, extraction, document understanding, agent automation, routine coding assistance, and cost-sensitive tool-using applications.

Context 1.05M
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

current

Input $0.20 per 1 million input tokens; cached input $0.02 per 1 million tokens; cache writes billed at 1.25x the uncached input rate. Requests with more than 272,000 input tokens are priced at 2x input for the full request.
Output $1.20 per 1 million output tokens. Requests with more than 272,000 input tokens are priced at 1.5x output for the full request.
OpenAI Reasoning

Complex reasoning, coding, research, cybersecurity, science, long-context analysis, document-heavy workflows, and tool-using agents

Context 1.05M
Reasoning 10/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Generally available

Input $4 per 1 million input tokens; cached input $0.40 per 1 million tokens
Output $20 per 1 million output tokens
OpenAI General Purpose

Cost-conscious reasoning, coding agents, long-context analysis, structured business automation, research workflows, and tool-enabled production applications

Context 1.05M
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Generally available

Input $2.00 per 1 million input tokens; $0.20 per 1 million cached input tokens
Output $12.00 per 1 million output tokens
OpenAI Reasoning

Complex reasoning, agentic coding, computer use, web research, scientific and professional workflows, and long-context document tasks

Context 1.05M
Reasoning 10/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current; rolling out through the OpenAI API and selected ChatGPT, Azure, and Amazon Bedrock offerings

Input $10.00 per 1 million input tokens for Standard short-context processing; $1.00 per 1 million cached input tokens; $12.50 per 1 million cache-write tokens. Long-context input is $20.00 per 1 million tokens.
Output $50.00 per 1 million output tokens for Standard short-context processing; $75.00 per 1 million output tokens for long-context processing. Batch and Flex processing are priced at 50% of Standard rates.
OpenAI Reasoning

High-volume reasoning, document analysis, coding assistance, retrieval-augmented generation, and repeatable agent workflows

Context 1.05M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current and available

Input $0.10 per 1 million input tokens; $0.01 cached input; $0.125 cache writes. Long-context rates are $0.20 input, $0.02 cached input, and $0.25 cache writes per 1 million tokens.
Output $0.50 per 1 million output tokens for short context; $0.75 per 1 million output tokens for long context.
OpenAI Reasoning

Complex coding, long-context reasoning, software engineering, research, computer use, and agentic workflows with tools.

Context 1.05M
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current; generally available through the OpenAI API

Input USD 2.00 per 1M input tokens for short context; USD 4.00 per 1M input tokens for long context. Cached input is USD 0.20 short-context or USD 0.40 long-context per 1M tokens. Cache writes are USD 2.50 short-context or USD 5.00 long-context per 1M tokens un
Output USD 10.00 per 1M output tokens for short context; USD 15.00 per 1M output tokens for long context under Standard processing.
OpenAI Multimodal
OpenAI
GPT-Audio

GPT-Audio

Audio-enabled chat applications, voice interfaces, spoken assistants, and applications requiring direct audio understanding and generation through Chat Completions.

Context 128K
Reasoning 5/10
Speed 7/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Streaming
Status

Deprecated; scheduled for shutdown on January 20, 2027

Input $2.50 per 1M text tokens; $32.00 per 1M audio tokens
Output $10.00 per 1M text tokens; $64.00 per 1M audio tokens
OpenAI Multimodal
OpenAI
GPT-Audio

GPT-Audio-1.5

Audio-in, audio-out conversational applications using the Chat Completions API, including voice assistants and tool-enabled spoken interfaces.

Context 128K
Reasoning 5/10
Speed 7/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Streaming
Status

Current; generally available

Input Text: $2.50 per 1M tokens; audio: $32.00 per 1M audio tokens
Output Text: $10.00 per 1M tokens; audio: $64.00 per 1M audio tokens
OpenAI Multimodal

Cost-sensitive, turn-based audio conversations, voice assistants, and audio-enabled applications using function calling

Context 128K
Reasoning 5/10
Speed 8/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use
Status

Deprecated; currently accessible but scheduled for API removal on 2027-01-20

Input $0.60 per 1 million text input tokens
Output $2.40 per 1 million text output tokens
OpenAI Multimodal

Precise image editing, detailed creative work, high-fidelity generation, infographics, layouts, and workflows where fewer retries matter more than minimum latency

Speed 6/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current

Input Text input: $5.00 per 1M tokens; image input: $8.00 per 1M tokens; cached text input: $1.25 per 1M tokens; cached image input: $2.00 per 1M tokens
Output Image output: $30.00 per 1M tokens
OpenAI Multimodal
OpenAI
GPT-Live

GPT-Live 1

Natural low-latency voice agents, customer support, conversational workflows, live assistance, and applications requiring interruption-aware speech interaction

Reasoning 7/10
Speed 10/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Streaming
Status

Current; available in the OpenAI API

Input $0.05 per voice-session minute, billed per second; backend model and tool usage billed separately
Output $0.05 per voice-session minute, billed per second; OpenAI documents this as the voice-session price rather than separate audio input and output token rates
OpenAI Other
OpenAI
GPT-Live-Transcribe

GPT-Live-Transcribe

Low-latency live captions, realtime call transcription, microphone streams, telephony audio, and voice-interface speech recognition

Reasoning 1/10
Speed 9/10
Outputs
Text
Capabilities
Audio input Multimodal input Streaming
Status

Current and generally available for realtime transcription

Input $0.017 per minute of realtime audio
Output Included in the per-minute realtime audio price; OpenAI does not list a separate output-token price
OpenAI Reasoning

Local and private reasoning applications, coding assistants, agentic workflows, on-device or edge inference, fine-tuning, and cost-sensitive deployments with suitable hardware.

Context 131K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Web search Fine-tuning Streaming
Status

Current open-weight model; downloadable and usable through self-hosted or third-party inference infrastructure. Not served through the OpenAI API or ChatGPT.

Input No official OpenAI API input price; self-hosting and third-party hosting costs vary.
Output No official OpenAI API output price; self-hosting and third-party hosting costs vary.
OpenAI Reasoning

Self-hosted reasoning, coding, agentic workflows, private deployments, research, and fine-tuning

Context 131K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Web search Fine-tuning Streaming
Status

Current open-weight model; downloadable and deployable locally or through third-party providers; not available through the OpenAI API

OpenAI Other
OpenAI
GPT-OSS-Safeguard

gpt-oss-safeguard-20b

Policy-based safety classification, LLM input/output filtering, content labeling, trust and safety review, and self-hosted moderation workflows

Context 131K
Reasoning 8/10
Speed 6/10
Outputs
Text
Status

Research preview; currently available as an open-weight model

OpenAI Safety
OpenAI
gpt-oss-safeguard

gpt-oss-safeguard-120b

Custom-policy safety classification, LLM input and output filtering, trust and safety labeling, nuanced moderation review, and offline safety analysis

Context 131K
Reasoning 8/10
Speed 3/10
Outputs
Text
Status

Research preview; open-weight and downloadable

OpenAI Realtime
OpenAI
GPT-Realtime

GPT-Realtime

Low-latency speech-to-speech voice agents, realtime customer support, education, accessibility, and conversational applications with function calling

Context 32K
Reasoning 5/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Multimodal input Tool use
Status

Deprecated; scheduled for API shutdown on January 20, 2027

Input Text: $4.00 per 1M tokens; cached text: $0.40 per 1M tokens; audio: $32.00 per 1M tokens; cached audio: $0.40 per 1M tokens; image: $5.00 per 1M tokens; cached image: $0.50 per 1M tokens
Output Text: $16.00 per 1M tokens; audio: $64.00 per 1M tokens
OpenAI Realtime Audio
OpenAI
GPT-Realtime

GPT-Realtime-1.5

Low-latency speech-to-speech voice agents, customer support, realtime assistants, and audio applications that need function calling.

Context 32K
Reasoning 6/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Multimodal input Tool use
Status

Active and currently available

Input $4.00 per 1M text tokens; $32.00 per 1M audio tokens; $5.00 per 1M image tokens. Cached input: $0.40 per 1M text or audio tokens and $0.50 per 1M image tokens.
Output $16.00 per 1M text tokens; $64.00 per 1M audio tokens.
OpenAI Multimodal
OpenAI
GPT-Realtime

GPT-Realtime-2

Reasoning voice agents, speech-to-speech applications, customer support, live assistants, tool-driven workflows, and long conversational sessions

Context 128K
Reasoning 9/10
Speed 7/10
Outputs
Text Speech
Capabilities
Image input Audio input Multimodal input Tool use Streaming
Status

Current

Input Text: $4.00 per 1M tokens; cached text: $0.40 per 1M; audio: $32.00 per 1M tokens; cached audio: $0.40 per 1M; image: $5.00 per 1M tokens; cached image: $0.50 per 1M
Output Text: $24.00 per 1M tokens; audio: $64.00 per 1M tokens
OpenAI Realtime
OpenAI
GPT-Realtime

GPT-Realtime-2.1

Low-latency speech-to-speech agents, customer-service voice workflows, realtime tool use, telephony, and multimodal assistants with image input

Context 128K
Reasoning 8/10
Speed 8/10
Outputs
Text Speech
Capabilities
Image input Audio input Multimodal input Tool use Streaming
Status

current

Input Text: $4.00 per 1M tokens; cached text: $0.40 per 1M; audio: $32.00 per 1M audio tokens; cached audio: $0.40 per 1M; image: $5.00 per 1M tokens; cached image: $0.50 per 1M
Output Text: $24.00 per 1M tokens; audio: $64.00 per 1M audio tokens
OpenAI Other

Low-latency spoken translation, multilingual calls, live interpretation, broadcasts, meetings, lessons, video rooms, captions, and translated audio experiences.

Context 16K
Reasoning 3/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Streaming
Status

Current

Input $0.034 per minute of realtime audio
Output $0.034 per minute of realtime audio
OpenAI Other

Low-latency live transcription, captions, meeting notes, call analysis, voice-agent input, and continuous speech-to-text workflows

Context 16K
Reasoning 1/10
Speed 9/10
Outputs
Text
Capabilities
Audio input Multimodal input Streaming
Status

Current and available through the OpenAI Realtime API for realtime transcription

Input $0.017 per minute of audio
Output Included in the audio-duration transcription price; no separate text-output price documented
OpenAI Realtime
OpenAI
GPT-Realtime

GPT-Realtime Mini

Cost-sensitive realtime voice agents, speech-to-speech applications, interactive assistants, and multimodal interfaces

Context 32K
Reasoning 5/10
Speed 8/10
Outputs
Text Speech
Capabilities
Image input Audio input Multimodal input Tool use
Status

Deprecated; currently accessible but scheduled for API shutdown on 2027-01-20

Input $0.60 per 1M text input tokens; $0.06 per 1M cached text input tokens
Output $2.40 per 1M text output tokens
OpenAI Lightweight
OpenAI
GPT-Realtime-2.1

GPT-Realtime-2.1 Mini

Lower-cost, low-latency realtime voice agents, speech-to-speech assistants, and tool-enabled conversational applications

Context 128K
Reasoning 7/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Multimodal input Tool use
Status

Current

Input Text: $0.60 per 1M tokens; cached text: $0.06 per 1M; audio: $10.00 per 1M tokens; cached audio: $0.30 per 1M; image: $0.80 per 1M tokens; cached image: $0.08 per 1M
Output Text: $2.40 per 1M tokens; audio: $20.00 per 1M tokens
OpenAI Reasoning
OpenAI
GPT-Rosalind

GPT-Rosalind

Governed biology, genomics, medicinal chemistry, protein analysis, drug discovery, literature synthesis, wet-lab troubleshooting, and scientific tool workflows

Reasoning 9/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use
Status

Generally available to eligible organizations through the trusted-access program; approved internal life sciences research only

Input $5 per 1M input tokens; $0.50 per 1M cached input tokens
Output $25 per 1M output tokens
OpenAI Other
OpenAI
GPT-Transcribe

GPT-Transcribe

High-accuracy transcription of recorded audio, streamed file transcripts, multilingual recordings, and domain-specific speech with keyword or language hints

Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Audio input Multimodal input Streaming
Status

Current

Input $0.0045 per audio minute
Output No separate output-token price; included in the per-minute transcription price
OpenAI Multimodal

Existing ChatGPT image-generation and image-editing integrations

Reasoning 1/10
Speed 7/10
Outputs
Text Image
Capabilities
Image input Multimodal input
Status

Deprecated; currently accessible; scheduled for shutdown on 2026-12-01

Input Text: $5.00 per 1M tokens; cached text: $1.25 per 1M tokens; image: $8.00 per 1M tokens; cached image: $2.00 per 1M tokens. Per-image generation: low $0.009-$0.013, medium $0.034-$0.05, high $0.133-$0.20 depending on size.
Output Text: $10.00 per 1M tokens; image: $32.00 per 1M tokens. Per-image generation: low $0.009-$0.013, medium $0.034-$0.05, high $0.133-$0.20 depending on size.
OpenAI Image Generation
OpenAI
GPT Image

GPT-Image-1

API-based image generation, image editing, reference-image workflows, inpainting, marketing assets, e-commerce imagery, and visual content production

Speed 6/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Deprecated; currently accessible and scheduled to shut down on 2026-10-23

Input Text input: $5.00 per 1M tokens; cached text input: $1.25 per 1M tokens; image input: $10.00 per 1M image tokens; cached image input: $2.50 per 1M image tokens
Output Image generation per image: low $0.011 at 1024x1024 or $0.016 at 1024x1536 and 1536x1024; medium $0.042 or $0.063; high $0.167 or $0.25. Image output tokens: $40.00 per 1M tokens.
OpenAI Other
OpenAI
GPT Image

GPT-Image-1.5

Production image generation, image editing, branded graphics, ecommerce product imagery, marketing assets, and workflows requiring preservation of important visual details

Reasoning 1/10
Speed 7/10
Outputs
Text Image
Capabilities
Image input Multimodal input
Status

Deprecated; currently accessible with API shutdown scheduled for 2026-12-01

Input $5.00 per 1M text tokens; $8.00 per 1M image tokens; cached input $1.25 per 1M text tokens and $2.00 per 1M image tokens
Output $10.00 per 1M text tokens; $32.00 per 1M image tokens; image generation $0.009-$0.20 per image depending on quality and resolution
OpenAI Multimodal
OpenAI
GPT Image

GPT-Image-2

High-quality text-to-image generation, reference-based image editing, text-heavy visual assets, product imagery, marketing creatives, and production design workflows.

Speed 8/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Active; GPT Image 2.5 models are available for newer workflows, but GPT-Image-2 remains accessible as a documented API model.

Input $8.00 per 1M image input tokens; $2.00 per 1M cached image input tokens; $5.00 per 1M text input tokens; $1.25 per 1M cached text input tokens
Output $30.00 per 1M image output tokens; $10.00 per 1M text output tokens where applicable
OpenAI Multimodal
OpenAI
GPT Image 1

GPT-Image-1 Mini

Cost-sensitive image generation and editing, high-volume variations, rapid ideation, previews, lightweight personalization, and draft creative assets.

Speed 8/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Deprecated; currently accessible but scheduled for API shutdown on 2026-12-01.

Input Text input: $2.00 per 1M tokens; cached text input: $0.20 per 1M tokens. Image input: $2.50 per 1M image tokens; cached image input: $0.25 per 1M image tokens.
Output Image output: $8.00 per 1M image tokens. Per-image generation: $0.005-$0.036 at 1024x1024 and $0.006-$0.052 at 1024x1536 or 1536x1024, depending on quality.
OpenAI Multimodal
OpenAI
GPT Image 2.5

GPT-Image-2.5 Flare

Fast, high-quality image generation and editing, creator content, product experiences, visual search, rapid prototyping, and high-volume workflows

Speed 9/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current; available through the OpenAI API

Input Text input: $5.00 per 1M tokens; cached text input: $1.25 per 1M tokens; image input: $8.00 per 1M image tokens; cached image input: $2.00 per 1M image tokens
Output Image output: $30.00 per 1M image tokens; text output is not billed because the model outputs images
OpenAI Reasoning
OpenAI
o-series

o3

Complex reasoning, advanced coding, mathematics, science, technical research, visual analysis and multi-step tool workflows

Context 200K
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current canonical alias; o3-2025-04-16 snapshot deprecated and scheduled for API shutdown on December 11, 2026

Input $2.00 per 1M input tokens; $0.50 per 1M cached input tokens. Batch: $1.00 input and $0.25 cached input per 1M tokens.
Output $8.00 per 1M output tokens. Batch: $4.00 per 1M output tokens.
OpenAI Reasoning
OpenAI
o1

o1

Complex reasoning, mathematics, science, coding analysis, visual reasoning, and high-accuracy multi-step tasks

Context 200K
Reasoning 9/10
Speed 4/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Deprecated; still documented in the OpenAI API model catalog

Input $15.00 per 1M input tokens; $7.50 per 1M cached input tokens
Output $60.00 per 1M output tokens
OpenAI Reasoning

Historically, difficult mathematics, science, coding, and other multi-step reasoning tasks requiring extended deliberation

Context 128K
Reasoning 9/10
Speed 4/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Retired; API access shut down on 2025-07-28

Input $15.00 per 1 million input tokens; cached input $7.50 per 1 million tokens
Output $60.00 per 1 million output tokens
OpenAI Reasoning

Cost-sensitive mathematics, science, algorithmic programming, debugging, and text-only reasoning

Context 128K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Streaming
Status

Deprecated

Input $1.10 per 1M input tokens; $0.55 per 1M cached input tokens
Output $4.40 per 1M output tokens
OpenAI Reasoning

Complex reasoning, difficult technical analysis, advanced programming, research workflows, and tasks where answer consistency matters more than latency or cost.

Context 200K
Reasoning 9/10
Speed 3/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use
Status

Deprecated in OpenAI's current model catalog; the dated snapshot o1-pro-2025-03-19 is also marked deprecated. No exact shutdown date for the canonical o1-pro alias was found in the reviewed official documentation.

Input $150 per 1 million input tokens
Output $600 per 1 million output tokens
OpenAI Reasoning

Complex multi-step research, source synthesis, legal and scientific analysis, market research, and large-scale internal-data investigation

Context 200K
Reasoning 10/10
Speed 3/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Deprecated

Input $10.00 per 1M input tokens; $2.50 per 1M cached input tokens
Output $40.00 per 1M output tokens
OpenAI Reasoning

Coding, mathematics, science, technical analysis, structured extraction, text-to-SQL, and multi-step reasoning

Context 200K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current canonical alias with deprecated snapshot; o3-mini-2025-01-31 is scheduled for API shutdown on 2026-10-23

Input $1.10 per 1M input tokens; $0.55 per 1M cached input tokens
Output $4.40 per 1M output tokens
OpenAI Reasoning

High-reliability reasoning, advanced mathematics, scientific analysis, complex coding, research, and multi-step professional work.

Context 200K
Reasoning 10/10
Speed 3/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use
Status

Current canonical alias; the dated snapshot o3-pro-2025-06-10 is marked deprecated in the model documentation.

Input $20 per 1 million input tokens
Output $80 per 1 million output tokens
OpenAI Reasoning

Fast, cost-sensitive reasoning; coding; mathematics; visual analysis; structured extraction; high-volume tool-using agents

Context 200K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Fine-tuning
Status

Deprecated; currently available through the API; scheduled for shutdown on 2026-10-23

Input $1.10 per 1 million input tokens; $0.275 per 1 million cached input tokens
Output $4.40 per 1 million output tokens
OpenAI Reasoning

Complex multi-step research, source synthesis, market analysis, legal or scientific research, and long-form evidence-based reports.

Context 200K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current canonical alias; the dated snapshot o4-mini-deep-research-2025-06-26 is deprecated.

Input $2.00 per 1M input tokens; $0.50 per 1M cached input tokens
Output $8.00 per 1M output tokens
OpenAI Moderation
OpenAI
omni-moderation

omni-moderation

Text and image safety classification, content filtering, AI-output screening, policy enforcement, and human-review routing

Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Current

Input Free
Output Free
OpenAI Other
OpenAI
omni-moderation

omni-moderation-latest

Text and image safety classification, content filtering, moderation queues, policy enforcement, and generated-content screening

Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Current default moderation model

Input Free through the Moderation API
Output Free through the Moderation API
OpenAI Multimodal
OpenAI
Sora 2

Sora 2

Rapid video concepting, social clips, image-to-video experiments, prototypes, rough cuts, and audiovisual creative iteration

Reasoning 1/10
Speed 8/10
Outputs
Video Speech
Capabilities
Image input Multimodal input
Status

Deprecated; currently accessible through the API as of September 23, 2026, with shutdown scheduled for September 24, 2026

Input Not token-priced; image and text inputs are included in video-generation requests
Output $0.10 per generated video second for 720x1280 portrait or 1280x720 landscape output
OpenAI Video Generation

Production-quality text-to-video and image-guided video generation, cinematic prototypes, marketing assets, and high-resolution short clips with synchronized audio.

Outputs
Video Speech
Capabilities
Image input Multimodal input
Status

Legacy; deprecated; currently accessible through September 23, 2026; scheduled for API shutdown on September 24, 2026

Input $0.30 per second at 720p; $0.50 per second at 1024p; $0.70 per second at 1080p. Batch pricing: $0.15, $0.25, and $0.35 per second respectively.
Output Video with synchronized audio; pricing is charged per generated second rather than per text or audio token.
OpenAI Embedding
OpenAI
text-embedding-3

text-embedding-3-large

High-quality semantic search, multilingual retrieval, RAG, recommendations, clustering, classification and similarity matching

Context 8K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Status

Current and available through the OpenAI API

Input $0.13 per 1 million input tokens
Output No separate output-token charge; the model returns embedding vectors
OpenAI Embedding
OpenAI
text-embedding-3

text-embedding-3-small

Cost-efficient semantic search, retrieval-augmented generation, clustering, recommendations, anomaly detection, and text or code similarity

Context 8K
Speed 9/10
Outputs
Embeddings
Status

Current

Input $0.02 per 1 million input tokens
Output Not applicable; embedding output is billed through input-token usage
OpenAI Embedding
OpenAI
text-embedding-ada-002

text-embedding-ada-002

Legacy semantic search, retrieval, clustering, recommendations, anomaly detection, and classification systems already built around ada-002 vectors

Context 8K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Status

Older embedding model; currently listed and accessible through the embeddings API

Input $0.10 per 1 million input tokens
OpenAI Other
OpenAI
TTS-1

TTS-1

Low-latency text-to-speech, realtime-oriented voice interfaces, narration, accessibility, and automated audio generation

Reasoning 1/10
Speed 9/10
Outputs
Speech
Capabilities
Streaming
Status

Current and accessible; optimized for low-latency text-to-speech

Input $15.00 per 1 million characters
OpenAI Other

High-quality text-to-speech generation, narration, accessibility audio, voice interfaces, and downloadable speech content

Reasoning 1/10
Speed 7/10
Outputs
Speech
Capabilities
Streaming
Status

Current; available through the OpenAI Audio API speech endpoint

Input $30 per 1 million characters
Output Included in speech-generation pricing; output is billed by input characters rather than output tokens
OpenAI Other
OpenAI
Whisper

Whisper

Multilingual audio transcription, English speech translation, language identification, subtitles, captions, and word- or segment-level timestamps.

Speed 8/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Deprecated; currently accessible through the OpenAI API and scheduled for shutdown on February 26, 2027.

Input $0.006 per minute of audio
Qwen Multimodal

Cost-sensitive cross-modal retrieval, image and video search, multimedia catalog indexing, and vector search

Context 1K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Capabilities
Image input Video input Multimodal input
Status

Current and available through Alibaba Cloud Model Studio International deployment

Input Image/video: $0.03 per 1 million input tokens; text: $0.09 per 1 million input tokens
Output Free; embedding output is not charged
Qwen Multimodal Embedding
Qwen
Multimodal Embedding

multimodal-embedding-v1

Cross-modal retrieval, text-to-image search, image similarity, video search, semantic classification, clustering, and multimodal vector indexing.

Context 512
Reasoning 0/10
Speed 0/10
Outputs
Embeddings
Capabilities
Image input Video input Multimodal input
Status

Current and accessible through Alibaba Cloud Model Studio in the China (Beijing) region; free trial pricing is listed.

Input Free trial
Qwen Other

Agent-environment simulation, tool-interaction modeling, terminal and software-engineering trajectories, and research on language world models

Context 262K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current open-weight model

Qwen Multimodal

Low-latency duplex voice assistants, real-time customer service, AI companions, and streamed speech-to-speech applications

Context 41K
Reasoning 6/10
Speed 8/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Web search Streaming
Status

Current and accessible; standard-edition real-time duplex speech model

Input Singapore: text $0.80/M tokens; audio $6.40/M tokens. China (Beijing): text $0.688/M tokens; audio $5.501/M tokens.
Output Singapore: text $6.40/M tokens; audio $24/M tokens. China (Beijing): text $5.501/M tokens; audio $20.628/M tokens.
Qwen Multimodal

Low-latency voice assistants, customer service, AI companions, full-duplex spoken interaction, and voice applications using tools or cloned voices.

Context 262K
Reasoning 4/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Web search Streaming
Status

Current and available

Input Singapore: $0.80 per 1M text-input tokens; $6.40 per 1M audio-input tokens. China (Beijing): $0.688 per 1M text-input tokens; $5.501 per 1M audio-input tokens.
Output Singapore: $6.40 per 1M text-output tokens; $24 per 1M audio-output tokens. China (Beijing): $5.501 per 1M text-output tokens; $20.628 per 1M audio-output tokens.
Qwen Multimodal
Qwen
Qwen-Audio-3.0-Realtime

qwen-audio-3.0-realtime-flash

Low-latency voice assistants, real-time customer service, duplex speech conversations, interactive voice agents, and applications requiring streaming audio responses.

Context 41K
Reasoning 6/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Web search Streaming
Status

Available; current API access confirmed, but newer models are recommended for some new projects

Input International/Singapore: text input $0.23 per 1M tokens; audio input $0.93 per 1M tokens. China (Beijing): text input $0.413 per 1M tokens; audio input $4.126 per 1M tokens.
Output International/Singapore: text-only output $0.70 per 1M tokens; combined text and audio output $1.87 per 1M tokens. China (Beijing): text-only output $4.126 per 1M tokens; combined text and audio output $13.752 per 1M tokens.
Qwen Speech Recognition

Real-time multilingual speech transcription, live captions, meeting transcription, voice interfaces, streaming subtitles, and Chinese-dialect recognition.

Context 8K
Reasoning 1/10
Speed 9/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current and available through Alibaba Cloud Model Studio in the International/Singapore and China (Beijing) regions.

Input International/Singapore: USD 0.93 per 1 million input tokens; China (Beijing): USD 0.848 per 1 million input tokens.
Output International/Singapore: USD 0.70 per 1 million output tokens; China (Beijing): USD 0.636 per 1 million output tokens.
Qwen Other

Expressive text-to-speech, audiobooks, film and video dubbing, content creation, premium voice services, multilingual speech, dialect synthesis, and voice cloning

Reasoning 1/10
Speed 8/10
Outputs
Speech
Capabilities
Streaming
Status

Current and available

Input USD 0.20 per 10,000 input characters in Singapore/International; USD 0.19253 per 10,000 input characters in China (Beijing)
Output Not separately billed; pricing is based on input characters
Qwen Other

Long-form offline transcription of meetings, interviews, calls, media files, and multilingual or dialect-rich recordings

Context 8K
Reasoning 2/10
Speed 7/10
Outputs
Text
Capabilities
Audio input
Status

Current and publicly available

Input USD 0.15 per 1 million input tokens in Singapore; USD 0.113 per 1 million input tokens in China (Beijing)
Output USD 0.47 per 1 million output tokens in Singapore; USD 0.382 per 1 million output tokens in China (Beijing)
Qwen Multimodal
Qwen
Qwen-Drive

Qwen-Drive-1.0

Autonomous-driving research, driving-scene VQA, 3D BEV perception, trajectory prediction, and embodied-AI experimentation

Reasoning 6/10
Speed 5/10
Outputs
Text Actions
Capabilities
Image input Multimodal input Streaming
Status

Current open-weight research release

Qwen Image Generation
Qwen
Qwen-Image

qwen-image-2.0

Text-to-image generation, image editing, text rendering in images, photorealistic scenes, creative design, and producing multiple image variants.

Reasoning 3/10
Speed 8/10
Outputs
Image
Capabilities
Image input Multimodal input Fine-tuning
Status

Current; accelerated model; functionally equivalent to qwen-image-2.0-2026-03-03

Input $0 per input image; only generated output images are billed
Output $0.035 per image internationally; $0.028671 per image in China (Beijing) and Singapore
Qwen Image Generation

Cost-sensitive text-to-image generation, posters, marketing graphics, illustrations, and images containing Chinese or English text

Speed 8/10
Outputs
Image
Status

Current and accessible; currently equivalent to qwen-image

Input Not applicable; image generation is billed per output image
Output $0.03/image internationally; $0.028671/image in China (Beijing)
Qwen Other
Qwen
Qwen-Image-Edit

qwen-image-edit

Natural-language single-image editing, bilingual text changes, object insertion or removal, style transfer, pose changes, and image fusion

Reasoning 2/10
Speed 6/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and accessible

Input $0.045 per image in Singapore for international deployment
Qwen Other
Qwen
Qwen-Image-Edit

qwen-image-edit-max

High-quality image editing, multi-image composition, industrial design concepts, geometric transformations, character-consistent edits, and controlled visual revisions

Reasoning 2/10
Speed 6/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and available; canonical model ID is functionally equivalent to qwen-image-edit-max-2026-01-16

Input $0.075 per output image in the International/Singapore region; $0.071677 per output image in China (Beijing)
Output $0.075 per output image in the International/Singapore region; $0.071677 per output image in China (Beijing)
Qwen Other
Qwen
Qwen-Image 2.0

qwen-image-2.0-pro

Professional text-to-image generation, image editing, posters, infographics, multilingual in-image text, photorealistic scenes, and reference-based creative production

Reasoning 6/10
Speed 5/10
Outputs
Image
Capabilities
Image input Multimodal input Fine-tuning
Status

Current rolling model; functionally equivalent to qwen-image-2.0-pro-2026-04-22

Output $0.075 per image internationally; $0.071676 per image in China (Beijing)
Qwen General Purpose
Qwen
Qwen-Plus

Qwen-Plus

General-purpose text generation, long-context analysis, multilingual applications, structured business workflows, function-calling agents, and applications that need optional reasoning mode.

Context 1M
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Current; qwen-plus currently resolves to the qwen-plus-2025-12-01 snapshot

Input $0.40 per 1M input tokens for 0–256K input; $1.20 per 1M input tokens above 256K in the International Singapore pricing tier. International Global pricing is $0.115 per 1M tokens for up to 128K, $0.345 for 128K–256K, and $0.689 for 256K–1M.
Output $1.20 per 1M non-thinking output tokens and $4.00 per 1M thinking output tokens for 0–256K input; $3.60 non-thinking and $12.00 thinking per 1M output tokens above 256K in the International Singapore pricing tier. International Global pricing differs by r
Qwen Multimodal

Complex image and video understanding, document analysis, chart interpretation, visual question answering, and structured extraction

Context 131K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Streaming
Status

Currently accessible legacy visual language model; current qwen-vl-max endpoint is functionally equivalent to qwen-vl-max-2025-08-13

Input $0.229 per 1M tokens in China (Beijing) and Singapore; $0.80 per 1M tokens for international deployment
Output $0.573 per 1M tokens in China (Beijing) and Singapore; $3.20 per 1M tokens for international deployment
Qwen General Purpose

Self-hosted chat assistants, multilingual text generation, coding and mathematics assistance, long-context document work, structured text generation, and cost-sensitive private deployments

Context 131K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight model; publicly available for download and self-hosted deployment

Qwen General Purpose

Self-hosted multilingual assistants, document processing, RAG, coding support, structured extraction, and cost-conscious production deployments.

Context 131K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Legacy open-weight model; downloadable and self-hostable, while Qwen2.5 API models marked deprecated are no longer callable through current Alibaba Cloud Model Studio pricing documentation.

Input Not applicable for the downloadable checkpoint; no current official per-token hosted price verified for this exact model.
Output Not applicable for the downloadable checkpoint; no current official per-token hosted price verified for this exact model.
Qwen General Purpose

Self-hosted assistants, multilingual text generation, long-document processing, coding support, structured extraction, RAG, and agent applications

Context 131K
Reasoning 7/10
Speed 5/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available; open-weight model

Qwen General Purpose

Self-hosted multilingual assistants, coding, mathematics, document analysis, structured text generation, and long-context workloads

Context 131K
Reasoning 8/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as an open-weight model; legacy relative to newer Qwen generations

Qwen Multimodal
Qwen
Qwen2.5-Omni

Qwen2.5-Omni-7B

Multimodal assistants, audio and video understanding, visual question answering, voice interaction, speech instruction following, and local multimodal AI research

Context 33K
Reasoning 7/10
Speed 7/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Streaming
Status

Current; open-weight model and available through Alibaba Cloud Model Studio

Input International: $0.10 per 1M text tokens; $6.76 per 1M audio tokens; $0.28 per 1M image/video tokens. China Beijing: $0.087 per 1M text tokens; $5.448 per 1M audio tokens; $0.287 per 1M image/video tokens.
Output International: $0.40 per 1M tokens for text-only input; $0.84 per 1M tokens for text after multimodal input; $13.51 per 1M tokens for text-and-audio output. China Beijing: $0.345 per 1M tokens for text-only input; $0.861 per 1M tokens for text after multi
Qwen Multimodal

Local or self-hosted image and video understanding, OCR, document extraction, chart and diagram analysis, visual question answering, visual grounding, and multimodal research

Context 33K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Fine-tuning Streaming
Status

Open-weight and downloadable; supported as a fine-tuning base model in Alibaba Cloud Model Studio; current hosted inference availability and standard pricing for this exact model are not clearly listed in the latest Model Studio inference-pricing catalog.

Qwen Multimodal

High-quality image, document, chart, screenshot, OCR, visual-grounding, and video analysis; multimodal agents and self-hosted experimentation

Context 131K
Reasoning 8/10
Speed 3/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Fine-tuning
Status

Available open-weight model; older Qwen2.5-VL generation and still listed by Alibaba Cloud Model Studio

Qwen Lightweight

Fast, high-volume text generation; long-context analysis; summarization; extraction; structured outputs; and applications needing optional reasoning.

Context 1M
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Current and available; canonical qwen-flash identifier is functionally equivalent to qwen-flash-2025-07-28

Input Tiered per 1M input tokens. China Beijing: CNY 0.15 up to 128K input, CNY 0.60 above 128K to 256K, CNY 1.20 above 256K to 1M. International Singapore: CNY 0.367 up to 256K, CNY 1.835 above 256K to 1M. Regional pricing and promotional offers may vary.
Output Tiered per 1M output tokens. China Beijing: CNY 1.50 up to 128K input, CNY 6 above 128K to 256K, CNY 12 above 256K to 1M. International Singapore: CNY 2.936 up to 256K, CNY 14.678 above 256K to 1M. Regional pricing and promotional offers may vary.
Qwen Lightweight

Lightweight local assistants, offline prototypes, embedded experimentation, education, simple text generation, and resource-constrained deployments

Context 33K
Reasoning 3/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight model

Qwen Lightweight

Efficient local inference, edge applications, multilingual chat, lightweight reasoning, coding assistance, and tool-enabled agents

Context 33K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight model

Input No official hosted API price; open-weight model
Output No official hosted API price; open-weight model
Qwen General Purpose

Local and self-hosted chat, compact reasoning, coding assistance, multilingual applications, retrieval-augmented generation, and lightweight tool-using agents

Context 33K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight model

Qwen General Purpose

Cost-efficient reasoning, coding, multilingual assistants, local deployment, structured text generation, tool-enabled agents, and fine-tuned applications

Context 131K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current and accessible; open-weight model available for self-hosting and Alibaba Cloud Model Studio API deployment

Input $0.072 per 1 million tokens for Global deployment; $0.18 per 1 million tokens for International deployment
Output $0.287 per 1 million tokens for non-thinking mode and $0.717 per 1 million tokens for thinking mode in Global deployment; International pricing is $0.70 non-thinking and $2.10 thinking per 1 million tokens
Qwen General Purpose

Local deployment, multilingual assistants, reasoning, mathematics, coding, structured text generation, research, and cost-sensitive agent workflows.

Context 131K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight model; also available as qwen3-14b through Alibaba Cloud Model Studio, with regional capability and pricing differences.

Input Alibaba Cloud Model Studio: China (Beijing) $0.144 per 1 million input tokens; Singapore $0.35 per 1 million input tokens. Regional prices and deployment terms may vary.
Output Alibaba Cloud Model Studio: China (Beijing) $0.574 per 1 million output tokens in non-thinking mode and $1.434 per 1 million output tokens in thinking mode; Singapore $1.4 per 1 million output tokens in non-thinking mode and $4.2 per 1 million output toke
Qwen Reasoning

Self-hosted assistants, coding, mathematical and logical reasoning, multilingual applications, long-context document processing, and agentic tool-use systems

Context 131K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current and accessible; open-weight release with hosted Alibaba Cloud Model Studio availability

Input US$0.108 per 1 million input tokens in several global deployments; US$0.20 per 1 million input tokens for the international Singapore deployment
Output US$0.431 per 1 million output tokens in several global deployments; US$0.80 per 1 million output tokens for the international Singapore deployment. Thinking output is separately priced at higher rates where exposed.
Qwen Reasoning

Self-hosted reasoning assistants, coding agents, mathematics, multilingual applications, structured text generation, and tool-calling workflows

Context 256K
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current; open-weight model and available through Alibaba Cloud Model Studio

Input $0.16 per 1 million tokens in international Singapore, Germany Frankfurt, and US Virginia deployments; $0.287 per 1 million tokens in China Beijing
Output $0.64 per 1 million tokens in international Singapore, Germany Frankfurt, and US Virginia deployments; China Beijing: $1.147 per 1 million non-thinking tokens or $2.868 per 1 million thinking tokens
Qwen Reasoning

Complex reasoning, mathematics, software development, multilingual applications, function calling, agentic workflows, research and self-hosted open-weight deployment

Context 131K
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current and accessible through Alibaba Cloud Model Studio; original Qwen3 open-weight release, with newer 2507 instruct and thinking variants available separately

Input $0.287 per 1M tokens for standard input in the United States, Germany and China; $0.700 per 1M tokens in Singapore. Thinking-mode input is priced the same.
Output $1.147 per 1M tokens for standard output in the United States, Germany and China; $2.868 per 1M tokens for thinking-mode output. Singapore pricing is $2.800 standard output and $8.400 thinking-mode output per 1M tokens.
Qwen Other
Qwen
Qwen3 Reranker

qwen3-rerank

Multilingual semantic search, RAG candidate reranking, enterprise document retrieval, knowledge-base search, and improving search-result relevance.

Context 4K
Reasoning 2/10
Speed 8/10
Status

Current and available

Input $0.10 per 1 million input tokens for the international Singapore deployment; pricing is region-dependent. Output is free.
Output Free
Qwen Other

Low-cost multilingual speech-to-text, language identification, offline transcription, real-time streaming ASR, and high-throughput deployments

Context 66K
Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Audio input Multimodal input Streaming
Status

Current, open-weight, Apache 2.0 licensed

Qwen Other

Multilingual speech transcription, language identification, long-audio processing, and self-hosted or streaming ASR applications

Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Audio input Multimodal input Fine-tuning Streaming
Status

Current; open-weight and downloadable

Input No official hosted API price identified for this exact open-weight checkpoint
Output No separate output charge for the self-hosted checkpoint
Qwen Coding
Qwen
Qwen3-Coder

Qwen3-Coder-Next

Repository-scale coding agents, code generation, code completion, debugging, refactoring, terminal workflows, and cost-sensitive self-hosted deployments.

Context 262K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current; open-weight model and available through Alibaba Cloud Model Studio

Input USD 0.144 per 1M tokens for China Beijing input up to 32K; USD 0.216 for 32K–128K; USD 0.359 for 128K–256K. International Singapore and Frankfurt pricing is USD 0.30, USD 0.50, and USD 0.80 per 1M input tokens across the same bands.
Output USD 0.574 per 1M tokens for China Beijing output up to 32K; USD 0.861 for 32K–128K; USD 1.434 for 128K–256K. International Singapore and Frankfurt pricing is USD 1.50, USD 2.50, and USD 4.00 per 1M output tokens across the same bands.
Qwen Coding
Qwen
Qwen3-Coder

Qwen3-Coder-Plus

Large-codebase analysis, code generation, refactoring, debugging, documentation and long-context coding-agent workflows

Context 1M
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current; canonical qwen3-coder-plus currently equivalent to qwen3-coder-plus-2025-09-23

Input US Virginia: $0.574 per 1M tokens up to 32K input; $0.861 for 32K-128K; $1.434 for 128K-256K; $2.868 for 256K-1M. Regional pricing varies.
Output US Virginia: $2.294 per 1M tokens up to 32K input; $3.441 for 32K-128K; $5.735 for 128K-256K; $28.671 for 256K-1M. Regional pricing varies.
Qwen Other
Qwen
Qwen3-LiveTranslate

Qwen3-LiveTranslate-Flash

Streaming translation of recorded or uploaded audio and video, multilingual subtitles, translated voice tracks, and applications requiring translated text or synthesized speech.

Context 53K
Reasoning 2/10
Speed 8/10
Outputs
Text Speech
Capabilities
Audio input Video input Multimodal input Streaming
Status

Current stable model

Input Audio input and output are billed at 12.5 tokens per second, with audio shorter than one second billed as one second. Video usage additionally consumes video tokens based on sampled frames and resolution. Monetary rates depend on the applicable Alibaba Cl
Output Audio output is billed at 12.5 tokens per second. Text output uses the applicable Model Studio token rate when charged separately. No standalone fixed monetary rate was verified for this exact model in the model documentation.
Qwen Other

Real-time multilingual speech interpretation, live voice translation, conference translation, streaming media, and audiovisual translation with text or synthesized speech output

Context 53K
Reasoning 1/10
Speed 8/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Streaming
Status

Legacy; still available; no longer recommended for new use

Input China (Beijing): CNY 64 per 1M input audio tokens and CNY 8 per 1M input image tokens. Singapore: CNY 73.392 per 1M input audio tokens and CNY 9.541 per 1M input image tokens.
Output China (Beijing): CNY 64 per 1M output text tokens and CNY 240 per 1M output audio tokens. Singapore: CNY 73.392 per 1M output text tokens and CNY 278.891 per 1M output audio tokens.
Qwen Reasoning
Qwen
Qwen3-Max

Qwen3-Max

Complex reasoning, coding assistance, web-grounded agents, function calling, structured extraction, and long-context text analysis

Context 262K
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Current

Input International: $1.20 per 1M input tokens up to 32K; $2.40 per 1M above 32K to 128K; $3.00 per 1M above 128K to 256K. Regional prices vary.
Output International: $6.00 per 1M output tokens up to 32K; $12.00 per 1M above 32K to 128K; $15.00 per 1M above 128K to 256K. Regional prices vary.
Qwen Multimodal

Local visual assistants, OCR, document and chart analysis, image question answering, lightweight video understanding, and multimodal prototyping

Context 256K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Fine-tuning
Status

Current open-weight model

Qwen Multimodal

Local image and video understanding, OCR, document extraction, visual question answering, visual coding, and lightweight multimodal agents

Context 262K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Fine-tuning
Status

Current open-weight model; available on Hugging Face and supported for supervised fine-tuning in Alibaba Cloud Model Studio

Qwen Multimodal

Local or hosted image and video understanding, OCR, document extraction, visual question answering, spatial reasoning, screenshot analysis, multimodal agents, and structured data extraction

Context 262K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Fine-tuning
Status

Current; open-weight model with hosted inference available through Alibaba Cloud Model Studio

Input $0.072 per 1 million input tokens
Output $0.287 per 1 million output tokens
Qwen Multimodal

Image and video understanding, OCR, document analysis, spatial reasoning, visual coding, long-context multimodal tasks, and visual-agent applications

Context 256K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use
Status

Current and available; open-weight release and Alibaba Cloud Model Studio API model

Input $0.20 per 1 million tokens for the listed international deployment; regional pricing may differ
Output $0.80 per 1 million tokens for the listed international deployment; regional pricing may differ
Qwen Multimodal

Document intelligence, OCR, image and video understanding, spatial reasoning, visual coding, and visual-agent applications

Context 131K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use
Status

Current; open-weight checkpoint and available through Alibaba Cloud Model Studio managed inference

Input $0.16 per 1 million input tokens in Alibaba Cloud Model Studio US (Virginia) global deployment; China Beijing pricing is $0.287 per 1 million input tokens
Output $0.64 per 1 million output tokens in Alibaba Cloud Model Studio US (Virginia) global deployment; China Beijing pricing is $1.147 per 1 million output tokens
Qwen Multimodal

High-quality image and video understanding, OCR, document intelligence, visual coding, spatial reasoning, long-context multimodal analysis and visual-agent applications

Context 131K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Current and accessible; open-weight release and Alibaba Cloud Model Studio API availability

Input $0.287 per 1 million tokens in the US Virginia global deployment; $0.400 per 1 million tokens in Singapore
Output $1.147 per 1 million tokens in the US Virginia global deployment; $1.600 per 1 million tokens in Singapore
Qwen Multimodal

Long-context multimodal analysis, document understanding, video and image interpretation, general reasoning, coding, tool-enabled assistants, and self-hosted deployment

Context 262K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Accessible; no longer recommended for new projects

Input $0.086 per 1M input tokens for requests up to 128K input tokens in US Virginia; $0.258 per 1M input tokens for 128K-256K requests
Output $0.688 per 1M output tokens for requests up to 128K input tokens in US Virginia; $2.064 per 1M output tokens for 128K-256K requests
Qwen Multimodal

Efficient multimodal assistants, coding, reasoning, long-context analysis, local deployment, and tool-using agents

Context 262K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Available; open-weight Apache 2.0 model with hosted API access through Alibaba Cloud Model Studio

Input USD 0.057 per 1 million input tokens for up to 128K input; USD 0.229 per 1 million input tokens for 128K–256K input in the Global deployment scope. International flat pricing is USD 0.25 per 1 million input tokens.
Output USD 0.459 per 1 million output tokens for up to 128K input; USD 1.835 per 1 million output tokens for 128K–256K input in the Global deployment scope. International flat pricing is USD 2 per 1 million output tokens.
Qwen Multimodal

Advanced multimodal reasoning, image and video understanding, coding, long-context analysis, document and chart interpretation, function-calling agents, and web-grounded workflows.

Context 262K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current and available through Alibaba Cloud Model Studio; released globally on February 24, 2026.

Input $0.115 per 1M tokens for input up to 128K; $0.287 per 1M tokens for input from 128K to 256K. Pricing varies by deployment scope and region.
Output $0.917 per 1M tokens for requests up to 128K input; $2.294 per 1M tokens for requests from 128K to 256K input. Pricing varies by deployment scope and region.
Qwen Multimodal

Advanced multimodal reasoning, coding, video and image understanding, long-context analysis, tool-using agents, and self-hosted open-weight deployments

Context 262K
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current; open-weight model and available through Alibaba Cloud Model Studio

Input $0.172 per 1M tokens for input up to 128K; $0.43 per 1M tokens for input above 128K and up to 256K in Beijing, Frankfurt, and Virginia. Singapore: $0.60 per 1M tokens.
Output $1.032 per 1M tokens for input up to 128K; $2.58 per 1M tokens for input above 128K and up to 256K in Beijing, Frankfurt, and Virginia. Singapore: $3.60 per 1M tokens.
Qwen Multimodal

Fast long-context text, image and video understanding; structured extraction; tool-enabled agents; web-grounded applications; and high-volume multimodal workloads.

Context 1M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current and available

Input $0.029 per 1M tokens for 0–128K input; $0.115 per 1M tokens for 128K–256K input; $0.172 per 1M tokens for 256K–1M input in US Virginia/Global pricing
Output $0.287 per 1M tokens for 0–128K input; $1.147 per 1M tokens for 128K–256K input; $1.72 per 1M tokens for 256K–1M input in US Virginia/Global pricing
Qwen Multimodal

Long-context reasoning, multimodal document and video analysis, coding, structured enterprise automation, function-calling agents, and web-grounded research.

Context 1M
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current; canonical rolling model identifier currently equivalent to qwen3.5-plus-2026-02-15

Input International: $0.40 per 1M input tokens for input up to 256K; $0.50 per 1M input tokens for input above 256K and up to 1M. Regional and global pricing varies.
Output International: $2.40 per 1M output tokens for input up to 256K; $3.00 per 1M output tokens for input above 256K and up to 1M. Regional and global pricing varies.
Qwen Multimodal

Fast multimodal analysis, long audio understanding, audiovisual question answering, voice assistants and spoken-response applications

Context 262K
Reasoning 7/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current; canonical model ID functionally equivalent to qwen3.5-omni-flash-2026-03-15

Input Singapore international: $0.40 per 1M tokens for text/image/video input and $3.00 per 1M tokens for audio input; China Beijing: $0.30 per 1M tokens for text/image/video input and $2.48 per 1M tokens for audio input
Output Singapore international: $2.20 per 1M tokens for text output and $11.90 per 1M tokens for text-and-audio output; China Beijing: $1.83 per 1M tokens for text output and $9.90 per 1M tokens for text-and-audio output
Qwen Multimodal

Low-latency voice assistants, speech-to-speech applications, realtime multimedia analysis, interactive agents, and multimodal conversations.

Context 262K
Reasoning 6/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current and available; rolling model identity functionally equivalent to qwen3.5-omni-flash-realtime-2026-03-15

Input International: $0.55 per 1M tokens for text/image/video input; $4.50 per 1M tokens for audio input. China mainland: $0.45 per 1M tokens for text/image/video input; $3.71 per 1M tokens for audio input.
Output International: $3.30 per 1M tokens for text output; $17.70 per 1M tokens for audio output. China mainland: $2.75 per 1M tokens for text output; $14.71 per 1M tokens for audio output.
Qwen Multimodal
Qwen
Qwen3.5-Omni

Qwen3.5-Omni-Plus

Multilingual voice assistants, speech-enabled multimodal applications, audio-visual analysis, spoken explanations, accessibility tools, and interactive media workflows.

Context 262K
Reasoning 7/10
Speed 7/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

current

Input $1.40 per 1M tokens for text/image/video input; $11.00 per 1M tokens for audio input in international deployments. China (Beijing) snapshot pricing is $0.96 per 1M tokens for text/image/video input and $7.29 per 1M tokens for audio input.
Output $8.30 per 1M tokens for text output; $44.00 per 1M tokens for text-and-audio output in international deployments. China (Beijing) snapshot pricing is $5.50 per 1M tokens for text output and $29.29 per 1M tokens for text-and-audio output.
Qwen Multimodal

Real-time voice assistants, speech-to-speech applications, multimodal customer service, visual conversational agents, live multimedia analysis, and interactive applications requiring controllable speech output.

Context 262K
Reasoning 6/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current and available; canonical model identifier functionally equivalent to snapshot qwen3.5-omni-plus-realtime-2026-03-15

Input China (Beijing): $1.38 per 1M text/image/video tokens; $11 per 1M audio tokens. Singapore: $2.10 per 1M text/image/video tokens; $16.50 per 1M audio tokens.
Output China (Beijing): $8.25 per 1M text-output tokens; $41.26 per 1M text-and-audio-output tokens. Singapore: $12.40 per 1M text-output tokens; $62 per 1M text-and-audio-output tokens.
Qwen Multimodal

Coding agents, repository-level software engineering, visual document analysis, video understanding, STEM reasoning, long-context assistants, and self-hosted multimodal applications

Context 262K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current and available; open-weight release and Alibaba Cloud Model Studio API model

Input $0.60 per 1M tokens internationally; $0.412564 per 1M tokens in China Beijing and Singapore
Output $3.60 per 1M tokens internationally; $2.475384 per 1M tokens in China Beijing and Singapore
Qwen Multimodal

Agentic coding, long-context software engineering, multimodal analysis, tool-using agents, and self-hosted deployments

Context 262K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current; open-weight and available through Alibaba Cloud Model Studio

Input USD 0.248 per 1M tokens in US Virginia, Germany Frankfurt, and China Beijing; USD 0.375 per 1M tokens in Singapore
Output USD 1.485 per 1M tokens in US Virginia, Germany Frankfurt, and China Beijing; USD 2.25 per 1M tokens in Singapore
Qwen Multimodal

Fast multimodal assistants, coding agents, visual document analysis, video understanding, tool-using workflows, object localization, and large-context applications

Context 1M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current

Input $0.25 per 1M input tokens up to 256K input tokens; $1.00 per 1M input tokens above 256K and up to 1M, Singapore international pricing
Output $1.50 per 1M output tokens up to 256K input tokens; $4.00 per 1M output tokens above 256K and up to 1M, Singapore international pricing
Qwen General Purpose

Advanced coding agents, front-end development, long-context analysis, structured API workflows, and text generation with web search

Context 262K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Preview; currently accessible; scheduled for deprecation on October 10, 2026

Input $1.30 per 1M tokens for input up to 128K; $2.00 per 1M tokens for input above 128K and up to 256K, international pricing
Output $7.80 per 1M tokens for requests up to 128K input; $12.00 per 1M tokens for requests above 128K and up to 256K input, international pricing
Qwen Multimodal

Long-context multimodal analysis, agentic coding, OCR, object localization, frontend development, visual reasoning, and tool-enabled enterprise assistants

Context 1M
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current; rolling model ID qwen3.6-plus is currently equivalent to qwen3.6-plus-2026-04-02

Input $0.276 per 1M tokens for 0-256K input tokens; $1.101 per 1M tokens for 256K-1M input tokens in Global deployment
Output $1.651 per 1M tokens for 0-256K input tokens; $6.602 per 1M tokens for 256K-1M input tokens in Global deployment
Qwen Multimodal

Fast multimodal agents, visual coding, tool-use workflows, search agents, and long-context document or screen analysis

Context 1M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current and available; rolling qwen3.7-flash identifier is currently equivalent to qwen3.7-flash-2026-07-15

Input Global: $0.028 per 1M input tokens for 0-32K; $0.083 for 32K-256K; $0.165 for 256K-1M. Singapore international: $0.030, $0.100, and $0.200 per 1M input tokens respectively.
Output Global: $0.110 per 1M output tokens for 0-32K; $0.330 for 32K-256K; $0.660 for 256K-1M. Singapore international: $0.130, $0.400, and $0.800 per 1M output tokens respectively.
Qwen Multimodal

Long-context reasoning, multimodal document and video analysis, coding, tool-using agents, structured extraction, and enterprise productivity workflows

Context 1M
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Current; rolling identifier currently equivalent to qwen3.7-plus-2026-05-26

Input $0.40 per 1M tokens up to 256K input tokens and $1.20 per 1M tokens from 256K to 1M input tokens in US Virginia; regional and promotional pricing varies
Output $1.60 per 1M tokens up to 256K input tokens and $4.80 per 1M tokens from 256K to 1M input tokens in US Virginia; thinking-mode output uses the same listed rates
Qwen Embedding

Multilingual semantic search, retrieval-augmented generation, code retrieval, recommendation, clustering, classification, and large-scale text vectorization

Context 131K
Speed 8/10
Outputs
Embeddings
Status

Current and available through Alibaba Cloud Model Studio internationally

Input $0.07 per 1 million input tokens in Singapore international deployment; China pricing may differ by region and billing mode
Qwen Reasoning

Advanced reasoning, coding, scientific and professional research, long-context analysis, and long-horizon agent workflows

Context 1M
Reasoning 10/10
Speed 4/10
Outputs
Text
Capabilities
Tool use Web search Streaming
Status

Current and available; open-weight release and hosted API model

Input $2 per 1 million tokens internationally; $1.65 per 1 million tokens in China (Beijing)
Output $6 per 1 million tokens internationally; $4.951 per 1 million tokens in China (Beijing)
Qwen Multimodal

Coding assistants, repository analysis, visual document workflows, long-context research, office automation, multimodal agents, and tool-using applications

Context 1M
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Active and currently available; open-weight release and Alibaba Cloud Model Studio API model

Input $0.424 per 1M tokens in China Beijing; $0.50 per 1M tokens in Singapore. Implicit cache input is $0.085 per 1M tokens in Beijing and $0.10 per 1M tokens in Singapore. Explicit cache creation is $0.53 per 1M tokens in Beijing and $0.625 per 1M tokens in Si
Output $1.696 per 1M tokens in China Beijing; $3 per 1M tokens in Singapore.
Qwen Multimodal

Fast long-context reasoning, coding assistance, visual document and chart analysis, video understanding, function-calling agents, and high-concurrency applications

Context 1M
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Current and accessible through Alibaba Cloud Model Studio

Input $0.113 per 1 million input tokens; implicit cache input $0.014 per 1 million tokens; explicit cache creation $0.177 per 1 million tokens; explicit cache read $0.014 per 1 million tokens
Output $0.382 per 1 million output tokens
Qwen Multimodal

High-volume coding, agentic workflows, long-context document and codebase analysis, office automation, and multimodal text-image-video understanding

Context 262K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Fine-tuning
Status

Current open-weight experimental preview

Qwen Multimodal

Complex coding, autonomous software engineering, long-horizon agent workflows, professional document analysis, visual reasoning, long videos, and demanding research tasks.

Context 1M
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Web search
Status

Stable official release; currently available through Alibaba Cloud Model Studio

Input US$1.65 per 1M input tokens and US$2.00 per 1M input tokens for International scope; regional pricing varies. Cached-input pricing starts at US$0.206 per 1M tokens for the listed regional/global scope and US$0.25 per 1M tokens for International scope.
Output US$4.951 per 1M output tokens for the listed regional/global scope and US$6.00 per 1M output tokens for International scope.
Qwen Multimodal

Real-time speech translation, multilingual meetings, live interpretation, translated voice communication, and audiovisual translation with low latency.

Context 53K
Reasoning 4/10
Speed 9/10
Outputs
Text Speech
Capabilities
Image input Audio input Multimodal input Streaming
Status

Current stable model

Input Singapore: audio input $7.50 per 1 million tokens; image input $0.55 per 1 million tokens. China (Beijing): audio input $5.653 per 1 million tokens; image input $0.466 per 1 million tokens.
Output Singapore: text output $20 per 1 million tokens; audio output $30 per 1 million tokens. China (Beijing): text output $14.133 per 1 million tokens; audio output $22.613 per 1 million tokens.
Qwen Multimodal

Long-form audio and video understanding, multimedia analysis, audio-visual agents, content summarization, and tool-using workflows

Context 1M
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current and available

Input USD 0.15 per 1 million input tokens for International deployment; USD 0.016 per 1 million cache-hit input tokens
Output USD 0.47 per 1 million output tokens for International deployment
Qwen Multimodal

Real-time voice assistants, speech-to-speech applications, interactive video agents, live media analysis, multimodal customer service, meeting and collaboration interfaces, and applications requiring tool or MCP integration.

Context 197K
Reasoning 7/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Video input Multimodal input Tool use Streaming
Status

Current and available; international deployment in Singapore

Input China (Beijing): CNY 1.5 per 1 million tokens for text/images/video input and CNY 6 per 1 million tokens for audio input. Singapore: CNY 1.677 per 1 million tokens for text/images/video input and CNY 6.781 per 1 million tokens for audio input.
Output China (Beijing): CNY 4.5 per 1 million tokens for text output and CNY 12 per 1 million tokens for audio output. Singapore: CNY 5.104 per 1 million tokens for text output and CNY 13.636 per 1 million tokens for audio output.
Qwen Lightweight
Qwen
Qwen Character

qwen-flash-character

Low-latency character dialogue, virtual companions, game NPCs, role-playing applications, IP character replication, and conversational smart devices

Context 33K
Reasoning 4/10
Speed 9/10
Outputs
Text
Capabilities
Web search Streaming
Status

Current; dynamically updated managed model

Input Regional pricing: USD 0.034 per 1M input tokens in Beijing and US Virginia; USD 0.05 per 1M input tokens in Singapore. Cached input is listed at USD 0.007 per 1M tokens in Beijing and US Virginia and USD 0.01 in Singapore. Alibaba Cloud also publishes loc
Output Regional pricing: USD 0.203 per 1M output tokens in Beijing and US Virginia; USD 0.40 per 1M output tokens in Singapore. Alibaba Cloud also publishes localized regional prices.
Qwen Image Generation

Complex text-to-image layouts, multilingual typography, posters, menus, storyboards, interface mockups, product visuals, and image editing with one to three reference images

Context 5K
Reasoning 2/10
Speed 6/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and available

Input $0.003 per image in the international Singapore deployment
Output $0.04 per 1K image or $0.075 per 2K image in the international Singapore deployment; regional pricing varies
Qwen Image Editing

Instruction-based image editing, multi-image fusion, character-consistent compositions, object replacement, style transfer, poster and text editing, and product-image variation.

Reasoning 5/10
Speed 8/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current; canonical model currently equivalent to qwen-image-edit-plus-2025-10-30

Input 0; image editing is billed by output image
Output $0.028671 per image in China (Beijing); $0.03 per image in Singapore/international
Qwen Image Generation
Qwen
Qwen Image

Qwen-Image-Max

Realistic text-to-image generation, creative concepts, marketing visuals, editorial artwork, product concepts, and general-purpose image creation.

Reasoning 1/10
Speed 5/10
Outputs
Image
Status

Current and available for text-to-image generation; the undated qwen-image-max identifier is currently equivalent to qwen-image-max-2025-12-30.

Input No separate input charge; text prompt input is included in per-image billing.
Output $0.075/image in Singapore international deployment; $0.071677/image in China (Beijing).
Qwen Other
Qwen
Qwen Image 3.0

qwen-image-3.0

Fast text-to-image generation, text-heavy layouts, image editing, reference-image compositing, marketing graphics, and high-volume image workflows.

Reasoning 3/10
Speed 8/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and available through Alibaba Cloud Model Studio

Input $0.003 per input image internationally; regional pricing varies
Output $0.03 per generated image internationally; regional pricing varies
Reka AI Multimodal
Reka AI
Reka Core

Reka Core

Complex multimodal analysis, long-context document understanding, image and video question answering, audio-aware workflows, coding, and enterprise applications requiring broad input support.

Context 128K
Reasoning 8/10
Speed 5/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Streaming
Status

Available/listed for API pricing, but not identified as a baseline model that is always publicly available; account-dependent availability should be verified.

Input $2.00 per 1M input tokens; $0.02 per image; $0.08 per video minute; $0.02 per audio minute
Output $6.00 per 1M output tokens
Reka AI Multimodal
Reka AI
Reka Edge

Reka Edge

Low-latency image and video understanding, object detection, physical AI, robotics, edge deployment, and on-device visual assistants

Context 16K
Reasoning 7/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Streaming
Status

Current and publicly available

Input $0.10 per 1 million tokens; $0.005 per image; $0.03 per video minute
Output $0.10 per 1 million tokens
Reka AI Multimodal
Reka AI
Reka Flash

Reka Flash

Fast multimodal applications, document and image analysis, short-video understanding, structured extraction, multilingual chat, coding, and tool-using agents

Context 128K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current and publicly available baseline model

Input $0.80 per 1 million input tokens; $0.01 per image; $0.06 per minute of video; $0.015 per minute of audio
Output $2.00 per 1 million output tokens
Reka AI Reasoning
Reka AI
Reka Flash

Reka Flash 3

Low-latency reasoning, coding assistance, function calling, local deployment, on-device applications and cost-sensitive inference

Context 66K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current and accessible; open-weight research-preview model with API availability through Reka Gateway

Input $0.10 per 1 million tokens through Reka Gateway
Output $0.20 per 1 million tokens through Reka Gateway
Reka AI Reasoning
Reka AI
Reka Flash

Reka Flash 3.1

Coding, mathematical reasoning, local inference, research experimentation, and fine-tuning for agentic workflows

Context 33K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available as an open-weight model; official API availability was announced but the current public baseline API catalog does not list the exact model identifier

SenseTime Video Generation

Audio-driven digital humans, lip-sync video, singing avatars, multilingual character animation, multi-person dialogue, and long-duration talking-video generation

Reasoning 1/10
Speed 9/10
Outputs
Video
Capabilities
Image input Audio input Multimodal input
Status

Current; ongoing project

SenseTime Multimodal
SenseTime
SenseNova-MARS

SenseNova-MARS-8B

Visual question answering, high-resolution image understanding, multimodal search, image-grounded research, and tool-assisted agentic reasoning

Context 262K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search
Status

Current open-weight research model

SenseTime Multimodal
SenseTime
SenseNova-MARS

SenseNova-MARS-32B

Visual deep-search, fine-grained image understanding, multimodal agent research, and self-hosted tool-using vision-language applications

Context 262K
Reasoning 8/10
Speed 4/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use
Status

Current open-weight model

SenseTime Multimodal

Native image generation, high-resolution visual creation, image editing, infographic and layout generation, visual understanding, and multimodal research or creative workflows

Reasoning 6/10
Speed 4/10
Outputs
Text Image
Capabilities
Image input Multimodal input
Status

Current open-weight flagship checkpoint

Input No official hosted API price verified; downloadable model weights are available
Output No official hosted API price verified; downloadable model weights are available
SenseTime Multimodal
SenseTime
SenseNova-Vision

SenseNova-Vision-7B-MoT

Unified computer-vision research, detection, OCR, segmentation, depth and normal estimation, visual grounding, and multi-view geometry

Reasoning 7/10
Speed 4/10
Outputs
Text Image
Capabilities
Image input Multimodal input Fine-tuning
Status

Current open-weight research model

SenseTime Multimodal

Long-horizon multimodal agents, data analysis, deep research, complex information presentation, office automation, and tool-driven workflows

Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Preview; currently available through the SenseNova Token Plan

Input Free during the public beta Token Plan; standard per-token API pricing is not publicly verified
Output Free during the public beta Token Plan; standard per-token API pricing is not publicly verified
SenseTime Image Generation
SenseTime
SenseNova Seko

SekoIDX

Character-consistent image generation for multi-episode videos, motion comics, short dramas, storyboards, and cross-shot visual production

Reasoning 2/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current; integrated into the SenseNova Seko series and Seko 2.0 platform

SenseTime Multimodal
SenseTime
SenseNova U

SenseNova U1 Pro

Professional image creation, infographics, advertising, e-commerce assets, presentations, educational diagrams, storyboards, and multi-step visual delivery workflows

Reasoning 8/10
Speed 5/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current production release; enterprise API available through whitelist access

SenseTime Multimodal
SenseTime
SenseNova U1

SenseNova U1

Open-source visual understanding, image generation, image editing, infographic creation, visual reasoning, and continuous image-text workflows.

Reasoning 7/10
Speed 7/10
Outputs
Text Image
Capabilities
Image input Multimodal input Fine-tuning
Status

Open-source model series; U1-8B-MoT and U1-A3B-MoT variants are available. SenseNova U1.5 is the newer successor for visual creation and editing, while U1 checkpoints remain documented and downloadable.

SenseTime Lightweight
SenseTime
SenseNova U1

SenseNova U1 Fast

Fast infographic generation, dense visual explanations, charts, diagrams, presentation graphics, and information-heavy layouts

Speed 8/10
Outputs
Image
Status

Current and accessible through SenseNova Token Plan and dedicated image-generation integrations

Input Free public beta with usage quota; commercial pricing not verified
Output Free public beta with usage quota; commercial pricing not verified
SenseTime Multimodal
SenseTime
SenseNova U1.5

SenseNova U1.5 Lite

Text-to-image generation, reference-image creation, visual design, posters, infographics, product imagery, and iterative image editing

Reasoning 5/10
Speed 8/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and available through the SenseNova Token Plan

StepFun Reasoning

Coding agents, software engineering, long-context reasoning, tool-using agents, private local inference, and work-centric automation

Context 262K
Reasoning 9/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current; open-weight model available for local deployment and accessible through StepFun's API platform

StepFun Multimodal

High-throughput coding agents, visual document and UI understanding, search-heavy research workflows, long-context analysis, and tool-using autonomous agents

Context 256K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search
Status

Current and available

Input $0.20 per 1 million tokens for cache misses; $0.04 per 1 million tokens for cache hits
Output $1.15 per 1 million tokens
StepFun Reasoning

Long-running agentic workflows, software engineering, coding, professional knowledge work, financial analysis, research, and vision-assisted tasks

Context 1M
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use
Status

Current preview; available through StepFun products and API. Open-weight release planned for October 15, 2026.

Input ¥7 per 1 million uncached input tokens; ¥0.35 per 1 million cached input tokens
Output ¥20 per 1 million output tokens
StepFun Speech Recognition

Multilingual transcription, meetings, subtitles, live content, customer-service recordings, specialized terminology, noisy audio, dialects, code-switching, and singing or music transcription

Reasoning 5/10
Speed 9/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current and available through the StepFun API

Input 2.8 CNY per audio hour
StepFun Other
StepFun
StepAudio 3

StepAudio 3 Gen

Zero-shot text-to-speech, natural-language voice design, singing and vocal generation, music, sound effects, ambience, and complete multi-element audio scenes

Outputs
Speech Music
Capabilities
Audio input Multimodal input
Status

Current research and platform-listed model; public API and access details are limited

StepFun Other

Song generation, instrumental music, lyric-to-song workflows, accompaniment, cover-style synthesis, and rapid music prototyping

Reasoning 2/10
Speed 3/10
Outputs
Music
Capabilities
Audio input Multimodal input
Status

Current; publicly showcased and available through StepFun's StepAudio 3 music experience

StepFun Realtime Audio

Natural realtime voice conversation, full-duplex interaction, interruption-aware assistants, emotional audio understanding and voice agents that use tools.

Reasoning 7/10
Speed 9/10
Outputs
Speech
Capabilities
Audio input Multimodal input Tool use Streaming
Status

Current and publicly listed

StepFun Multimodal
StepFun
Step Edge

Step Edge

On-device visual assistants, GUI grounding, screen understanding, spatial reasoning, automotive interaction, and low-latency local agent workflows

Context 131K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use
Status

Current; officially announced as an edge-deployment base model

StepFun Multimodal

On-device speech recognition, audio understanding, voice assistants, in-vehicle interaction, privacy-sensitive audio processing, and low-latency edge AI.

Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Audio input
Status

Current; edge-deployment model with public capability information but no verified public hosted API pricing or general model-download documentation located.

StepFun Other
StepFun
Step Edge

Step Edge Gen

On-device text-to-image generation, local image editing, privacy-sensitive creative features, and low-latency edge-device applications.

Reasoning 1/10
Speed 9/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current research model; public technical information is available, but official public API availability and downloadable weights were not verified.

StepFun Other
StepFun
Step Edge

Step Edge GUI

Low-latency desktop and mobile GUI automation, visual grounding, local computer-use agents, and privacy-sensitive edge workflows

Reasoning 6/10
Speed 9/10
Outputs
Actions
Capabilities
Image input Multimodal input Tool use
Status

Current; edge-deployment GUI agent model

Self-hosted text generation, research, quantization, and domain-specific fine-tuning

Context 2K
Reasoning 4/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Available open-weight model; older Falcon generation with newer successors

Input No official hosted API price; self-hosted model weights
Output No official hosted API price; self-hosted model weights

Research, self-hosted text generation, model fine-tuning, summarization, and general language-model experimentation

Context 2K
Reasoning 4/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as open weights; legacy-generation model

Input No official hosted API pricing; self-hosted open weights
Output No official hosted API pricing; self-hosted open weights

Research, self-hosted text generation, language-model evaluation, and domain-specific fine-tuning when substantial GPU infrastructure is available

Context 2K
Reasoning 6/10
Speed 3/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight model; legacy-generation checkpoint

Input No official first-party hosted API pricing; self-hosted/open-weight distribution
Output No official first-party hosted API pricing; self-hosted/open-weight distribution

Memory-efficient local text generation, edge-device experimentation, research, and downstream fine-tuning

Context 33K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Current open-weight model; downloadable from Hugging Face

Input No official hosted API pricing; self-hosted model weights
Output No official hosted API pricing; self-hosted model weights

Efficient local text generation, multilingual language modeling, long-context applications, model adaptation, and compact reasoning research.

Context 131K
Reasoning 6/10
Speed 8/10
Outputs
Text
Status

Current open-weight model

Input No official hosted API pricing; downloadable weights are available for self-hosted deployment.
Output No official hosted API pricing; downloadable weights are available for self-hosted deployment.

Efficient local inference, compact conversational assistants, multilingual text generation, long-context applications, and resource-constrained deployments

Context 131K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Streaming
Status

Current; open-weight

Input No official hosted API pricing; self-hosted/open-weight model
Output No official hosted API pricing; self-hosted/open-weight model

Local text generation, multilingual research, domain adaptation, fine-tuning, long-context experiments, and resource-conscious deployment.

Context 131K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current open-weight model; downloadable from Hugging Face and usable with supported local inference frameworks.

Input No official hosted API pricing; self-hosted model weights
Output No official hosted API pricing; self-hosted model weights

Private or local multilingual text generation, long-context experimentation, domain adaptation, and fine-tuning from a 7B-scale open-weight foundation

Context 262K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning
Status

Current open-weight model; pretrained base checkpoint

Input No official hosted API pricing; downloadable open-weight checkpoint
Output No official hosted API pricing; deployment cost depends on self-hosted infrastructure

Long-context text generation, multilingual instruction following, document processing, retrieval-augmented generation, coding and self-hosted enterprise deployments

Context 262K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current open-weight model; publicly downloadable and usable with Transformers, vLLM and llama.cpp

Input No official hosted API price; self-hosted model weights
Output No official hosted API price; self-hosted model weights

Ultra-lightweight local text generation, edge deployment, offline instruction following, rewriting, extraction, and embedded AI experiments

Context 262K
Reasoning 2/10
Speed 10/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Current open-weight model

Input No official hosted API price; downloadable weights
Output No official hosted API price; downloadable weights

Lightweight local text generation, instruction-following, edge devices, embedded applications, experimentation, and privacy-sensitive deployments.

Context 262K
Reasoning 2/10
Speed 9/10
Outputs
Text
Status

Current open-weight model

Input No official hosted API pricing found; model weights are available for local deployment under the Falcon-LLM License.
Output No official hosted API pricing found; local inference costs depend on hardware and deployment framework.
Technology Innovation Institute (TII)
Falcon 2

Falcon2-11B

Research, multilingual text generation, fine-tuning, quantization, and self-hosted inference

Context 8K
Reasoning 4/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as an open-weight pretrained base model

Input No official hosted API price; downloadable weights for self-hosted deployment
Output No official hosted API price; infrastructure-dependent

Local inference, multilingual text completion, research, continued pretraining, domain adaptation, and fine-tuning on constrained hardware

Context 4K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight pretrained base model

Input No official hosted API price; self-hosted/open-weight model
Output No official hosted API price; self-hosted/open-weight model

Lightweight local assistants, multilingual instruction following, extraction, classification, education, and resource-conscious deployments

Context 8K
Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Streaming
Status

Current open-weight model; downloadable from Hugging Face

Input No official first-party hosted API pricing
Output No official first-party hosted API pricing

Fine-tuning, multilingual text generation, research, compact local inference, and edge-oriented deployments

Context 8K
Reasoning 5/10
Speed 8/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight model

Input No official hosted API price; downloadable weights for self-hosted deployment
Output No official hosted API price; downloadable weights for self-hosted deployment

Self-hosted multilingual chat, instruction following, reasoning, mathematics, coding, research, and long-context text generation

Context 33K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available open-weight model

Input No official hosted API price; downloadable weights
Output No official hosted API price; downloadable weights

Fine-tuning, multilingual text generation, language-model research, code and mathematics experimentation, and self-hosted inference

Context 33K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available open-weight model

Input No official hosted API pricing; self-hosted/open-weight deployment
Output No official hosted API pricing; self-hosted/open-weight deployment

Self-hosted multilingual assistants, STEM and mathematics tasks, coding, instruction following, research, and local function-calling systems.

Context 33K
Reasoning 7/10
Speed 5/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Available open-weight model; official repository remains accessible. No official deprecation or shutdown date found.

Input No official hosted API pricing; downloadable weights are provided for self-managed inference.
Output No official hosted API pricing; downloadable weights are provided for self-managed inference.
Tencent AI
AuK

AuK

Open-source text-to-speech, reference-voice generation, speech and lyric editing, emotion and timbre transformation, speech enhancement, and source separation

Reasoning 2/10
Speed 5/10
Outputs
Speech
Capabilities
Audio input Multimodal input Fine-tuning
Status

Current open-source release

Tencent AI Lightweight

Efficient local text generation, long-context analysis, lightweight reasoning, coding assistance, and agent-oriented applications

Context 256K
Reasoning 6/10
Speed 9/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight model; publicly available for download and local deployment

Tencent AI General Purpose

Local and self-hosted text generation, Chinese and multilingual instruction following, long-context analysis, mathematics, coding, reasoning, and lightweight agent workloads

Context 262K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current open-weight model; publicly available for self-hosted deployment

Input No official hosted API price published; self-hosted model weights are available
Output No official hosted API price published; self-hosted inference costs depend on infrastructure
Tencent AI General Purpose

Chinese and multilingual text generation, reasoning, coding, long-context workloads, private local inference, quantized deployment, and domain fine-tuning.

Context 262K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Fine-tuning
Status

Current open-weight downloadable model

Input No first-party hosted API token price identified; downloadable weights are intended for self-hosted or third-party deployment.
Output No first-party hosted API token price identified; downloadable weights are intended for self-hosted or third-party deployment.
Tencent AI Reasoning
Tencent AI
Hunyuan-A13B

Hunyuan-A13B

Open-weight reasoning, long-context analysis, mathematics, science, coding, agent workflows, and cost-conscious self-hosted inference

Context 262K
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current and accessible as an open-weight model; also listed as the Tencent Cloud API model hunyuan-a13b. Tencent Cloud documentation notes an ongoing migration of Hunyuan services toward TokenHub.

Input ¥0.50 per 1 million input tokens on Tencent Cloud postpaid API
Output ¥2 per 1 million output tokens on Tencent Cloud postpaid API
Tencent AI General Purpose
Tencent AI
Hunyuan-Large

Hunyuan-Large

Large-scale Chinese and English text generation, reasoning, mathematics, coding, long-context analysis, research, and self-hosted experimentation

Context 256K
Reasoning 8/10
Speed 4/10
Outputs
Text
Capabilities
Fine-tuning Streaming
Status

Available as an open-weight release; legacy for the historical Tencent Cloud hunyuan-large API identifier

Tencent AI
Hunyuan3D

Hunyuan3D-2.1

Image-to-3D asset creation, game and virtual-world content, product visualization, design prototyping, and self-hosted 3D generation

Speed 4/10
Capabilities
Image input Multimodal input Fine-tuning
Status

Current open-weight release; self-hosted deployment

Input No official hosted token or per-request API price; model weights are available for self-hosting
Output No official hosted output price; generated 3D assets are produced through self-hosted inference
Tencent AI
Hunyuan3D

Hunyuan3D 2.0

Local image-to-3D asset generation, textured mesh creation, game and design prototypes, Blender workflows, and research on open 3D generative models

Speed 6/10
Capabilities
Image input Multimodal input Fine-tuning
Status

Available open-weight model system; superseded by the newer Hunyuan3D-2.1 release but still publicly accessible

Tencent AI
Hunyuan 3D

HY-3D-3.0

Text-to-3D, image-to-3D, sketch-to-3D, rapid game and e-commerce asset creation, 3D printing, and production-oriented asset prototyping.

Speed 8/10
Capabilities
Image input Multimodal input
Status

Current and accessible through Tencent Cloud APIs; newer HY-3D-3.1 is also available as a separate model version.

Input Credit-based pricing: 25 credits for a default Professional normal textured generation; 15 credits for Geometry mode; 30 credits for LowPoly mode. Additional features such as PBR, multi-view, and custom face count consume extra credits.
Output Not token-priced; generated 3D assets are billed by generation credits.

Local English- and Chinese-language text-to-image generation, creative prototyping, ComfyUI workflows, LoRA customization, and ControlNet-based image conditioning

Reasoning 1/10
Speed 6/10
Outputs
Image
Capabilities
Fine-tuning
Status

Open-weight and downloadable; no official deprecation or shutdown notice found

Input No official hosted API input pricing published; intended primarily for local deployment
Output No official hosted API output pricing published; image generation uses local compute rather than a documented token price
Tencent AI
HunyuanImage

Hy-Image-3.0

Text-to-image generation, reference-guided image creation, marketing graphics, e-commerce content, visual ideation, and image-generation applications requiring custom aspect ratios

Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and accessible through Tencent Cloud TokenHub; also available as an open-weight HunyuanImage-3.0 release

Output Approximately 0.20 CNY per generated image in Tencent Cloud mainland China TokenHub pricing; approximately 0.032 USD per image in cited international pricing
Tencent AI Multimodal
Tencent AI
HunyuanOCR

HunyuanOCR-1.5

Multilingual OCR, document parsing, text spotting, table and formula extraction, structured information extraction, and local visual-document processing

Context 131K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning Streaming
Status

Current open-weight release

Tencent AI Multimodal
Tencent AI
Hunyuan turbos-vision

HY-Vision-Video

Video description, video question answering, video summarization, content review, scene analysis, and video metadata generation.

Context 32K
Reasoning 3/10
Speed 8/10
Outputs
Text
Capabilities
Video input Multimodal input
Status

Online and currently listed in Tencent Cloud TokenHub; the older Hunyuan platform entry was retired on 2026-06-22, while the TokenHub model remains available.

Input CNY 3 per 1 million input tokens
Output CNY 9 per 1 million output tokens
Tencent AI
HunyuanVideo

HunyuanVideo

Research and production experimentation with high-quality local text-to-video generation

Reasoning 1/10
Speed 3/10
Outputs
Video
Status

Open-source and currently accessible; original HunyuanVideo model, with HunyuanVideo-1.5 released later as a lighter successor

Tencent AI
HunyuanVideo

HunyuanVideo-1.5

Local text-to-video and image-to-video generation, open-source video research, creative prototyping, and developers needing a comparatively lightweight high-quality video model.

Reasoning 1/10
Speed 6/10
Outputs
Video
Capabilities
Image input Multimodal input Fine-tuning
Status

Current and publicly accessible open-weight model

Input No official hosted API price identified; model weights are available for local deployment under the Tencent Hunyuan Community License.
Output No official hosted API price identified; local inference costs depend on hardware and infrastructure.
Tencent AI
HunyuanVideo

HunyuanVideo-I2V

Image-to-video generation, reference-image animation, visual effects, creative video prototyping, and locally hosted open-weight video workflows

Speed 4/10
Outputs
Video
Capabilities
Image input Multimodal input Fine-tuning
Status

Available as an open-weight model and official open-source repository; no official shutdown date found

Tencent AI Multimodal
Tencent AI
Hunyuan Vision 1.5

HY-Vision-1.5-Thinking

Image-grounded reasoning, OCR, chart and document analysis, visual localization, educational problem solving, and multilingual visual question answering

Context 40K
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Online and currently available through Tencent Cloud TokenHub

Input CNY 3 per 1M input tokens
Output CNY 9 per 1M output tokens
Tencent AI Reasoning

Coding agents, long-context analysis, complex reasoning, productivity automation, structured workflows, and multi-step tool use

Context 256K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Web search Fine-tuning
Status

Current; open-weight and available through Tencent Cloud TokenHub

Input CNY 1 per 1 million input tokens; cached input CNY 0.25 per 1 million tokens on Tencent Cloud TokenHub
Output CNY 4 per 1 million output tokens on Tencent Cloud TokenHub

Fast text-to-3D and image-to-3D asset generation, prototyping, and automated 3D content pipelines

Reasoning 1/10
Speed 8/10
Capabilities
Image input Multimodal input
Status

Current and available through Tencent Cloud TokenHub

Input 15–25 Tencent Cloud points per generation request; 1 point is listed as CNY 0.12

Automated rigging and skinning of human or animal 3D characters for animation, games, virtual characters, and asset prototyping

Reasoning 1/10
Speed 5/10
Status

Current and accessible through Tencent Cloud TokenHub

Input 10 Tencent Cloud credits per request
Output Included in the 10-credit per-request charge; output is a rigged 3D model file

Reference-guided texturing of existing OBJ or GLB meshes, including PBR material generation and automated asset preparation

Reasoning 1/10
Speed 6/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and available through Tencent Cloud TokenHub API

Output 30 Tencent Cloud points per generation

Automated UV unwrapping and preparation of 3D assets for texturing, rendering, game development, and digital-content workflows.

Reasoning 1/10
Speed 7/10
Status

Current and accessible through Tencent Cloud TokenHub

Output 10 points per call; Tencent Cloud states that 1 point equals CNY 0.12

Real-time and asynchronous Mandarin, English, mixed-language, and Chinese-dialect transcription, captions, subtitles, and short voice-command recognition

Reasoning 4/10
Speed 8/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Preview; internal-test availability

Input Usage-based Tencent Cloud ASR pricing; exact model price not verified in the reviewed official documentation
Output No separate output price; transcription is billed through the applicable Tencent Cloud ASR or large-model 2.0 pricing scheme
Tencent AI Multimodal

High-resolution text-to-image generation, reference-image creation, multi-turn image editing, posters, UI concepts, marketing assets, and product-image workflows

Context 100K
Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input Web search
Status

Current preview model

Input 10 CNY per million tokens; reference-image generation uses the published TokenHub token rules
Output 15,000 tokens per 1K or 2K image; 20,000 tokens per 4K image, equivalent to approximately CNY 0.15 and CNY 0.20 respectively at the listed CNY 10 per million-token rate
Tencent AI Lightweight

Low-latency multilingual translation, localized content, structured translation instructions, and cost-sensitive production workflows

Context 8K
Reasoning 2/10
Speed 9/10
Outputs
Text
Capabilities
Streaming
Status

Current and available through Tencent Cloud TokenHub

Input ¥0.3 per 1 million tokens
Output ¥1.2 per 1 million tokens
Tencent AI Translation

Professional multilingual translation, localization, terminology-sensitive workflows, and context-aware business translation

Context 8K
Reasoning 2/10
Speed 8/10
Outputs
Text
Capabilities
Streaming
Status

Current and available through Tencent Cloud TokenHub

Input ¥0.5 per 1 million tokens in China; US$0.074 per 1 million tokens on Tencent Cloud international pricing
Output ¥2 per 1 million tokens in China; US$0.295 per 1 million tokens on Tencent Cloud international pricing
Tencent AI Specialized

Professional, domain-specific and high-quality multilingual translation with contextual disambiguation and instruction following

Context 8K
Reasoning 6/10
Speed 8/10
Outputs
Text
Capabilities
Streaming
Status

Online and currently available through Tencent Cloud TokenHub

Input 0.5 CNY per 1 million input tokens
Output 2 CNY per 1 million output tokens
Tencent AI
Hy-Role

Hy-Role

Chinese role-play, character simulation, fictional dialogue, AI avatars, and emotionally oriented conversational experiences

Context 32K
Reasoning 3/10
Speed 7/10
Outputs
Text
Capabilities
Streaming
Status

Current and available through Tencent Cloud TokenHub

Input CNY 2.4 per 1 million input tokens
Output CNY 9.6 per 1 million output tokens
Tencent AI Multimodal

Image understanding, OCR, chart and diagram analysis, STEM visual reasoning, visual question answering, and multi-image comparison

Context 44K
Reasoning 7/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input
Status

Current and available through Tencent Cloud TokenHub

Input ¥7.5 per 1 million input tokens
Output ¥17.5 per 1 million output tokens

Generating explorable 3D environments, Gaussian-splat scenes, point clouds, and collision meshes from text or reference images

Reasoning 1/10
Speed 2/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and accessible through Tencent Cloud TokenHub

Input ¥10 per 1 million tokens; reference usage is 8,000,000 tokens per generated scene
Output Approximately ¥80 per generated 3D scene at the documented reference usage
Tencent AI Multimodal

Text-to-3D, image-to-3D, multi-view reconstruction, game assets, product visualization, digital-human content, e-commerce assets and 3D printing workflows

Reasoning 1/10
Speed 7/10
Capabilities
Image input Multimodal input
Status

Current and available through Tencent Cloud HY-3D APIs and TokenHub

Input Not token-priced; text and image inputs are included in the per-generation task charge
Output 15–60 credits per generation; CNY 0.12 per credit, approximately CNY 1.80–7.20 per generation
Tencent AI General Purpose

Long-context coding agents, complex tool-use workflows, productivity automation, document analysis, game development, and scientific reasoning

Context 1M
Reasoning 8/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Preview; currently available as an open-weight model and through Tencent products, Tencent Cloud TokenHub, and OpenRouter

Input CNY 6 per 1M tokens; cached input CNY 0.3 per 1M tokens
Output CNY 18 per 1M tokens

Text-to-panorama and image-to-panorama generation for immersive environments, virtual tours, games, simulations, visualization, and 3D-world pipelines.

Reasoning 1/10
Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and accessible through Tencent Cloud TokenHub

Input USD 1.60 per million tokens; reference usage 481,250 tokens per panorama
Output Approximately USD 0.77 per generated panorama
Tencent AI Embedding

High-volume semantic retrieval, vector search, FAQ matching, text clustering, classification, and cost- or latency-sensitive knowledge-base applications

Context 33K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Status

Current and available through Tencent Cloud TokenHub

Input $0.07 per million text-input tokens
Tencent AI Embedding

High-quality multilingual semantic search, retrieval-augmented generation, enterprise knowledge bases, similarity matching, and text classification.

Context 32K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Status

Current and available through Tencent Cloud TokenHub and the Tencent Cloud embeddings API.

Input USD 0.084 per million input tokens internationally; RMB 0.6 per million input tokens on the China TokenHub pricing page.
Output Not applicable; the model is billed for text input tokens and returns embedding vectors.
Tencent AI Multimodal

Fast cross-modal image-text retrieval, multimodal semantic matching, and video search

Context 33K
Reasoning 1/10
Speed 8/10
Outputs
Embeddings
Capabilities
Image input Video input Multimodal input
Status

Current and available through Tencent Cloud TokenHub

Input USD 0.07 per million text-input tokens; USD 0.098 per million image-input tokens; USD 0.21 per million video-input tokens
Tencent AI Multimodal Embedding
Tencent AI
Kinfra-VL-Embedding

Kinfra-VL-Embedding-8b

High-precision multimodal retrieval, cross-modal image-text search, video search, and semantic matching across text, image, and video collections

Context 33K
Reasoning 1/10
Speed 6/10
Outputs
Embeddings
Capabilities
Image input Video input Multimodal input
Status

Current and available through Tencent Cloud TokenHub

Input USD 0.084 per million text-input tokens; USD 0.126 per million image-input tokens; USD 0.252 per million video-input tokens
Tencent AI Multimodal

Video, image, audio, and text understanding; multimedia summarization; content tagging; structural video analysis; object localization; enterprise media workflows

Context 128K
Reasoning 5/10
Speed 7/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Streaming
Status

Currently accessible; scheduled for shutdown on October 15, 2026 at 00:00 Beijing time on Tencent Cloud TokenHub and ADP

Input ¥1.2 per million tokens
Output ¥3.5 per million tokens
Tencent AI
WAND-Dubbing-Clone

WAND-Dubbing-Clone-V1

Multilingual video translation, voice-preserving dubbing, subtitle translation, online courses, films, and short-form video localization.

Reasoning 2/10
Speed 5/10
Outputs
Video Speech
Capabilities
Audio input Video input Multimodal input
Status

Available

Input CNY 10 per 1 million tokens; usage is resolution-dependent and billed according to video duration and token consumption.
Tencent AI
WAND-Dubbing-Clone

WAND-Dubbing-Clone-v2

Multilingual video localization, translated online courses, short-form video dubbing, film and media localization, and long-video voice-preserving translation

Reasoning 2/10
Speed 6/10
Outputs
Video Speech
Capabilities
Video input Multimodal input
Status

Current and available through Tencent Cloud TokenHub as an asynchronous AI dubbing model

Input ¥0.2061 per second for 720p or lower; ¥0.2311 per second at 1080p; ¥0.2811 per second at 2K or higher. TokenHub lists the underlying rate as ¥10 per million tokens.
Tencent AI Multimodal

Fast everyday image creation, social-media graphics, marketing materials, short-video covers, and reference-guided image generation

Reasoning 1/10
Speed 8/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and available through Tencent Cloud TokenHub

Input 10 CNY per million tokens; input images are free for the Flash tier
Output Reference consumption: 45,000 tokens per 1K image (0.45 CNY), 67,500 tokens per 2K image (0.675 CNY), and 100,800 tokens per 4K image (1.008 CNY)

Low-cost, high-volume text-to-image and reference-to-image generation, including e-commerce product imagery, batch visual assets, and marketing content

Reasoning 1/10
Speed 8/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Currently available managed image-generation model

Input 10 CNY per 1 million tokens; reference pricing is approximately 0.162 CNY per 1K image, 0.18 CNY per 2K image, and 0.225 CNY per 4K image
Output Approximately 0.162 CNY per 1K image, 0.18 CNY per 2K image, and 0.225 CNY per 4K image
Tencent AI Image Generation
Tencent AI
WAND-Vega-Image 1.0

WAND-Vega-Image-1.0-Pro

Professional text-to-image and reference-to-image generation, brand visuals, refined product imagery, high-quality design assets, and 1K-to-4K creative production.

Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and available through Tencent Cloud TokenHub

Input Reference images: first 3 images free; from the 4th image, 0.10 CNY per image
Output 1K: 0.95 CNY per image; 2K: 0.95 CNY per image; 4K: 1.71 CNY per image
Tencent AI
WAND ASR

WAND-ASR-v1

Audio and video transcription, subtitle generation, sentence-level timestamps, and short-form speech-to-text workflows

Speed 7/10
Outputs
Text
Capabilities
Audio input Video input Multimodal input
Status

Current and available through Tencent Cloud TokenHub

Input 10 CNY per 1 million tokens; reference usage is 50 tokens per second, approximately 0.0005 CNY per second

Multi-view and video-to-3D reconstruction, camera and depth estimation, point-map prediction, surface-normal estimation, novel-view synthesis, and 3D Gaussian Splatting workflows

Reasoning 2/10
Speed 7/10
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Legacy but publicly available open-weight model; superseded by WorldMirror-2.0 in Tencent's HY-World 2.0 catalog

xAI Multimodal
xAI
Grok 4

Grok 4.3

Long-context analysis, enterprise agents, research, coding assistance, structured extraction and tool-enabled workflows

Context 1M
Reasoning 9/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

current

Input $1.25 per 1M tokens below 200K prompt tokens; $2.50 per 1M tokens at or above 200K prompt tokens
Output $2.50 per 1M tokens below 200K prompt tokens; $5.00 per 1M tokens at or above 200K prompt tokens
xAI Coding
xAI
Grok 4

Grok 4.5

Software engineering, codebase analysis, technical reasoning, long-context work, tool-using agents, document analysis, and structured workflow automation

Context 500K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

current

Input $2.00 per 1M tokens; $0.30 per 1M cached tokens; long context of 200K tokens or more: $4.00 input and $0.60 cached input per 1M tokens
Output $6.00 per 1M tokens; long context of 200K tokens or more: $12.00 per 1M tokens
xAI Reasoning
xAI
Grok 4

Grok 4.7

Advanced software engineering, long-context reasoning, agentic tool use, research with web or X search, and professional knowledge work.

Context 500K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current; available through the xAI API, Grok Build, Cursor, and selected model gateways.

Input $2.00 per 1 million input tokens below 200,000 prompt tokens; $0.50 per 1 million cached input tokens; $4.00 per 1 million input tokens above 200,000 prompt tokens; $1.00 per 1 million cached input tokens above 200,000 prompt tokens.
Output $6.00 per 1 million output tokens below 200,000 prompt tokens; $12.00 per 1 million output tokens above 200,000 prompt tokens.
xAI Reasoning
xAI
Grok 4.6

Grok 4.6

Agentic coding, long-context software engineering, research, knowledge work, visual analysis, structured extraction, and tool-using workflows.

Context 500K
Reasoning 9/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Current and available; superseded as xAI's flagship by Grok 4.7 but still supported on the xAI API.

Input $2.00 per 1M tokens below 200k prompt tokens; $4.00 per 1M tokens at 200k tokens or more. Cached input is $0.50 per 1M tokens below 200k and $1.00 per 1M tokens at 200k or more.
Output $6.00 per 1M tokens below 200k prompt tokens; $12.00 per 1M tokens at 200k tokens or more.
xAI General Purpose

Low-latency coding, interactive development, and agentic workflows in Cursor or Grok Build

Context 500K
Reasoning 9/10
Speed 10/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use
Status

Available as a faster-serving variant in Cursor and Grok Build; not available as a separate public xAI API model

Input 2x standard Grok 4.7 input-token rate where Fast pricing applies; standard Grok 4.7 API rate is $2 per 1M input tokens
Output 2x standard Grok 4.7 output-token rate where Fast pricing applies; standard Grok 4.7 API rate is $6 per 1M output tokens
xAI Reasoning

Deep research, parallel investigation, long-context analysis, and tool-assisted synthesis

Context 1M
Reasoning 9/10
Speed 5/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

Beta; currently available through the xAI API

Input $1.25 per 1M tokens; $2.50 per 1M tokens for requests at or above 200K input tokens
Output $2.50 per 1M tokens; $5.00 per 1M tokens for requests at or above 200K input tokens
xAI General Purpose

Fast general-purpose text generation, image-aware analysis, coding assistance, structured extraction, tool-calling agents and large-context workflows

Context 1M
Reasoning 6/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

current

Input $1.25 per 1M tokens; $0.20 per 1M cached tokens. Long-context requests at or above 200K prompt tokens: $2.50 per 1M input tokens; $0.40 per 1M cached tokens.
Output $2.50 per 1M tokens; $5.00 per 1M tokens for long-context requests at or above 200K prompt tokens.
xAI Reasoning

Complex reasoning, coding, technical research, long-context document analysis, image understanding, structured responses, and tool-enabled agentic workflows

Context 1M
Reasoning 8/10
Speed 8/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Web search Streaming
Status

current

Input $1.25 per 1M tokens for prompts under 200K tokens; $2.50 per 1M tokens for prompts at or above 200K tokens; cached input is $0.20 or $0.40 per 1M tokens respectively
Output $2.50 per 1M tokens for prompts under 200K tokens; $5.00 per 1M tokens for prompts at or above 200K tokens
xAI Coding
xAI
Grok Build

Grok Build 0.1

Agentic coding, web development, debugging, software engineering workflows, MCP integrations, and fast tool-calling applications.

Context 256K
Reasoning 8/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Current; public beta

Input $1.00 per 1M input tokens; $0.20 per 1M cached input tokens. Long-context pricing for prompts exceeding 200K tokens is $2.00 per 1M input tokens and $0.40 per 1M cached input tokens.
Output $2.00 per 1M output tokens; $4.00 per 1M output tokens when the prompt exceeds 200K tokens.
xAI Other

API-based text-to-image generation, image editing, visual prototyping, creative applications, and per-image billing

Context 1K
Reasoning 1/10
Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Current and available through the xAI API as of September 24, 2026

Input $0.002 per input image; text prompt pricing is not separately stated
Output $0.02 per generated image for 1K and 2K resolution
xAI Multimodal

High-quality text-to-image generation, image editing, multi-reference compositing, marketing graphics, product imagery, typography, and design assets

Outputs
Image
Capabilities
Image input Multimodal input
Status

Generally available

Input $0.01 per input image
Output $0.04-$0.08 per generated image, depending on resolution and quality
xAI Multimodal
xAI
Grok Imagine Image

Grok Imagine Image Quality

High-quality image generation and editing, realistic product and marketing imagery, detailed scenes, creative assets, and images requiring stronger text rendering or prompt adherence.

Reasoning 1/10
Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input
Status

Currently available; scheduled for API retirement on November 2, 2026. After retirement, requests using the slug are served by Grok Imagine Image 2.0 with quality set to low.

Input $0.01 per input image
Output $0.05 per 1K image; $0.07 per 2K image
xAI Other
xAI
Grok Imagine Video

Grok Imagine Video 1.5

Short-form text-to-video, image-to-video, reference-guided video, cinematic prototyping, marketing clips, and audiovisual creative workflows

Speed 8/10
Outputs
Video Speech
Capabilities
Image input Audio input Multimodal input
Status

Generally available

Input $0.01 per image; preset audio input is free
Output $0.08/sec at 480p; $0.14/sec at 720p; $0.25/sec at 1080p
xAI Lightweight
xAI
Grok Imagine Video 1.5

Grok Imagine Video 1.5 Lite

Low-cost text-to-video and image-to-video drafts, social clips, rapid creative iteration, and high-volume generation

Reasoning 0/10
Speed 8/10
Outputs
Video
Capabilities
Image input Multimodal input
Status

Current and available through the xAI Imagine API

Input $0.01 per image input; video input is charged per second where applicable
Output $0.02 per second at 480p; $0.03 per second at 720p; $0.14 per second at 1080p
xAI Multimodal
xAI
Grok Voice Think Fast

Grok Voice Think Fast 2.0

Realtime voice agents, customer support, telephony, sales, multilingual conversations, and tool-enabled spoken workflows

Reasoning 8/10
Speed 9/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input Tool use Web search Streaming
Status

Current and available through the xAI Speech to Speech API

Input $0.08 per minute of audio; $0.004 per text input
Output $0.08 per minute of audio
xAI Speech Recognition
xAI
Grok Voice Transcribe

Grok Voice Transcribe 1.0

Batch and real-time multilingual speech transcription, dictation, voice assistants, accessibility, meetings, and customer-support audio

Speed 8/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Currently accessible; original speech-to-text model; scheduled for deprecation in the coming weeks

Input $0.10 per hour of audio for REST batch transcription; $0.20 per hour of audio for streaming
Output Included in the transcription service price; output is text transcript data
xAI Other
xAI
Grok Voice Transcribe

Grok Voice Transcribe 2.0

Batch and real-time transcription of multilingual, noisy, conversational, telephony, and voice-agent audio

Reasoning 0/10
Speed 9/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current

Input $0.10 per hour of audio for batch REST transcription; $0.20 per hour of audio for streaming transcription
Xiaomi HyperAI Multimodal

Local audio understanding, speech-to-text dialogue, spoken conversational agents, and controllable text-to-speech research

Context 8K
Reasoning 7/10
Speed 4/10
Outputs
Text Speech
Capabilities
Audio input Multimodal input
Status

Current open-weight model; locally downloadable and usable

Input No official hosted API price published; self-hosted weights
Output No official hosted API price published; self-hosted weights
Xiaomi HyperAI Multimodal
Xiaomi HyperAI
MiMo-Embodied

MiMo-Embodied-7B

Embodied AI research, autonomous-driving perception and planning, spatial reasoning, affordance prediction, robot navigation, and multimodal visual analysis

Context 128K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input
Status

Active open-weight model

Xiaomi HyperAI
MiMo-V2.5

MiMo-V2.5-ASR

Speech-to-text transcription for Chinese, English, regional Chinese dialects, code-switched speech, lyrics, noisy recordings, and multi-speaker conversations

Reasoning 3/10
Speed 7/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current and accessible through the Xiaomi MiMo API; open-source code and weights available

Input $0.074 per hour of input audio for overseas API usage; ¥0.5 per hour on the domestic pricing page
Output No separate output charge documented; billing is based on input-audio duration
Xiaomi HyperAI
MiMo-V2.5-TTS

MiMo-V2.5-TTS

Expressive text-to-speech, audiobooks, podcasts, dubbing, character dialogue, voice interfaces, narrated content, and stylized speech or singing.

Context 8K
Reasoning 1/10
Speed 7/10
Outputs
Speech Music
Capabilities
Streaming
Status

Available

Input Free for a limited time
Output Free for a limited time
Xiaomi HyperAI Multimodal

Local coding agents, terminal automation, cybersecurity experimentation, visual coding and agentic reinforcement-learning research

Context 262K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Image input Multimodal input Tool use Streaming
Status

Current open-weight SFT checkpoint

Input No official hosted API price; open-weight download
Output No official hosted API price; open-weight download
Xiaomi HyperAI Multimodal

High-volume multimodal API workloads, coding assistants, agent automation, long-context document and repository analysis, tool-using workflows, and cost-sensitive professional applications.

Context 1M
Reasoning 8/10
Speed 9/10
Outputs
Text Actions
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current; open-weight and available through the Xiaomi MiMo API, MiMo Studio, MiMo Desktop, MiMo Code, and other supported integrations.

Input $0.14 per million tokens for cache-miss input; $0.0028 per million tokens for cache-hit input.
Output $0.28 per million tokens.
Xiaomi HyperAI Multimodal
Xiaomi HyperAI
MiMo-V2.6

MiMo-V2.6-Pro

Long-horizon agents, coding, cybersecurity, research, computer use, multimodal analysis, and complex multi-step workflows

Context 1M
Reasoning 9/10
Speed 7/10
Outputs
Text Actions
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current; open-source weights and API access available

Input $0.435 per 1M tokens cache miss; $0.0036 per 1M tokens cache hit; batch input $0.2175 per 1M tokens
Output $0.87 per 1M tokens; batch output $0.435 per 1M tokens

Latency-sensitive multimodal reasoning, coding, tool-use, long-context research, and interactive agent workflows

Context 1M
Reasoning 9/10
Speed 10/10
Outputs
Text
Capabilities
Image input Audio input Video input Multimodal input Tool use
Status

Current; high-speed inference variant of MiMo-V2.6-Pro

Input $0.036 per million cached-input tokens; $4.35 per million uncached-input tokens
Output $8.70 per million output tokens
Xiaomi HyperAI Multimodal

Local image and video understanding, multimodal reasoning, visual question answering, OCR-oriented tasks, visual grounding, and GUI analysis

Context 128K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Image input Video input Multimodal input
Status

Current open-weight model; publicly available for download and local deployment

Input No official first-party hosted API pricing found; open-weight download
Output No official first-party hosted API pricing found; open-weight download
Xiaomi HyperAI Multimodal

Self-hosted image and video understanding, multimodal reasoning research, supervised fine-tuning, and reinforcement-learning experimentation

Context 128K
Reasoning 7/10
Speed 6/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Fine-tuning
Status

Current open-weight checkpoint; available for download

Input No official hosted API pricing; downloadable weights are available under the MIT license
Output No official hosted API pricing; downloadable weights are available under the MIT license
Xiaomi HyperAI
MiMo V2.5

MiMo-V2.5-Pro

Long-horizon agent workflows, repository-scale coding, complex software engineering, tool-driven automation, and very large documents

Context 1M
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Web search Fine-tuning Streaming
Status

Deprecated; API model name scheduled to expire on 2026-10-21 at 10:00 Beijing Time

Input ¥0.025 per million tokens cache hit; ¥3 per million tokens cache miss; equivalent USD prices $0.0036 and $0.435 per million tokens
Output ¥6 per million tokens; equivalent USD price $0.87 per million tokens
Yandex AI General Purpose

Russian-language conversational assistants, long-context dialogue, RAG, document analysis, information extraction, reporting, and complex business text generation

Context 131K
Reasoning 8/10
Speed 7/10
Outputs
Text
Capabilities
Streaming
Status

Current and available in Yandex AI Studio

Input $0.00204918 per 1,000 input tokens in asynchronous mode, before VAT
Output $0.0083606544 per 1,000 output tokens in asynchronous mode, before VAT
Yandex AI Multimodal
Yandex AI
Alice AI ART

Alice AI ART

Text-to-image generation, Russian-language visual content, illustrations, graphic design, advertising assets, presentations, landing pages, and product cards

Context 500
Reasoning 1/10
Speed 7/10
Outputs
Image
Capabilities
Image input Multimodal input Streaming
Status

Current

Input $0.0182786856 per image-generation request, without VAT
Yandex AI General Purpose

Russian-language enterprise text generation, document analysis, RAG, structured extraction, rewriting, classification, reporting, and tool-augmented business assistants

Context 33K
Reasoning 7/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Web search Fine-tuning Streaming
Status

Current and available through Yandex Cloud AI Studio; explicit model URI recommended. The RC alias remains available until the end of support.

Input 0.8 RUB per 1,000 input tokens including VAT; 0.8 RUB per 1,000 cached input tokens; 0.2 RUB per 1,000 tool tokens. Asynchronous mode: 0.41 RUB per 1,000 input tokens including VAT.
Output 0.8 RUB per 1,000 output tokens including VAT. Asynchronous mode: 0.41 RUB per 1,000 output tokens including VAT.
Z.ai Video Generation
Z.ai
CogVideoX

CogVideoX-3

Short-form text-to-video, image animation, start-and-end-frame transitions, advertising, marketing, realistic scenes, and 3D-style video generation

Speed 7/10
Outputs
Video Audio
Capabilities
Image input Multimodal input
Status

Current and available through the Z.ai video-generation API

Input $0.20 per video
Output $0.20 per video
Z.ai Reasoning
Z.ai
GLM-5

GLM-5.2

Long-horizon software engineering, repository-scale coding, code agents, complex refactoring, large-document analysis, research implementation, and tool-using workflows.

Context 1M
Reasoning 9/10
Speed 7/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Active; still listed in Z.AI API pricing and available as downloadable open weights, although GLM-5.3 is the newer flagship successor.

Input $1.40 per 1 million input tokens; cached input $0.26 per 1 million tokens
Output $4.40 per 1 million output tokens
Z.ai Reasoning
Z.ai
GLM-5

GLM-5.3

Complex software engineering, large codebases, long-horizon coding agents, terminal workflows, technical research, and authorized cybersecurity analysis

Context 1M
Reasoning 9/10
Speed 6/10
Outputs
Text
Capabilities
Tool use Fine-tuning Streaming
Status

Current and available

Z.ai Multimodal

Multimodal coding agents, screenshot and interface understanding, long-context document work, visual verification, tool-assisted research, and cost-sensitive agent workflows

Context 1M
Reasoning 9/10
Speed 9/10
Outputs
Text
Capabilities
Image input Video input Multimodal input Tool use Streaming
Status

Current and available through the Z.ai API, GLM Coding Plan, and publicly released open weights

Input $0.15 per 1 million input tokens; cached input $0.03 per 1 million tokens
Output $0.50 per 1 million output tokens
Z.ai Other

Hosted multilingual speech-to-text, short real-time transcriptions, captions, meeting snippets, voice input, customer-service audio, and terminology-aware transcription

Speed 8/10
Outputs
Text
Capabilities
Audio input Streaming
Status

Current and publicly available through the Z.AI audio transcription API

Input Approximately CNY 0.06 per minute; verify current account pricing
Output Not separately priced; transcription is billed by audio duration according to reported listings
Z.ai Image Generation
Z.ai
GLM-Image

GLM-Image

Text-heavy posters, presentation graphics, educational diagrams, social-media layouts, image editing, and open-weight image-generation research

Speed 3/10
Outputs
Image
Capabilities
Image input Multimodal input Fine-tuning
Status

Current; available through the Z.ai API and as downloadable open weights

Input $0.015 per image
Output $0.015 per generated image
Z.ai Other
Z.ai
GLM-OCR

GLM-OCR

High-volume OCR, PDF parsing, table and formula recognition, handwriting, receipts, forms, information extraction, and local document-processing pipelines

Reasoning 2/10
Speed 9/10
Outputs
Text
Capabilities
Image input Multimodal input Fine-tuning
Status

Current; open-weight; available through hosted API and self-hosted deployment

Input $0.03 per 1 million tokens
Output $0.03 per 1 million tokens