Yi-6B
Local bilingual text generation, research, fine-tuning, offline applications, and resource-conscious deployment
Available open-weight model; older first-generation Yi base model
Explore AI models across providers and compare what they are designed to do. Browse language and reasoning models, coding models, multimodal systems, image and video models, audio and speech models, embeddings, realtime models and other specialized AI systems.
Local bilingual text generation, research, fine-tuning, offline applications, and resource-conscious deployment
Available open-weight model; older first-generation Yi base model
Long-document completion, local deployment, English-Chinese text generation, research, and downstream fine-tuning
Available as downloadable open weights; legacy-generation model
Local bilingual English-Chinese chat, personal projects, academic experimentation, and fine-tuning on modest open-weight infrastructure
Open-weight and downloadable; legacy historical model
Local text generation, code completion, mathematics, bilingual English-Chinese applications, research, and downstream fine-tuning
Available as an open-weight downloadable model; older Yi-generation model with no verified first-party hosted API offering for this exact model
Long-context document processing, code generation, mathematics, bilingual English-Chinese text generation, local inference, and domain fine-tuning
Available open-weight base model; legacy generation but still publicly downloadable
Self-hosted bilingual text generation, research, custom fine-tuning, coding experiments, and English-Chinese applications
Available open-weight model; older-generation base checkpoint
Long-document analysis, English-Chinese generation, retrieval experiments, research, and self-hosted fine-tuning
Available as downloadable open-weight model; no official hosted API availability or retirement date verified
Self-hosted bilingual assistants, English-Chinese dialogue, open-weight LLM research, private inference, quantization, and fine-tuning
Open-weight; downloadable; legacy-generation model with no verified provider-managed hosted API availability
Long-context chat, complex text analysis, multilingual generation, prediction, and general-purpose enterprise language applications.
Available through 01.AI API documentation; current lifecycle details are not explicitly stated by the provider.
Cost-sensitive hosted chat, Chinese-English generation, coding, mathematics, reasoning, summarization and high-volume API workloads
Proprietary hosted API model; current public availability and lifecycle status are not clearly documented
Local text generation, self-hosted applications, experimentation, domain adaptation, and fine-tuning
Available as an open-weight downloadable model for self-hosted and local inference; no hosted inference provider is currently listed on its Hugging Face model page.
Local conversational applications, instruction following, lightweight coding, bilingual English-Chinese text generation, experimentation, and self-hosted inference
Open-weight; accessible for local deployment; no explicit retirement date found
Local text generation, research, fine-tuning, Chinese-English applications, and cost-sensitive self-hosted deployments
Open-weight and downloadable; no official deprecation or shutdown date found
Self-hosted conversational assistants, local text generation, lightweight coding help, multilingual experimentation, and fine-tuning research.
Available open-weight model; official weights remain accessible, with no exact provider-published deprecation or shutdown date verified.
Local bilingual text generation, research, domain adaptation, and fine-tuning
Open-weight and currently downloadable; standard 4K-context base checkpoint
Self-hosted bilingual assistants, English-Chinese text generation, coding support, mathematics, reasoning, fine-tuning, and privacy-sensitive deployments
Open-weight model remains available for self-hosted deployment; 01.AI hosted model-platform API service ended on 2026-09-03
Local code completion, code generation, multilingual programming tasks, long-context source-code analysis, and downstream fine-tuning
Available as downloadable open-weight model; no exact first-party hosted API availability verified
Local code generation, completion, debugging, code explanation, lightweight IDE assistants, and model fine-tuning experiments
Available open-weight model
Local code generation, completion, editing, repository-scale context, multilingual programming, and fine-tuning
Available open-weight base model
Self-hosted coding assistance, code generation, debugging, code explanation, code translation, and long-context repository analysis
Open-weight and publicly available; no verified official hosted API listing for this exact model
Local bilingual image understanding, visual question answering, OCR-oriented image analysis, and image-to-text applications
Open-weight multimodal model family; current accessibility and active maintenance are not clearly documented by 01.AI
Local bilingual image understanding, visual question answering, OCR-assisted extraction, image summarization, and lightweight multimodal experimentation
Available as an open-weight self-hosted model; no current official hosted inference provider is listed on its model page
Self-hosted bilingual image understanding, visual question answering, image text recognition, and research applications with substantial GPU capacity
Available as open-weight downloadable model; no current first-party hosted API availability verified
Function calling, tool selection, agent orchestration, and structured workflow automation
Documented by 01.AI; current live availability not independently verified
Low-cost short-context chat, text generation, summarization, classification, and general language applications
Documented in the 01.AI API catalog; current operational availability should be verified through the provider
Long-context text generation, open-weight research, experimentation, and private or self-hosted deployment
Legacy open-weight model; downloadable and accessible through AI21 Labs' official Hugging Face repository
Long-document analysis, grounded generation, enterprise RAG, document summarization, information extraction, private deployment, and multilingual text workflows.
Current and available; the API endpoint jambalarge-1.7 points to the dated snapshot jambalarge-1.7-2025-07.
Long-context document analysis, RAG, structured generation, function calling, multilingual enterprise assistants, and private deployment
Legacy; superseded by newer Jamba releases, including AI21-Jamba-Mini-1.7. Public model weights remain available.
Long-context document analysis, retrieval-augmented generation, enterprise assistants, structured text generation, multilingual workflows, and self-hosted deployments
Legacy and superseded by newer Jamba Large releases; downloadable open-weight checkpoint remains available
Long-context retrieval-augmented generation, enterprise document analysis, grounded question answering, structured extraction, classification, and private deployment
Available; older Jamba 1.6 generation with Jamba 1.7 available as a newer successor
Long-context RAG, grounded question answering, enterprise document processing, classification, structured text generation, function calling and privacy-sensitive private deployments.
Legacy/open-weight model; downloadable from AI21's official Hugging Face repository and available for self-managed deployment, but not featured among AI21's current primary Jamba models as of September 25, 2026.
Long-document analysis, enterprise RAG, grounded question answering, structured text generation, private deployment, and cost-sensitive text workflows
Available as an AI21 open-weight model; hosted API availability is not currently verified
Long-context RAG, grounded enterprise question answering, document extraction, local inference, on-device assistants, and lightweight agent workflows
Current and available; open-weight release
Long-context enterprise question answering, grounded generation, document analysis, instruction-heavy workflows, RAG systems and self-hosted deployments
Current open-weight model; available through Hugging Face and AI21 Studio
Local reasoning, long-context document analysis, private RAG, extraction, coding assistance, and lightweight agent controllers
Current open-weight model; available for download and local inference
Fast German-language document quality classification and large-scale web-data filtering
Available; open-weight research model; not deployed by a Hugging Face Inference Provider
Fast local classification of German documents for grammar-oriented corpus filtering and data curation
Available as an open-weight model repository; not deployed by Hugging Face Inference Providers
German grammar-related text-quality filtering, dataset curation, and binary classification of web text
Available as an open-weight Hugging Face model
German web-document quality scoring, dataset filtering, ranking, and pretraining-data curation
Available open-weight model
Research on tokenizer-free language modeling, English-German text generation, multilingual NLP, open-weight deployment, and custom model adaptation
Available as open-weight research software through Hugging Face; not deployed by an inference provider
English and German text generation, instruction-following research, tokenizer-free language-model research, and self-hosted non-commercial experimentation
Available as an open-weight research model
English and German instruction following, tokenizer-free language-model research, text-compression experiments, and self-hosted open-weight deployments
Available open-weight research release
Historical multilingual text completion, language-model research, and semantic representation workflows
Legacy; current public availability is not clearly documented
Legacy multilingual text generation, conversational applications, and explainability-oriented workflows using Aleph Alpha's first Luminous generation
Legacy; current public availability not verified
Semantic search, information retrieval, query-document matching, clustering, classification, similarity scoring, and text feature extraction
Legacy; current API availability unverified
Historical multilingual text-completion research and general language-task experiments
Historical or legacy; current availability not verified
Steerable multilingual text generation, summarization, classification, question answering, and explainability-oriented enterprise workflows
Available; first-generation Luminous control model
Historical multilingual text completion, language understanding research, and compatibility work involving Aleph Alpha’s original Luminous API
Legacy generation; historical API and Playground availability documented, current public availability unverified
Zero-shot multilingual text generation, instruction following, classification, conversational prototypes, and explainability-oriented enterprise workflows.
Legacy model; documented in Aleph Alpha SDK materials, with current endpoint availability dependent on provider access and deployment status.
Vision-language research, image captioning, visual question answering, and experiments with adapter-based multimodal fine-tuning
Legacy research/demo model; publicly released checkpoint and source code remain available, but it is not documented as a current hosted commercial model
Multilingual information retrieval, semantic search, reranking, clustering, and instruction-guided text embeddings
Available as downloadable open-weight model; customer and on-premises deployment options documented
Compact vector representations for semantic search, information retrieval, reranking, clustering, and similarity-based classification
Available open-weight checkpoint; no hosted inference deployment or provider API pricing was verified
Concise multilingual generation, summarization, extraction, domain-specific text workflows, and self-hosted deployment
Available as an open-weight model; also referenced in Aleph Alpha API tooling
Multilingual text generation, classification, summarization, question answering, engineering and automotive applications, and safety-conscious research deployments
Available open-weight release; safety-aligned variant
English and German instruction following, multilingual research, tokenizer-free language-model experimentation, and self-hosted deployment
Available as downloadable open-weight research software
English and German instruction following, multilingual text generation, German-language applications, and research into tokenizer-free language models
Current open-weight research release; downloadable from Hugging Face
Research on tokenizer-free language modeling, English and German text generation, long-context experiments, and self-hosted non-commercial applications
Current publicly downloadable open-weight research model
Local image understanding, visual question answering, captioning, document and chart analysis, counting, pointing, and multimodal research
Available as an open-weight downloadable checkpoint
Local image-and-text understanding, visual question answering, document and chart analysis, image captioning, and multimodal research
Available open-weight checkpoint; preview release
Self-hosted image understanding, visual question answering, captioning, image grounding, pointing, counting, and multimodal research.
Available open-weight legacy research model; newer Molmo 2 models are Ai2's current successor family.
Local image understanding, visual question answering, image captioning, document and chart analysis, counting, and experimentation with open multimodal models
Available open-weight preview checkpoint
Efficient local image and video understanding, visual grounding, pointing, captioning, counting, tracking, and multimodal research
Current open-weight model
Open multimodal research, video question answering, visual grounding, pointing, counting, captioning, object tracking, document understanding, and robotics perception
Current open-weight model; downloadable checkpoint
Open multimodal research, local image and video understanding, visual grounding, pointing, counting, captioning, and tracking
Current open-weight model
Open research and downstream fine-tuning for vision-guided robotic manipulation, spatial reasoning, trajectory planning, and robot action prediction
Available open-weight research model; superseded by MolmoAct2 as Ai2's newer MolmoAct generation
Robotic manipulation research, action reasoning, downstream mid-training, and reproducing zero-shot SimplerEnv experiments
Available open-weight preview checkpoint
Open robotics research, visual action reasoning, robot-manipulation experiments, and fine-tuning on custom robot datasets
Available open-weight preview checkpoint; intended for fine-tuning and downstream post-training
Open robotics research, embodied visual reasoning, robot-policy fine-tuning, and manipulation tasks on supported or closely related hardware
Current open-weight foundation checkpoint
Depth-aware robot manipulation research, embodied reasoning, and fine-tuning vision-language-action policies for target robot embodiments.
Current open-weight foundation checkpoint
Language-guided 3D trajectory forecasting, robotics planning research, and motion-conditioned video generation
Documented research variant; no separately released public checkpoint identified in the current first-party model collection
Open research and applications requiring image or video grounding, visual pointing, object localization, counting, tracking, spatial reasoning, and multimodal analysis.
Current open-weight model
GUI screenshot grounding, screen-element localization, computer-use perception, and research on pointing models
Current open-weight research model
Video object pointing, temporal grounding, object counting, tracking, and research on language-guided video understanding
Current open-weight research model
Open research, reproducible language-model experiments, local deployment, English text generation, and fine-tuning
Available open-weight model; not formally deprecated or retired
Local inference, open-model research, benchmarking, continued pretraining, and task-specific fine-tuning.
Available as a downloadable open-weight model; no official first-party hosted API identified.
Fully open language-model research, local inference, reproducible training experiments, evaluation, and downstream fine-tuning
Current downloadable open-weight model; no first-party hosted API deployment verified
Continued pretraining, supervised fine-tuning, reinforcement-learning research, language-model research, custom deployment, and applications requiring a transparent open training pipeline.
Current and downloadable open-weight model
Open, self-hosted chat assistants, instruction following, local coding help, tool-calling agents, and model research
Current open-weight model; available for download and self-hosted inference
Open research, mathematical reasoning, code generation, logic, long-context experimentation, and local self-hosted inference
Active open-weight model
Open-weight research, continued pretraining, fine-tuning, programming, mathematics, reading comprehension, and long-context language-model experiments
Current; open-weight and downloadable
Open-weight chat, tool-using assistants, multi-turn dialogue, synthetic data generation, self-hosting, and model research
Available open-weight checkpoint; superseded by Olmo 3.1 32B Instruct for newer instruction-tuned deployments
Open-model reasoning research, mathematics, coding, long-context analysis, and self-hosted deployment
Available as an open-weight model; superseded by the newer Olmo 3.1 32B Think model
Math-focused RLVR research, reinforcement-learning experiments, verifiable reward design, open-model evaluation, and further fine-tuning
Current open-weight research checkpoint
Mathematical reasoning, coding, complex multi-step tasks, open-model research, local deployment, and reinforcement-learning research
Current open-weight model
Open coding-model research, reinforcement learning from verifiable rewards, code-generation experiments, automated evaluation, and fine-tuning
Current, open-weight, experimental RL-Zero coding checkpoint
Self-hosted chat assistants, instruction following, tool-use agents, synthetic data, and open-model research
Current; open-weight and downloadable
English audio transcription, captions, meetings, lectures, podcasts, accessibility applications, and open ASR research
Current open-weight model
Open English speech-to-text transcription, captioning, meetings, lectures, calls, podcasts, and local ASR research
Available open-weight model
Open, self-hosted English transcription for meetings, lectures, calls, podcasts, accessibility, and speech-recognition research
Current open-weight model; downloadable and usable for local inference
English short- and long-form speech transcription, meeting and call transcription, lecture captioning, podcast processing, broadcast transcription, and speech analytics
Current open-weight model
Open English speech transcription, meeting and podcast transcription, captioning, timestamped audio indexing, and ASR research
Current open-weight model; available for download and self-managed inference
Efficient self-hosted English transcription, speech research, edge-oriented experiments, and applications where a small open ASR checkpoint is preferred.
Available open-weight checkpoint
Open-model research, efficient local text generation, language-model evaluation, and fine-tuning
Available open-weight release
Efficient satellite-image and Earth-observation embeddings, remote-sensing research, and downstream classification or segmentation
Current; downloadable open-weight research model
Satellite-image and Earth-observation embeddings, remote-sensing representation learning, geospatial classification, segmentation, and downstream fine-tuning
Current open-weight model
Efficient Earth observation embeddings, satellite image and time-series representation learning, geospatial classification, segmentation, and large-scale remote-sensing inference
Current open-weight model
Satellite-image embeddings, remote-sensing representation learning, geospatial segmentation, land-cover analysis, and Earth observation research
Current open-weight Earth observation foundation model
Open-weight language-model research, long-context text generation, local deployment, continued pretraining, and fine-tuning
Current open-weight model
Self-hosted instruction following, open post-training research, mathematics, coding, evaluation and domain adaptation
Available for download; superseded by Llama-3.1-Tulu-3.1-8B
Self-hosted instruction following, reasoning, mathematics, coding, open-model research, and reproducible post-training experiments
Available open-weight model
Large-scale research, instruction-following evaluation, open-weight post-training research, mathematical reasoning, coding benchmarks, and self-hosted experimentation.
Available as an open-weight research and educational model; no hosted inference provider deployment is currently listed on its Hugging Face model page.
Multimodal research, vision-language experiments, image generation, visual question answering, dense computer-vision tasks, and academic benchmarking
Open-weight research release; publicly available inference code and checkpoints; legacy relative to Ai2's newer multimodal models
Multimodal research, image understanding and generation, audio and video understanding, spatial prediction, embodied AI and robotic-manipulation experiments, and self-hosted academic prototyping.
Open-weight research release; publicly accessible checkpoints and source code; no official hosted inference API or commercial token pricing identified.
Open-vocabulary monocular 3D detection, spatial perception, robotics research, augmented reality, and lifting 2D prompts into metric 3D boxes
Current open-weight research model
Enterprise image generation and editing, product visualization, advertising and marketing assets, image variations, background removal, virtual try-on, and brand or subject-consistent visual content.
Legacy; currently accessible as of 2026-09-25; scheduled for end of life on 2026-09-30
Low-cost multimodal document analysis, image and video understanding, visual question answering, summarization, RAG, and tool-enabled agents.
Active
High-volume, low-latency text classification, summarization, translation, extraction, routing, FAQs and narrowly defined fine-tuned tasks
active
Long-context multimodal analysis, enterprise document workflows, complex tool calling, agentic orchestration, codebase analysis, and teacher-model distillation before retirement.
End-of-Life
Enterprise multimodal applications, document analysis, visual question answering, video understanding, long-context summarization, RAG, and tool-using assistants
Active
Short-form advertising, marketing concepts, product visualization, storyboards, social video drafts, and image-guided cinematic clips.
Active through the current Nova Reel 1.1 workflow; the original Nova Reel v1.0 model is legacy and scheduled for end of life on 2026-09-30.
Real-time voice assistants, customer-service automation, interactive education, language learning, and speech-enabled enterprise workflows
Retired; legacy model with official end-of-life date of September 14, 2026
High-volume multimodal applications, document and video analysis, customer service, business automation, software engineering, long-context workflows, and cost-sensitive AI agents.
Active; generally available through Amazon Bedrock. AWS states EOL is no sooner than 2026-12-02.
Real-time voice assistants, customer-service automation, telephony, interactive learning, multilingual conversations, and tool-enabled speech agents.
Active; EOL no sooner than 2026-12-02
Browser automation, visual UI navigation, repetitive web workflows, agentic QA, tool-oriented tasks, and human-supervised enterprise processes
Generally available
Cross-modal semantic search, multimodal RAG, digital asset discovery, recommendations, classification, and clustering
Active; generally available through Amazon Bedrock
Multimodal search, text-to-image retrieval, image similarity, visual recommendations, personalization, and image-text matching.
Active
Text-to-image generation, image editing, reference-guided composition, background removal, color-controlled visuals, image variations, and subject-consistent branded content.
Retired; AWS lists the model as legacy with an end-of-life date of 2026-06-30.
Semantic search, vector indexing, retrieval-augmented generation, personalization, clustering, classification, and recommendation pipelines.
Available; older G1/V1 text-embedding generation
Cost-efficient text embeddings for semantic search, RAG, document retrieval, classification, clustering, reranking, and recommendations
Active
Semantic search, vector retrieval, recommendation, semantic matching, knowledge bases, and retrieval-augmented generation
Current and accessible through Baidu Qianfan and AI Studio embedding APIs
Multimodal understanding, long-context Chinese and English applications, complex reasoning, coding, tool-enabled agents, and enterprise workloads
Current and accessible through Baidu Qianfan as of September 25, 2026
Realistic text-to-image generation, reference-grounded visual creation, commercial-style imagery, and applications needing reduced generative-artificiality.
Available; listed in Baidu AI Cloud's inference-service model updates and supported for offline batch inference in Qianfan ModelBuilder documentation.
Object removal, masked image repainting, image variation, and batch image-editing workflows.
Available with access dependent on Baidu Qianfan service configuration; listed for Qianfan batch inference. Related AI作画-iRAG版 API sales stopped on 2026-04-30, but that product notice does not explicitly state that the Qianfan ERNIE-iRAG-Edit endpoi
Lightweight local text generation, Chinese and English language experimentation, compact conversational systems, and domain adaptation
Current open-weight model
Fast, low-cost Chinese text generation, long-context applications, enterprise agents, content creation, reasoning, and code assistance.
Current and accessible through Baidu Qianfan; canonical API endpoint is ernie-4.5-turbo-128k.
Cost-efficient multimodal understanding, image and video analysis, OCR, document comprehension, translation, visual question answering, and code-related tasks
Current and available
Local or self-hosted multimodal applications, visual question answering, document and chart understanding, image analysis, video understanding, and efficient vision-language inference
Current open-weight model; publicly released under the Apache 2.0 license
Agentic workflows, web-search-assisted tasks, reasoning, Chinese-language knowledge work, creative writing, and general-purpose assistant applications.
Current; officially released and accessible through Baidu's ERNIE website and AI Studio playground. Public Qianfan API availability for an exact ERNIE 5.1 endpoint was not verified.
Code completion, code generation, unit-test generation, code optimization, code explanation, and fine-tuning for software-development workflows.
Retired; Qianfan ModelBuilder listed August 14, 2025 as the retirement date.
Chinese-language reasoning, long-form analysis, complex calculations, literary and document writing, agent workflows, and function calling
Ready; currently accessible through Baidu Qianfan as ernie-x1-turbo-32k
Deep reasoning, Chinese and English question answering, mathematics, coding, factual responses, tool calling, web-grounded applications, and agent workflows
Current preview model
Low-cost image-to-video generation, short social clips, product animation, marketing assets, and turning still images into dynamic scenes
Current and available through Baidu Qianfan API
Chinese image-to-video generation, audiovisual storytelling, marketing videos, multi-person dialogue, synchronized speech, sound effects, and cinematic short-form content
Current model family; concrete Qianfan variants include Turbo, Lite, Pro, Turbo-I2V-Audio, and Turbo-I2V-Effect
Low-cost text-to-image generation, marketing visuals, creative assets, illustrations, and rapid image prototyping
Current and available
Multilingual OCR and structured parsing of complex documents, including tables, formulas, charts, seals, scanned pages, warped documents, and screen photographs
Available; superseded by PaddleOCR-VL-1.6
Document parsing, OCR, layout analysis, table extraction, and structured document understanding
Fast enterprise question answering, summarization, workflow nodes, agent response generation, and text processing with a 32K context window
Available through Baidu’s documented Wenxin Workshop API; not found in the current Qianfan V2 model-list documentation, so new integrations should verify legacy endpoint compatibility.
Long-context Agent planning, task decomposition, component selection, and function-calling workflows.
Current platform-listed planning model; standalone lifecycle status is not separately published.
Enterprise agent workflows, intent recognition, instruction routing, and tool-calling tasks
Legacy; current public availability not verified
Fast enterprise question answering, lightweight agent planning, component selection, and streaming text applications
Current and listed as available in Baidu Qianfan documentation; fast agent-oriented model
Long-horizon, high-precision dexterous robot manipulation and real-world VLA policy specialization
Current research model/framework; public API availability not documented
Protein and biomolecular complex structure prediction, computational biology research, molecular design workflows, and self-hosted scientific inference
Active open-source biomolecular structure prediction project with multiple model variants, including Protenix-v1 and Protenix-v2
Multimodal agent workflows, search and information retrieval, coding agents, GUI interaction, image and video understanding, complex instruction following, and long-context business tasks.
Beta; accessible through BytePlus ModelArk as seed-1-8-251228. The separate LAS multimodal deep-thinking operator using Seed1.8 ended service on 2026-09-20.
Self-hosted general language modeling, long-context research, reasoning experiments, coding assistance, summarization, and foundation-model fine-tuning
Available open-weight model
Self-hosted long-context reasoning, coding, agentic tool use, research, question answering, summarization, and general text generation
Current open-weight model; downloadable under the Apache-2.0 license
Research, custom post-training, long-context processing, coding experiments, and self-hosted foundation-model deployment
Current open-weight downloadable foundation model
General-purpose Chinese and multilingual assistance, coding, reasoning, image and document understanding, and voice-interaction applications.
Retired; the primary doubao-1-5-pro-32k-250115 deployment was scheduled to shut down on 2026-09-21 at 14:00 China Standard Time.
Visual reasoning, image and video understanding, OCR, visual grounding, GUI-agent research, gameplay analysis, and multimodal benchmark evaluation
Retired; Volcano Engine service ended on 2026-03-31
Cross-modal semantic search, text-image retrieval, video retrieval, multimodal knowledge bases, classification, clustering, and recommendation
Current and available through Volcano Engine
Multimodal document analysis, visual question answering, coding, mathematics, general reasoning, long-context analysis and adaptive-thinking applications
Deprecated; new endpoint creation stopped on 2026-09-24; existing service scheduled for automatic migration or replacement on 2026-11-24
Deep reasoning, coding, mathematics, logical analysis, document understanding, and visual reasoning over images or videos
Deprecated; scheduled for retirement and migration to a Seed2.0 model
Cost-conscious production applications requiring long-context multimodal understanding, document and video analysis, coding assistance, tool use, GUI automation, and structured extraction.
Current; latest documented release seed-2-0-lite-260428
High-concurrency inference, batch generation, classification, extraction, summarization, and cost-sensitive multimodal workloads
Current; available through ByteDance's Volcano Engine model API
Complex multimodal reasoning, long-chain agent workflows, visual and video analysis, document understanding, scientific research support, coding, and enterprise automation.
Active and currently accessible through Volcano Engine Ark; canonical deployment version 260215
Complex agent workflows, high-value office and research tasks, long-horizon coding, document and visual analysis, video understanding, and tool-enabled productivity automation.
Current and officially released; available through ByteDance Seed, Doubao, and Volcano Engine API channels.
Fast multimodal agents, coding assistants, document and video analysis, tool-calling workflows, structured data extraction, and cost-sensitive production applications
Current and available through Volcano Engine Ark; API model identifier doubao-seed-2-1-turbo-260628
Long-form audiovisual storytelling, text-to-video, reference-based video generation, creative production, advertising, education, industrial simulation, and video editing.
Current; available through ByteDance platforms including Jimeng AI, Doubao Pro, and the Seed platform. API access through BytePlus ModelArk was announced as forthcoming.
Multimodal text-to-video and reference-based video creation, cinematic short clips, video editing and extension, multi-shot storytelling, and synchronized audio-video production.
Current official model page remains available; older generation superseded by Seedance 2.5
Full-scene audio creation, expressive voice generation, dubbing, dialogue, sound effects, ambience, advertising, games, podcasts, and multilingual audio production
Currently available through BytePlus
High-throughput code generation research, diffusion-language-model evaluation, and code-editing experiments
Experimental research preview
Embodied robotics research, long-horizon manipulation, bimanual control, dexterous object handling, and adapting robot policies to new objects and tasks
Current officially documented research model; no public hosted API, commercial pricing, or downloadable weights verified
Real-time audio-visual assistants, scene-aware guidance, live explanation, interactive learning, accessibility, and proactive multimodal collaboration
Current; fully rolled out for large-scale audio-visual full-duplex deployment
Professional image generation and editing, high-density infographics, advertising and e-commerce creative, multilingual visual content, precise spatial edits, layer separation, and multi-reference image workflows
Current and publicly documented; available through BytePlus ModelArk and ByteDance Seed services
Open-weight computer-use research, GUI grounding, browser automation prototypes, screenshot-based interface interaction, and visual action-model experimentation
Available open-weight research release; superseded by UI-TARS-2 for the provider's newer UI-TARS development line
Research, local text generation, fine-tuning experiments, benchmarking, and lightweight English NLP prototypes
Available open-weight research model
Open-weight language-model research, local text generation, fine-tuning experiments, scaling-law studies, and educational or reference implementations.
Available as an open-weight research model
Open LLM research, scaling-law experiments, language-model evaluation, fine-tuning, and self-hosted English text generation
Open-weight and publicly downloadable; research-oriented base model
Open LLM research, local text generation, reproducible training experiments, and downstream fine-tuning
Open-weight, publicly downloadable; no official retirement date identified
Research, small-scale language-model experimentation, causal text generation, fine-tuning studies, and local deployment
Available open-weight model; legacy research release
Compact local language-model experiments, text completion, education, benchmarking, and fine-tuning research
Available as downloadable open weights; research-oriented and not instruction-tuned
Local text generation, language-model research, benchmarking, fine-tuning experiments, and studying compute-efficient scaling at small model sizes
Available open-weight research checkpoint
Multi-turn document retrieval and retrieval-augmented generation pipelines
Available as open-weight query and context encoder checkpoints
Document-grounded conversational question answering, local retrieval-augmented generation, and open-weight research deployments
Available as a public open-weight checkpoint; not identified as a current Cerebras-hosted API model
Low-latency assistants, customer support, high-volume text processing, coding assistance, image understanding, and parallel subagent workloads
Active (latest); retirement not sooner than October 15, 2026
Complex reasoning, long-running autonomous agents, advanced coding, multi-stage research, document-heavy analysis, and high-value enterprise workflows
Active; tentative retirement not sooner than 2027-06-09
Long-running agentic coding, demanding reasoning, multistep research, complex document analysis, spreadsheets, presentations, and high-stakes knowledge work
Active (latest)
Advanced cybersecurity, vulnerability research, biology, healthcare, life-sciences research, and long-running technical workflows requiring extensive context and reasoning
Active; invite-only; preview/beta on Amazon Bedrock
Vetted cybersecurity defense, advanced biology and life-sciences research, complex coding, long-running agents, technical investigation, and high-value research workflows
Active; invitation-only through Project Glasswing and trusted-access programs
Defensive cybersecurity research, vulnerability discovery, attack-surface auditing, autonomous coding, long-running agents, and large-context analysis.
Deprecated; deprecated June 9, 2026. Anthropic recommends migrating to Claude Mythos 5.
Complex reasoning, agentic coding, repository-scale software work, long-context analysis, research, document and spreadsheet workflows, and enterprise automation
Active legacy; retirement not sooner than 2027-02-05
Complex agentic coding, code review, enterprise analysis, long-context research, document workflows, financial and legal reasoning, and multi-step tool-using applications.
Active legacy; Anthropic recommends migration to Claude Opus 5.5. Retirement is scheduled no sooner than 2027-07-24.
Long-running agentic coding, repository-scale software engineering, code review, knowledge work, multimodal document analysis, and tool-using applications
Active (latest)
Complex software engineering, coding agents, multi-step research, computer-use workflows, enterprise analysis, visual document understanding, and high-value tool-using applications.
Active (legacy)
Complex reasoning, agentic coding, repository-scale software engineering, research, document analysis, computer-use workflows, and high-stakes knowledge work
Active legacy; migration to Claude Opus 5.5 recommended
Complex reasoning, agentic coding, long-horizon workflows, enterprise knowledge work, document analysis, vision tasks, and computer-use agents
Active legacy model; migration to newer Opus models is recommended for new deployments
Complex coding, software agents, computer-use workflows, visual document analysis, research, and long-running multi-step tasks
Legacy; still available as of 2026-09-24; migration to Claude Sonnet 5 recommended
Agentic coding, software engineering, browser and computer-use workflows, long-context analysis, tool-driven automation, and high-volume assistants
Current and generally available
Agentic coding, computer use, long-context document analysis, enterprise knowledge work, structured extraction, and high-volume applications needing strong reasoning at Sonnet-tier pricing
Active legacy; Anthropic recommends migration to Claude Sonnet 5
Fast coding agents, long-context analysis, image-aware workflows, document creation, and tool-enabled business applications
Active (latest)
Multilingual generation, translation, summarization, customer support, global communication, and multilingual research
Live
Multilingual image understanding, OCR, image captioning, visual question answering, image-based translation, visual reasoning and research deployments using open weights.
Live
Multilingual audio transcription, enterprise speech archives, meeting and interview transcription, call-center audio, and low-latency ASR workflows
Live; open-source research release
Arabic speech transcription, multilingual Arabic-English audio, regional dialects, code-switched speech, call-center audio and high-throughput ASR workloads.
Live
Enterprise RAG, long-context document analysis, multilingual applications, tool use, agentic workflows, financial text processing and structured text generation
Live
Enterprise agents, multimodal document and image analysis, multilingual workflows, reasoning-intensive automation, retrieval-augmented generation, and tool-using applications.
Live
Complex enterprise agents, tool use, retrieval-augmented generation, multilingual reasoning, long-context analysis, and workflow automation.
Live
High-quality multilingual text translation, enterprise document translation, and privacy-sensitive translation workflows.
Live
Enterprise document intelligence, OCR, chart and table analysis, visual question answering, and multilingual image understanding
Live
Low-cost enterprise RAG, multilingual document workflows, long-context chat, structured extraction, and tool-using agents
Live
Cost-sensitive RAG, enterprise chat, tool use, coding assistance, and fast multi-step agents
Live
Complex enterprise RAG, long-context document analysis, multilingual assistants, citations, structured data tasks and multi-step tool-use agents.
Live
Multilingual semantic search, multimodal RAG, PDF and document retrieval, image-to-text retrieval, classification, clustering, and enterprise vector indexing.
Generally available
Fast, storage-efficient English semantic search, retrieval, classification, clustering, and large-scale embedding workloads
Current and available
English semantic search, retrieval-augmented generation, vector indexing, classification, clustering, and similarity matching
Live
Fast multilingual semantic search, cross-lingual retrieval, RAG, clustering, classification features, and compact vector indexes
Active
Multilingual semantic search, cross-lingual retrieval, RAG indexing, classification, clustering, and image-text similarity
Active
Agentic software engineering, repository-level code changes, terminal-based coding agents, code review, local inference, and private deployment.
Live
Enterprise machine translation, multilingual documentation, localization, internal communications, safety procedures, and private or self-hosted translation workflows
Live
Compact multilingual image understanding, OCR, documents, charts, visual question answering, and specialized fine-tuning
Current open-weight release
High-volume enterprise document parsing, table and form extraction, search indexing, RAG ingestion, and document context for AI agents
Live; generally available
Multilingual semantic reranking, enterprise search, hybrid retrieval, and RAG pipelines
Active
English semantic reranking for enterprise search, hybrid retrieval, RAG pipelines, FAQs, knowledge bases, documents, code retrieval, and semi-structured records
Active and currently listed by Cohere; older Rerank 3.0 English model with newer Rerank alternatives available
Multilingual semantic reranking for enterprise search, hybrid retrieval, cross-language search, and RAG pipelines
Current older-generation model; superseded by newer Rerank model generations
Low-latency multilingual search reranking, high-throughput retrieval, enterprise search, hybrid search and RAG pipelines
Current and available
High-quality multilingual reranking for enterprise search, RAG pipelines, semantic retrieval, and semi-structured document ranking.
Generally available
Multilingual translation, conversation, summarization, and text generation focused on African and West Asian languages; local and edge deployment
Live
South Asian multilingual conversation, translation, target-language generation, local inference, and edge or on-device applications
Live
Multilingual translation, cross-lingual text generation, localized assistants, education, research, and efficient local or edge deployment
Live
Efficient multilingual translation, text generation, localization, language learning, and local or edge deployment for European and Asia-Pacific languages
Live; available through the Cohere Chat API and as an open-weight model
Self-hosted English text completion, code completion, model research, and fine-tuning experiments using an open-weight mixture-of-experts model.
Retired from Databricks Foundation Model APIs and Foundation Model Fine-tuning; open weights remain available for self-hosted use.
Self-hosted general English-language instruction following, coding assistance, text generation, and domain-specific fine-tuning
Open-weight model; Databricks Foundation Model API serving retired
Local experimentation, research, instruction-tuning studies, and organizations needing an openly downloadable model with commercial-use licensing
Legacy open-weight model; publicly downloadable; no official shutdown date found
Self-hosted general-purpose text generation, experimentation, and fine-tuning with open model weights.
Open weights; legacy model
Short-form instruction following, local inference, open-model research, and domain fine-tuning
Legacy open-weight model; downloadable and usable for self-hosted deployment
Self-hosted code generation, completion, repository analysis, debugging, code translation, and programming-focused research.
Legacy open-weight model series; downloadable checkpoints remain available, while the hosted Coder API line was superseded and merged into DeepSeek-V2.5.
Local code completion, fill-in-the-middle generation, IDE integrations, repository-level coding, and private or self-hosted inference
Open-weight and downloadable; accessible through the official Hugging Face repository, with no verified provider-published shutdown date
Local code generation, code completion, debugging, code explanation, repository-scale prompts, and developers needing an open-weight coding model
Available as an open-weight downloadable model; older but still accessible
Local bilingual text generation, research, experimentation, instruction tuning, and lightweight self-hosted assistants
Legacy open-weight model; downloadable and usable through compatible local inference tools
Local deployment, bilingual English-Chinese generation, language-model research, mathematics, coding experiments, and custom fine-tuning workflows
Legacy open-weight model; downloadable checkpoints remain available, but it is superseded by newer DeepSeek model families
Advanced mathematical reasoning, natural-language theorem proving, proof generation, proof verification, and research on self-correcting reasoning systems
Available as an open-weight research model; no official DeepSeek-hosted API deployment or public model-specific API pricing verified
Local OCR, document digitization, PDF and image parsing, layout-aware markdown conversion, table extraction, figure parsing, and visual-text compression research
Current open-weight model; publicly available for local and self-hosted inference
Local OCR, scanned-document transcription, layout-aware document parsing, table extraction, and document-to-Markdown workflows
Current open-weight model; self-hosted checkpoint
Lean 4 theorem proving, formal mathematics, automated proof generation, and theorem-proving research
Open-weight research release; legacy predecessor to DeepSeek-Prover-V1.5 and DeepSeek-Prover-V2
Lean 4 proof completion, formal mathematics research, automated theorem proving, and verifier-guided proof search
Open-weight, downloadable, legacy/superseded
Lean 4 theorem proving, formal mathematics, automated proof synthesis, proof-search research, and verifier-guided reasoning
Open-weight and downloadable; currently accessible from the official model repository; no first-party hosted API identified for this exact checkpoint
Lean 4 theorem proving, formal mathematics research, proof completion, proof-search experiments, and open-weight model fine-tuning
Legacy open-weight model; downloadable and usable, but superseded by newer DeepSeek-Prover releases
Mathematical reasoning, coding, technical analysis, research, and self-hosted reasoning applications
Open-weight checkpoint available; original hosted API identity superseded and scheduled for discontinuation
Self-hosted mathematics, coding, research, complex reasoning, and long-form analysis
Available open-weight model
Reasoning research, mathematics, coding experiments, reinforcement-learning studies, open-weight evaluation, and model distillation
Open-weight and downloadable; experimental/legacy research model; not a current first-party hosted API model
Local mathematical reasoning, compact reasoning experiments, educational applications, lightweight coding assistance, and self-hosted inference on limited hardware
Active open-weight model; downloadable for self-hosted inference
Local or third-party deployment for general text generation, translation, mathematics, research, and code generation
Legacy open-weight model; downloadable and usable through local or third-party inference, but not listed in DeepSeek's current first-party hosted API catalog
Local text generation, Chinese and English language tasks, MoE research, fine-tuning, and efficient self-hosted inference
Open-weight and downloadable; legacy self-hosting model
Open-weight general language generation, coding assistance, code completion, and self-hosted experimentation
Legacy open-weight model; the V2.5 series was superseded by newer DeepSeek model families, while the model weights remain available
Open-weight general language generation, coding, long-context text tasks, research, and cost-sensitive third-party inference.
Superseded and no longer current as a first-party hosted API model; open-weight checkpoint remains available for self-hosted and third-party deployment.
Open-weight reasoning, coding, tool-calling, long-context analysis, and self-hosted agent systems
Legacy hosted API generation; official open-weight release remains available
Open-weight deployment, coding assistance, long-context text processing, reasoning workflows, search agents, and terminal-oriented automation
Open-weight checkpoint available; dedicated DeepSeek API endpoint retired on 2025-10-15
Difficult mathematics, advanced coding, scientific reasoning, long-form analysis, benchmark evaluation, and research deployment
Retired hosted API; open-weight model remains available for self-hosting and third-party deployment
Local image understanding, visual question answering, diagram and document analysis, multimodal research, and compact deployments
Available open-weight checkpoint; legacy first-generation model
Open-weight reasoning, coding, long-context analysis, tool-using agents, research workflows, and cost-sensitive deployments
Legacy open-weight model; former DeepSeek API aliases deepseek-chat and deepseek-reasoner were scheduled for discontinuation on 2026-07-24
Long-context reasoning, coding, agent workflows, and cost-sensitive API applications
Retired; legacy API identifier temporarily routed to DeepSeek-V4.1-Flash from September 10, 2026
Self-hosted language-model research, custom post-training, domain adaptation, and large-context text generation
Current downloadable open-weight base checkpoint; no Hugging Face Inference Provider deployment listed
Image understanding, screenshot and chart analysis, multimodal coding agents, visual tool-use workflows, and text-plus-image reasoning
Retired as an independent model on 2026-09-10; legacy API identifier temporarily routes requests to DeepSeek-V4.1-Flash
Complex reasoning, coding agents, long-context analysis, tool-using workflows, and large document or codebase processing
Deprecated for independent serving; deepseek-v4-pro API requests are routed to DeepSeek-V4.1-Flash until V4.1-Pro launches
Research, continued pretraining, fine-tuning, foundation-model evaluation, and custom large-scale inference
Current open-weight base checkpoint; downloadable under the MIT License
Low-cost, high-throughput reasoning and coding, long-context analysis, agentic workflows, tool calling, and text-plus-image understanding.
Current and available through the DeepSeek API; the canonical API identifier is deepseek-flash.
Local research, image understanding, visual question answering, multimodal prototyping, and lightweight text-to-image experimentation
Available as an open-weight research model; superseded in the Janus series by newer Janus-Pro variants but still downloadable and usable.
Local research, visual question answering, image interpretation, and compact text-to-image experimentation
Available as an open-weight downloadable checkpoint; no official hosted inference API identified
Browser automation, visual UI interaction, repetitive web workflows, form filling, and user-interface testing
Legacy preview; currently documented and accessible, with newer Gemini 3.x models recommended for new computer-use applications
Large-scale multimodal processing, low-latency reasoning, coding, data extraction, tool-using agents, and applications requiring a very large context window.
Stable and currently served through the Gemini API with restricted access for users who have actively used Gemini 2.5 models; not deprecated; no shutdown date announced.
High-volume classification, simple extraction, lightweight multimodal analysis, routing, tagging, summarization, and extremely latency-sensitive applications.
Stable; currently served through the Gemini API with access limited to users who have actively used Gemini 2.5 models. Google recommends newer models for new projects.
Real-time voice and video agents, speech-to-speech assistants, interactive customer support, tutoring, coaching, and multimodal Live API applications
Preview; currently listed by Google with limited access for users who have actively used Gemini 2.5 models; no shutdown date announced
Low-latency controllable text-to-speech, voice assistants, narration, read-aloud features, and multi-speaker audio generation
Preview; currently available with limited access conditions
Advanced coding, complex reasoning, mathematics, STEM analysis, long documents, large codebases, multimodal analysis, and tool-using agents.
Stable and generally available through the Gemini API; access is currently limited to users who have actively used Gemini 2.5 models. Google states that the model is not deprecated and will continue to be served until further notice.
High-fidelity single-speaker and multi-speaker narration, audiobooks, podcasts, professional voiceovers, and scripted creative audio
Preview; currently listed and accessible through the Gemini API, with restricted access to Gemini 2.5 models
Fast image generation, conversational image editing, image transformation, and high-volume visual workflows
Deprecated; scheduled to shut down on 2026-10-02
Agentic workflows, everyday coding, reasoning and planning, multimodal analysis, long-context document work, and cost-sensitive tool-using applications
Preview
Complex reasoning, advanced software engineering, long-context multimodal analysis, and agentic workflows requiring reliable tool use
Preview
Fast multimodal applications, coding assistants, long-context document and video analysis, tool-using agents, enterprise workflows, and rapid agentic execution
Stable; previous-generation Flash model; currently accessible; no shutdown date announced
Coding, software engineering, tool-using agents, long-context multimodal analysis, structured extraction and high-throughput enterprise workflows
Generally available; previous-generation Flash model, currently supported
Long-horizon software engineering, autonomous agents, multimodal document workflows, enterprise knowledge work, and tool-using applications.
Generally available
High-volume translation, classification, extraction, summarization, document processing, and lightweight tool-using agent workflows
Deprecated; generally available and accessible until scheduled shutdown on May 7, 2027
Low-latency voice agents, real-time dialogue, multimodal live sessions, and interactive audio applications
Legacy preview; currently accessible; Google recommends Gemini 3.8 Live for most new low-latency voice-agent deployments
Controllable expressive speech, narration, accessibility, scripted audio, and multi-speaker TTS prototypes
Legacy preview; currently accessible; no shutdown date announced
Fast, low-cost 1K image generation and editing, rapid visual prototyping, interactive applications, high-volume image variations, storyboarding, and lightweight creative workflows
Generally available; scheduled retirement June 28, 2027 or later
Agentic workflows, coding agents, long-context multimodal analysis, tool use, and scaled production applications
Generally available; stable
High-volume, latency-sensitive agentic workflows, document parsing, translation, classification, data extraction, search-backed applications, and multimodal sub-agents.
General availability
Low-latency, real-time speech-to-speech translation for calls, meetings, travel, customer support, and multilingual voice applications
Preview
Pre-recorded audio transcription, multilingual speech recognition, speaker-labeled transcripts, timestamped transcripts, smart dictation, and domain-specific vocabulary
Generally available (GA)
High-volume text-to-speech production, low-latency voice-agent cascades, read-aloud applications, voice replication, and everyday single-speaker speech
Generally available
Studio-quality narration, audiobooks, expressive voice acting, complex multi-speaker dialogue, regional accents, difficult pronunciations, long-form narration, voice design, and voice replication.
Generally available (GA); no shutdown date announced
Low-latency voice agents, real-time audio-to-audio dialogue, multimodal assistants, and interactive tool-using applications
Stable; generally available
Complex real-time voice agents, multi-step problem solving, asynchronous tool workflows, technical support, travel coordination, and spoken STEM or coding tutoring
Stable; generally available
Autonomous market research, due diligence, literature reviews, competitive analysis, source-heavy investigations, and cited research reports
Preview
Comprehensive market research, competitive analysis, due diligence, literature reviews, and source-rich investigative reports
Preview; currently available through the Interactions API in the Gemini API and Google AI Studio
Text semantic search, RAG retrieval, document matching, classification, clustering, and recommendation systems
Current, scheduled for shutdown on 2028-05-14
Cross-modal semantic search, multimodal RAG, vector retrieval, recommendations, classification, clustering, and indexing mixed text and media collections.
Generally available
Fast, high-volume image generation and editing, visual iteration, marketing assets, diagrams, infographics, localization, and applications requiring image-search grounding
Generally available; the stable Gemini API model is gemini-3.1-flash-image
Professional image generation and editing, complex compositions, product mockups, infographics, branded creative, multilingual localization, and high-fidelity visual prototyping
Current stable model
Fast text-to-video, image-to-video, conversational video editing, video extension, interpolation, marketing content, and short-form cinematic production
Current; stable model available as gemini-omni-1.1-flash, with gemini-omni-flash-preview also documented
High-level robot planning, spatial reasoning, video progress tracking, tool orchestration, and multi-robot collaboration
Public preview
Low-latency robotic agents, continuous audio/video monitoring, function-based robot orchestration, warehouse workflows, and multi-robot coordination
Public preview
Full-length AI-generated songs, vocal music, instrumental arrangements, songwriting experiments, soundtracks, and image-inspired music creation
Generally available
Interactive instrumental music generation, live musical improvisation, prompt-driven DJ tools, MIDI-controlled experiences, and real-time creative audio applications
Experimental and currently documented through the Gemini API as lyria-realtime-exp
Generating short music or audio clips
Full-length AI music, soundtrack creation, songwriting, structured compositions, advertising audio, games, and creative production workflows
Public Preview
Cinematic text-to-video and image-to-video generation, short-form storytelling, storyboarding, advertising concepts, visual effects exploration, and creative previsualization with synchronized audio
Preview; currently accessible through Google APIs and Google products
Fast, high-volume video generation, creative iteration, social content, advertising concepts, and automated production workflows
Generally available on Vertex AI; preview model on the Gemini API
High-volume text-to-video and image-to-video generation, rapid creative iteration, social content, advertising variations, and cost-sensitive production workflows
Preview; currently available through the Gemini API and Google Cloud Vertex AI
Zero-shot forecasting of regularly sampled numerical time series across changing temporal resolutions
Current open-weight model; r1.1 revision available
Japanese text generation, summarization, classification, extraction, question answering, and Japanese-English translation in legacy IBM enterprise deployments
Deprecated; withdrawn from IBM Software Hub 5.2.0 and removed from watsonx.ai in 2.2.0
Multilingual enterprise question answering, retrieval-augmented generation, summarization, extraction, classification, and text generation in English, German, Spanish, French, and Portuguese
Retired; deprecated January 15, 2025 and withdrawn from standard watsonx.ai availability April 16, 2025
Fine-tuning, domain adaptation, text generation, summarization, extraction, classification, and question answering
Available for deployment on demand in IBM watsonx.ai; also available as an Apache 2.0 open-weight model through IBM's Hugging Face organization.
Fine-tuning, long-context document processing, enterprise text classification, extraction, summarization, question answering, and self-managed deployment
Available; base model intended for fine-tuning and dedicated deployment
Self-hosted enterprise assistants, long-context document analysis, RAG, summarization, extraction, multilingual text workflows, and function calling
Available as an open-weight model; superseded by Granite-3.3-8B-Instruct for newer deployments
Long-context enterprise RAG, document and meeting summarization, information extraction, multilingual dialogue, code-related tasks, function calling, and self-hosted deployments
Available; legacy in some IBM watsonx catalogs
Local or dedicated enterprise assistants, RAG, long-document summarization, multilingual text workflows, coding assistance, reasoning, and resource-conscious inference
Available; open-weight model with local deployment and dedicated IBM watsonx deployment options
Self-hosted enterprise assistants, long-context RAG, summarization, multilingual text tasks, coding assistance, function calling, and cost-sensitive deployments.
Retired from IBM watsonx.ai deploy-on-demand service on 2026-02-22; open-weight checkpoint remains available for self-hosted and compatible third-party deployments.
Low-latency local inference, long-context text generation, RAG, multilingual assistants, coding assistance, and lightweight tool-calling agents
Current open-weight instruct model
Enterprise RAG, multi-tool agents, function calling, customer-support automation, multilingual instruction following, and long-context workloads
Available; open-weight instruct model
Efficient enterprise assistants, multilingual text generation, RAG, classification, extraction, summarization, coding assistance, fill-in-the-middle completion, structured JSON, and tool-calling workflows
Current; open-weight instruct model; available for download and listed for deploy-on-demand use in IBM watsonx.ai
Efficient local or private deployment, multilingual enterprise text processing, RAG, summarization, extraction, coding assistance, function calling, and lightweight AI assistants.
Currently available open-weight model; earlier-generation Granite model superseded by the newer Granite 4.2 family, with no verified deprecation or shutdown date.
Self-hosted enterprise assistants, multilingual text generation, RAG, coding assistance, structured extraction, and tool-calling agents
Current open-weight instruction model
Self-hosted enterprise assistants, long-context RAG, multilingual applications, coding, structured extraction, and tool-calling agents
Current open-weight instruct model; publicly available
Efficient reasoning, coding assistance, tool calling, multilingual dialogue, local deployment, and lightweight enterprise agents
Current; publicly available open-weight model
Local or self-hosted reasoning, coding assistants, tool calling, multilingual dialogue, retrieval-augmented generation, and agentic workflows
Current; open-weight and downloadable
Self-hosted enterprise reasoning, coding agents, multilingual applications, tool calling, long-context workflows, and organizations requiring Apache 2.0 licensing
Current; open-weight and available for download
Self-hosted English text generation, summarization, extraction, classification, experimentation with LAB-aligned open-weight models, and resource-conscious deployments
Withdrawn from IBM watsonx.ai on 2025-01-07; open-weight checkpoint remains available for self-hosted deployment
English enterprise chat, retrieval-augmented generation, question answering, summarization, extraction, and classification
Withdrawn; access ended January 19, 2025
Self-hosted coding assistants, code generation, code explanation, code conversion, repository-scale prompts, and historical or reproducible research
Retired; withdrawn from IBM watsonx.ai on 2025-07-17
Local code generation, explanation, repair, translation, and research with an open Apache 2.0 model
Deprecated; still publicly available for historical and scientific use
Schema linking and relevant-table or relevant-column selection in text-to-SQL pipelines
Available; deploy on demand only
Natural-language-to-SQL generation over structured databases, analytics assistants, and the SQL-generation stage of text-to-SQL pipelines.
Available; deploy-on-demand model in IBM watsonx.ai
Self-hosted coding assistants, code generation, code conversion, code explanation, and programming experiments where an Apache 2.0 open-weight model is preferred.
Available as an open-weight model and documented for IBM watsonx.ai deploy-on-demand use; deprecated and withdrawn from the watsonx.ai multitenant offering according to IBM's 2025 lifecycle notice.
Self-hosted code generation, code explanation, code repair, code conversion, coding assistants, and legacy Granite Code compatibility
Legacy/deprecated; available in some IBM watsonx environments and retained for historical or scientific use
Fast, low-footprint English semantic search, RAG retrieval, similarity matching, and vector indexing
Legacy; downloadable and usable, but superseded by Granite Embedding Small English R2
Low-latency multilingual semantic search, retrieval-augmented generation, vector search, document similarity, long-document retrieval, and cross-lingual code retrieval
Current; open-weight model
Low-cost multilingual semantic search, cross-lingual retrieval, vector databases, similarity matching, and RAG pipelines
Retired from IBM watsonx.ai; open-weight checkpoint remains available through IBM's Hugging Face organization
English semantic search, vector retrieval, RAG, similarity matching, enterprise knowledge-base search, and local embedding deployment
Legacy/superseded but downloadable and usable
Multilingual semantic search, retrieval-augmented generation, cross-lingual retrieval, long-document search, similarity, and code retrieval
Current open-weight model
English semantic search, vector retrieval, retrieval-augmented generation, document similarity, clustering, and enterprise information retrieval
Current open-weight model
Compact English semantic search, retrieval-augmented generation, document similarity, and private vector-search deployments
Current open-weight model
Multilingual semantic search, vector retrieval, RAG, clustering, similarity matching, and text classification features
Retired; shutdown completed 2026-08-08
Above-ground biomass mapping, forest monitoring, carbon-stock estimation, ecological analysis, and satellite-based remote-sensing research
Available open-weight model
Canopy-height mapping, forest monitoring, vegetation analysis, carbon-cycle research, ecological assessment, and remote-sensing experimentation.
Available as downloadable open-weight model; not deployed by an inference provider; IBM repository disclosure states that the project is not maintained as an IBM product.
Land surface temperature estimation, urban heat island analysis, satellite-based environmental monitoring, and temporal gap filling
Available as an open-weight model for local inference
Remote-sensing feature extraction, Earth-observation transfer learning, and regional flood-segmentation workflows using multispectral and SAR satellite imagery
Available as an open-weight research model; no hosted inference provider is currently listed
Spatial downscaling of weather forecasts, reanalysis data, and climate simulations
Current and available as downloadable open-weight model
Prompt and response safety classification, jailbreak detection, RAG groundedness and relevance checks, hallucination evaluation, and enterprise AI guardrails
Deprecated in IBM watsonx.ai documentation; open-weight model and official model materials remain available
AI safety guardrails, jailbreak detection, RAG groundedness and relevance checks, function-call hallucination detection, custom criteria evaluation, and best-of-N response ranking
Current and available as an open-weight model
Multilingual speech-to-text with speaker labels, word-level timestamps, keyword biasing, and self-hosted enterprise transcription
Current open-weight model
Multilingual offline transcription, speech-to-text, speech translation, subtitle generation, and domain-specific recognition with keyword biasing
Current open-weight model
High-throughput multilingual speech transcription where low inference latency is more important than maximum recognition accuracy.
Current; open-weight; Apache 2.0
High-throughput, low-latency English speech-to-text transcription on local, edge, and enterprise systems
Current; open-weight model
Fast local English speech-to-text transcription, edge-device ASR, research, browser demos, and high-throughput noncommercial batch processing
Current; research and noncommercial use only
Lightweight multivariate forecasting at minute- and hour-level resolutions, especially when 512 historical observations and a 96-point forecast horizon are appropriate.
Available
Zero-shot forecasting of demand, prices, energy loads, traffic, telemetry, and other regularly sampled numerical time series; probabilistic forecasts and uncertainty intervals; local or self-managed inference.
Current; open-weight model
High-throughput multivariate time-series forecasting, zero-shot or few-shot forecasting, probabilistic prediction, exogenous-variable forecasting, and CPU-friendly production deployment.
Current open-weight model family; available through the IBM Granite Hugging Face repository. No hosted inference provider deployment was listed on the reviewed model page.
Efficient multivariate forecasting for regularly sampled energy, traffic, manufacturing, network, sales, and sensor data
Available through IBM watsonx.ai; downloadable model branch
Multivariate forecasting when at least 1,536 historical observations per channel are available, especially demand, traffic, electricity, manufacturing, finance, and other minute- or hour-level forecasting tasks.
Available
Chart extraction, table parsing, semantic key-value extraction, visual document processing, and local enterprise RAG pipelines
Current and downloadable; newer Granite Vision 4.1 4B version available
Enterprise document understanding, chart and table extraction, OCR-oriented image analysis, visual question answering, and multimodal RAG
Available; legacy relative to Granite Vision 4.0 3B Vision
Structured extraction from charts, tables, invoices, forms, and enterprise document images
Current open-weight model
English semantic search, dense retrieval, vector indexing, duplicate-question matching, and lightweight retrieval-augmented generation
Deprecated; withdrawn in Dallas on 2026-09-08 and scheduled for withdrawal in other listed regions on 2027-01-12.
English semantic search, vector database indexing, retrieval-augmented generation, document matching, and query-passage retrieval
Deprecated; scheduled for withdrawal on 2027-01-12
Bilingual English-Korean chat, instruction following, Korean-language applications, local inference, and open-weight LLM research
Available open-weight model; older EXAONE generation and superseded by newer EXAONE releases
Lightweight local text generation, English-Korean assistants, summarization, rewriting, classification, and deployment on resource-constrained devices
Current downloadable open-weight model; research use permitted, commercial use requires a separate license
Bilingual English-Korean assistants, long-document processing, local deployment, coding support, summarization, translation, and research
Available as an open-weight research model
Local or self-hosted English-Korean text generation, instruction following, long-context processing, coding, mathematics, and organizations requiring downloadable model weights.
Available open-weight model; older EXAONE generation
On-device assistants, local text generation, Korean and multilingual conversational applications, lightweight reasoning, and resource-constrained deployments
Current open-weight model
Self-hosted Korean, English, and Spanish language tasks requiring a combination of general-purpose generation, reasoning, coding, long-context processing, or agentic tool use
Available open-weight model; no official deprecation or shutdown date found
Open-weight multimodal reasoning, Korean-language tasks, document understanding, OCR, visual question answering, and self-hosted deployments
Current open-weight model
Local mathematics, coding, reasoning experiments, compact bilingual English-Korean applications, and resource-conscious deployment
Current and publicly downloadable open-weight model
Local research, mathematical reasoning, science problem solving, coding evaluation, Korean and English text generation, and self-hosted experimentation
Current and accessible open-weight research release
Mathematical reasoning, scientific problem solving, coding evaluation, and local research deployments
Current open-weight research model; downloadable and locally deployable
Zero-shot probabilistic forecasting of financial and other real-valued time series, especially research and educational workflows involving uncertainty estimates.
Current open-weight research release
Whole-slide digital pathology research, molecular subtyping, and image-based mutation prediction
Available research release; a later EXAONE Path 2.0 model has also been released
Computational pathology research, whole-slide image representation learning, cancer biomarker prediction, and biomedical image analysis
Current open-source research model
Research prediction of EGFR mutation status from lung adenocarcinoma whole-slide pathology images
Current research release; gated Hugging Face access; locally deployable
Research on EGFR mutation prediction, computational pathology, molecular subtyping, and whole-slide image biomarker analysis
Available as a gated open-source research release; not deployed by a hosted inference provider
In-context classification and regression on structured tabular datasets, especially low-retraining workflows and field data with missing values
Current; released for research and education
Long-context reasoning, multilingual generation, Korean-language workloads, coding, tool-using agents, research and distributed self-hosted inference
Current open-weight model
Self-hosted multilingual reasoning, Korean-language applications, coding, long-context document analysis, and tool-enabled agents
Current open-weight model
Image embeddings, dense feature extraction, image retrieval, classification, segmentation, depth estimation, object discovery, video tracking pipelines, and geospatial computer vision
Current; downloadable open-weight research model suite
Local inference, fine-tuning, private deployment, text generation, retrieval-augmented generation, research, and cost-sensitive applications
Current open-weight static model; downloadable and usable through compatible local or hosted inference deployments
High-quality open-weight research, multilingual applications, coding, reasoning, synthetic-data generation, model distillation and self-hosted deployments with substantial infrastructure
Available as an open-weight model; older generation superseded by newer Llama releases
Private local inference, mobile and edge assistants, summarization, rewriting, retrieval-supported generation, and lightweight multilingual applications
Available downloadable open-weight model; static checkpoint
Private local inference, multilingual text generation, edge applications, model adaptation, and fine-tuning
Available; static pretrained open-weight model
Visual question answering, image reasoning, chart and document understanding, image captioning, multimodal research, and self-hosted or partner-hosted AI applications.
Available open-weight model; static model trained on an offline dataset
Self-hosted visual question answering, image captioning, document analysis, visual reasoning, and multimodal assistants
Available open-weight static model
Self-hosted or hosted multilingual chat, coding assistance, long-context text generation, tool calling, synthetic data, and applications requiring open model weights
Available open-weight model; static model trained on an offline dataset
Open-weight multimodal assistants, image understanding, visual question answering, coding, multilingual applications, creative writing and long-context text processing.
Available; open-weight static checkpoint
Long-context document and code analysis, visual question answering, multimodal assistants, multilingual applications, self-hosted inference, and customized deployments
Available; open-weight model
Text and image moderation for prompts and generated responses in generative AI systems
Current and available as an open-weight model
Low-cost prompt and response safety classification, local moderation, mobile and edge deployments, and customizable LLM guardrails
Current open-weight model; downloadable subject to access approval
Self-hosted multilingual input and output moderation for LLM applications
Available open-weight safety classifier
Safety classification of mixed text-and-image prompts and text responses in multimodal LLM systems
Available open-weight multimodal safety model; Meta's current model repositories continue to list it.
Low-latency detection of prompt injections and jailbreak attempts in LLM applications, agents, retrieved documents, and other untrusted text
Current open-weight safety classifier
Multilingual prompt-injection detection, jailbreak screening, agent security, and filtering untrusted text before it reaches an LLM
Current open-weight model; gated download access
Local agents, long-running tool workflows, coding assistants, multimodal document and screenshot understanding, private on-device inference, and model customization
Current open-weight model; self-hosted and available through selected third-party hosted inference providers
Text-to-image generation, precise image editing, multi-image composition, anchored visual series, product imagery, creative assets, and grounded visual content.
Current and available through Meta Model API; also available in selected Meta AI consumer experiences
Agentic workflows, coding agents, computer-use automation, multimodal document and media analysis, long-context reasoning, tool orchestration, and web-grounded applications.
Current but superseded by Muse Spark 1.2 and Muse Spark 1.3; available on the Meta Model API Standard tier in public preview for US developers.
Long-horizon coding agents, repository-scale software engineering, multimodal code generation, debugging, refactoring, and tool-driven workflows
Available; previous version, with Muse Spark 1.3 recommended for new work
Long-horizon coding agents, software engineering, browser and computer-use workflows, tool orchestration, large repositories, document analysis, and multimodal reasoning
Current; available through Meta Model API and Muse Code
Real-time speech-to-text, live captions, voice agents, meeting transcription, call intelligence, dictation, and speaker-aware transcription
Current and available through Meta Model API
Multilingual speech representation learning, audio embeddings, low-resource language research, and custom downstream speech systems
Current; open-source research model family
Cross-modal audio-video-text retrieval, audiovisual embeddings, sound-event understanding, media indexing, and multimodal perception systems.
Current and openly available
Single-image reconstruction of textured 3D objects from natural scenes, 3D computer-vision research, Gaussian-splat workflows, and rapid asset prototyping
Available research release; gated model checkpoints
Prompted audio separation, speech and noise isolation, instrument and vocal extraction, audiovisual sound segmentation, and audio-editing research
Current open research release; downloadable checkpoints and public demo available
Reference-free evaluation and benchmarking of text-guided audio-separation outputs
Current gated research release
Multilingual automatic speech recognition, speech-to-text translation, text translation, text-to-speech translation, and speech-to-speech translation
Available open-weight research model; noncommercial research use
Open-vocabulary object detection, pixel-level image segmentation, and multi-object video tracking
Current and available; hosted through Meta Model API and available as released research checkpoints
Computational neuroscience, fMRI response prediction, brain encoding, multisensory research, and in-silico experiment design
Open-weight research release
Text-guided biomedical image segmentation, annotation assistance, organ and tumor delineation, pathology-cell analysis, and research-oriented medical imaging pipelines
Current; open-weight research model and available for managed deployment through Microsoft Foundry
First-pass chest X-ray report drafting, structured findings extraction, radiologist workflow assistance, research evaluation, and institution-specific fine-tuning
Limited preview; registration and eligibility approval required
Local image captioning, object detection, phrase grounding, region description, OCR and lightweight computer-vision pipelines
Current open-weight model; downloadable from Hugging Face and usable for local or self-hosted inference
Local image captioning, object detection, visual grounding, OCR, region annotation and multi-task computer-vision pipelines.
Current open-weight model; publicly accessible through Hugging Face for local or self-hosted inference.
Open-weight reasoning, mathematics, coding, research, and general text-generation applications requiring DeepSeek-R1-style reasoning with Microsoft post-training
Available as an open-weights model and through Microsoft Foundry hosted API
High-quality text-to-image generation, photorealistic imagery, product and marketing visuals, presentation graphics, and precise image-to-image editing.
Preview
High-quality text-to-image generation, controlled image editing, commercial imagery, product and branding visuals, photorealistic scenes, and multi-reference creative workflows
Public Preview
Fast, cost-conscious text-to-image generation, image editing, creative production, concept visualization, and high-volume image workflows
Public preview; scheduled for retirement on 2026-10-01
High-fidelity text-to-image generation, precise image editing, hero imagery, commercial and photorealistic creative work, accurate in-image typography, and visually dense scenes requiring consistent objects, characters, materials, and spatial relationship
Public preview
Fast, high-volume text-to-image generation, image editing, marketing assets, product imagery and production design
Public Preview
Complex mathematical reasoning, software engineering, quantitative enterprise analysis, long-context document work, and application-managed agent workflows.
Public preview
Multilingual speech-to-text, captions, meeting transcription, accessibility, call analysis, content workflows, voice-agent audio understanding, and domain-specific terminology
Preview; currently accessible through Microsoft Foundry and Azure Speech
Multilingual audio transcription, meeting and contact-center records, captions, clinical notes, accessibility, media search, voice-agent evaluation, and domain-specific transcription with speaker labels and timestamps
Public preview
Expressive long-form narration, audiobooks, podcasts, educational content, voice-over, accessibility, and high-fidelity branded audio
Public preview
Low-latency expressive speech for voice agents, assistants, call centers, IVR systems, and interactive multilingual applications
Public preview
Medical image and text embeddings, similarity search, multimodal retrieval, downstream classification, outlier detection, dataset curation, and healthcare AI development workflows
Limited preview
Text-prompted segmentation of complete CT or MRI volumes, organ and lesion delineation, volumetry, annotation assistance, and medical-imaging research
Available through Microsoft Foundry classic managed compute; research and model-development use only
Enterprise applications with mixed-complexity workloads, model selection automation, agent workflows, cost optimization, and configurable quality-versus-latency trade-offs.
Current; latest documented version is 2025-11-18
Local and private text generation, coding assistance, mathematics, reasoning, summarization, and latency-sensitive applications
Retired from Microsoft Foundry; open-weight repository remains available for independent deployment
Long-context chat, retrieval-augmented generation, document summarization, coding, mathematics, reasoning, and self-hosted text-generation applications
Generally available in Microsoft Foundry; open-weight checkpoint available for download and self-hosting
Local or low-latency text generation, lightweight assistants, mathematics, coding, summarization, and resource-constrained deployments
Retired from Microsoft Foundry on August 30, 2025
Long-document analysis, local assistants, code and math tasks, retrieval-augmented generation, private or offline inference, and resource-constrained deployments.
Retired from Microsoft Foundry on August 30, 2025; open-weight checkpoint remains downloadable and usable for self-hosted inference.
Local and private text generation, instruction following, code and mathematics assistance, summarization, extraction, and resource-constrained deployments
Retired from Microsoft Foundry on 2025-08-30; downloadable open-weight checkpoint remains available
Long-context text generation, document analysis, summarization, local assistants, coding support, mathematics, and cost-sensitive self-hosted applications
Retired from Microsoft Foundry on 2025-08-30; open-weight checkpoint remains available for self-hosted deployment
Local multilingual chat, summarization, document analysis, coding assistance, long-context retrieval, and resource-constrained deployments
Available open-weight model; also supported in Microsoft and third-party local or hosted inference environments
Long-context multilingual assistants, coding, mathematics, reasoning, retrieval-augmented generation, and controlled local deployment.
Retired from Azure Foundry on August 30, 2025; downloadable open-weight checkpoint remains available for self-hosted deployment.
Lightweight image and text reasoning, OCR, charts, tables, diagrams, documents, screenshots, and multi-image comparison
Current and accessible; open-weight model and available through Microsoft Foundry deployment options
Efficient local or hosted text generation, STEM reasoning, mathematics, coding assistance, technical question answering, and research on small language models
Preview in Microsoft Foundry; open-weight model available under the MIT license
Efficient mathematical reasoning, math tutoring, automated assessment, lightweight reasoning agents, edge deployment, mobile applications, and latency-sensitive local inference
Current open-weight model; available through Microsoft’s Hugging Face repository and documented Azure AI Foundry availability
Efficient local or cloud text generation, multilingual applications, mathematics, coding, reasoning, retrieval-augmented generation, and edge deployment
Generally available
Mathematical reasoning, educational tutoring, formal proof assistance, symbolic computation, local inference, edge deployment, and latency-sensitive applications
Preview; currently available through Microsoft Foundry and open-weight distribution
Compact multimodal assistants, OCR, document and chart analysis, image question answering, speech recognition, speech translation, audio summarization, and private or local deployment
Available; open-weight; listed in Microsoft Foundry and Hugging Face
Mathematical reasoning, coding, science, logic, algorithmic problem solving, research, and resource-constrained local deployments
Preview; open-weight and currently listed in Microsoft Foundry
Mathematical reasoning, scientific problem solving, coding assistance, algorithmic tasks, local deployment, and reasoning-model research
Current downloadable open-weight model
Short text-to-video and image-to-video clips, cinematic experiments, advertising concepts, social content, and scenes with complex motion.
Legacy or superseded; no longer prominent in MiniMax's current first-party model catalog
Short cinematic videos, image animation, realistic human motion, stylized scenes, visual effects, advertising concepts, and social-media content.
Legacy or superseded video model; still referenced in MiniMax consumer and platform offerings, while MiniMax H3 is the current primary video model in developer documentation.
Fast image-to-video generation, high-volume short-form content, social media clips, advertisements, and rapid creative iteration
Legacy; current availability should be verified
Short text-driven video concepts, storyboards, and early Hailuo-style cinematic experiments
Legacy and superseded; no longer listed in MiniMax's current primary video-generation catalog
Prompt-based image generation, reference-guided variations, commercial visuals, and batch creative production
Current and accessible through MiniMax's image-generation console and API platform
Multilingual audio transcription, meeting transcription, speaker-labeled transcripts, live captions, call analysis, and subtitle generation
Current public model
Fast short-form text-to-video, image-to-video, reference-guided generation, and synchronized-audio production
Current; commercially hosted by fal
Multimodal commercial video generation, reference-based editing, product and advertising content, short cinematic clips, and locally deployed 768p workflows
Current; open-weight release and hosted API available
Coding agents, multi-step tool workflows, long-context codebase analysis, research automation, and self-hosted experimentation
Legacy/open-weight model; current hosted availability and pricing are not confirmed in MiniMax's latest public model catalog
Multilingual software engineering, coding agents, tool-using workflows, application development, and cost-sensitive automation
Legacy but currently available
Coding agents, software engineering, search and browser agents, tool-calling workflows, office automation, and long-context technical work
Current and accessible; open-weight model; also available through the MiniMax API and MiniMax Agent
Agentic software engineering, repository-level coding, production debugging, complex tool workflows, office document automation, and long-context professional tasks
Current; available through MiniMax API, MiniMax Agent, and downloadable open weights
Low-latency coding assistants, multilingual software development, tool-using agents, long-horizon workflows, and interactive office automation
Current and available through the MiniMax Open Platform API; historical model variant
Low-latency coding assistants, software-engineering agents, tool-using workflows, search tasks, and long-context productivity automation
Current and available; highspeed variant of MiniMax M2.5
Low-latency coding assistants, software-engineering agents, tool-calling workflows, and interactive developer applications
Current and available through the MiniMax API Platform; highspeed variant of MiniMax M2.7
Long-context coding agents, autonomous tool-using workflows, multimodal document and video analysis, and private deployment
Active; open-weight and available through the MiniMax API
Generating complete songs with expressive vocals, lyrics, melodies, instrumental arrangements, duets, a cappella passages, and cinematic musical soundscapes.
Legacy or limited availability; MiniMax states that music models were no longer available through Token Plan from 2026-08-20, but a full model retirement date was not verified.
Text-to-music, instrumental generation, game and video scoring, detailed musical direction, and genre reinterpretation with Cover mode
Legacy or superseded; current public MiniMax navigation highlights Music 3.0, and Music 2.6 availability should be verified before use
Local generation of complete songs from lyrics and structured musical descriptions
Open-weight model; MiniMax Token Plan access discontinued on 2026-08-20
Reinterpreting existing songs in new genres, vocal styles, arrangements, and production directions while preserving the source melody
Current; specialized music-cover model introduced with MiniMax Music 2.6
High-quality multilingual voiceovers, audiobooks, narration, digital characters, advertising, education, and zero-shot voice cloning
Legacy or older generation; still accessible through some partner platforms, but not listed as a core model on MiniMax's current global pricing page
Low-latency multilingual text-to-speech, streaming voice agents, interactive applications, expressive narration, and voice cloning
Legacy or superseded; exact current first-party availability is unclear
Real-time voice agents, conversational assistants, customer-service automation, interactive characters, multilingual speech, and low-latency text-to-speech
Legacy or superseded; current first-party availability is unverified
High-quality voiceovers, audiobooks, narration, localization, e-learning, game dialogue, accessibility audio, and production speech
Legacy or transition-era model; not prominently listed in MiniMax's current first-party speech catalog as of September 25, 2026
High-quality expressive narration, audiobooks, podcasts, advertising, character voices, multilingual speech, and applications prioritizing audio fidelity over the lowest latency
Current and available through the MiniMax API
Real-time text-to-speech, voice assistants, conversational agents, interactive applications, multilingual narration, gaming characters and expressive voice experiences
Current and accessible through the MiniMax Open Platform API
Text-to-video generation with explicit cinematic camera-movement direction, short advertising concepts, storyboards, and controlled visual experiments
Legacy or limited availability; current official catalog status and continued first-party access are not clearly verified
Low-latency IDE autocomplete, fill-in-the-middle completion, code generation, code editing, test generation and developer assistants
Active
Semantic code search, repository retrieval, coding-agent RAG, code similarity, duplicate detection, clustering, and code analytics
Active
Lean 4 theorem proving, formal verification, autoformalization, proof debugging, and agentic proof engineering
Public Preview; scheduled for retirement on 2026-09-30
Low-cost edge and local inference, image-aware assistants, document analysis, structured extraction, lightweight agents, task routing, and privacy-sensitive deployments.
Active; generally available
Efficient edge and local inference, image understanding, document workflows, structured extraction, lightweight agents, and high-volume text generation.
Active; generally available
Private assistants, local vision-language applications, multilingual workloads, document and image analysis, and cost-efficient agentic systems
Active; generally available
Semantic search, retrieval-augmented generation, vector databases, document classification, clustering, duplicate detection and general text retrieval
Generally available
Long-context enterprise assistants, multilingual applications, image-aware document analysis, agentic workflows, coding, RAG and self-hosted sovereign deployments
Active; generally available; open-weight
Agentic coding, software engineering, long-context analysis, multimodal document workflows, structured outputs and multi-step tool use
GA; currently available; open weights
Text moderation, conversational safety classification, content filtering, policy enforcement, guardrails, and jailbreaking detection
Active; generally available; Premier
High-volume OCR, structured document extraction, enterprise search, RAG ingestion, invoice processing, compliance workflows, and document automation
Generally available; superseded by OCR 4.1 as Mistral's latest OCR model
Cost-efficient general chat, multimodal document analysis, coding, agentic workflows, and configurable reasoning
Active; generally available
High-volume document extraction, scanned forms, handwriting, invoices, complex tables, archival digitization, and document-to-knowledge pipelines.
Legacy; available for existing integrations and production workloads. OCR 4 is the newer model.
OCR, document parsing, structured extraction, enterprise search, RAG ingestion, invoice processing, and document AI workflows
Generally Available
Policy-based moderation of prompts, model responses, refusals, text, images, and text-image combinations
Public Preview; open weights; self-hosted
Batch transcription of meetings, interviews, calls, subtitles, compliance recordings, and searchable audio
Active; GA; Premier hosted model
Live speech transcription, realtime captions, voice interfaces, realtime note-taking and speech-to-speech pipelines
GA
Production-scale audio understanding, multilingual transcription, audio Q&A, meeting and call summarization, speech translation, and voice-driven function calling
Active; generally available
Multilingual voice generation, expressive voice agents, zero-shot voice cloning, custom voice adaptation, and low-latency speech output
GA; currently available through the Mistral API and Mistral Studio
Fine-tuning and research on speech recognition, audio understanding, audio classification, audio question answering, and speech-audio generation
Current open-weight base model; not instruction-tuned
Self-hosted speech recognition, audio understanding, audio question answering, audio captioning, and spoken conversational agents
Available as an open-weight self-hosted model
Repository-level software issue resolution, code repair, test generation, coding agents, and self-hosted software engineering workflows
Available open-weight model
Self-hosted image and video understanding, OCR, long-document analysis, visual question answering, screenshot perception, and efficient multimodal applications
Current open-weight/downloadable model; a newer Kimi-VL-A3B-Thinking-2506 variant is recommended for stronger multimodal reasoning
Local multimodal reasoning, mathematical visual question answering, OCR, document understanding, image and video analysis, and research on open-weight vision-language models
Available as downloadable open weights; superseded by Kimi-VL-A3B-Thinking-2506
Open-weight image and video reasoning, OCR, chart interpretation, visual mathematics, long PDFs, high-resolution screenshots and GUI-agent grounding
Current open-weight model; publicly available for self-hosted and third-party inference
Fine-tuning, foundation-model research, custom language systems, coding experiments, and self-hosted inference
Open-weight and downloadable; legacy relative to newer Kimi K2.x models but still accessible from the official Hugging Face repository
Open-weight coding assistants, tool-using agents, general-purpose chat, and self-hosted research deployments
Open-weight checkpoint available; legacy relative to Moonshot AI's current hosted API catalog
Self-hosted reasoning agents, autonomous research, long-horizon tool workflows, coding, and complex multi-step analysis
Retired from Moonshot's direct API; open-weight checkpoint remains available for self-hosted or third-party deployment
Long-horizon software engineering, agentic coding, visual document understanding, tool-using workflows, and multi-agent orchestration
Current; open-source model, available through Kimi, Kimi API, Kimi Code, and downloadable model weights
Long-horizon software engineering, repository-level coding, multi-file refactoring, debugging, coding agents, and tool-driven development workflows
Current and accessible through Kimi API; Kimi Code default service has moved to Kimi K2.8 Preview, while Kimi K2.7 Code remains available through API and Kimi K2.7 Code HighSpeed service.
Long-context software development, repository analysis, code completion, multi-file refactoring, and agentic coding workflows
Preview; fully rolled out in Kimi Code
Multimodal coding, visual debugging, long-context analysis, tool-using agents, and complex research or office workflows
current
Long-context coding, software engineering, multimodal document and video understanding, agentic workflows, technical research, and complex reasoning
Current and available; open-weight model
Long-context text generation, local inference, continued pretraining, research, and fine-tuning
Current open-weight model; publicly downloadable
Long-context text generation, self-hosted assistants, document analysis, research, and efficient local inference
Current open-weight instruction-tuned checkpoint
Self-hosted text generation, language-model research, efficient MoE inference, code and mathematics experimentation, and fine-tuning research
Available open-weight pretrained checkpoint
Self-hosted instruction following, general text generation, research, experimentation, and cost-conscious local inference
Available open-weight checkpoint
Native-resolution image feature extraction, vision-language model backbones, high-resolution document and image understanding, and multimodal research
Current open-weight model; available for local use through Hugging Face Transformers; not deployed by a Hugging Face Inference Provider
General-purpose semantic search, retrieval, document similarity, clustering, and text classification
Current and accessible through CLOVA Studio's Embedding API
Sentence similarity, semantic search, document relatedness, clustering, and text classification features
Current and available through the CLOVA Studio Embedding API
Fast, cost-sensitive Korean text generation, classification, summarization, report drafting, data expansion and customized enterprise chatbots
Current and available through CLOVA Studio and related NAVER Cloud services
Korean-focused text generation, summarization, extraction, classification, tuning, and batch text workflows
Current and available through CLOVA Studio
Korean-language business applications, image understanding, document and visual analysis, instruction following, and API-based assistants.
current
Complex reasoning, mathematics, science, language reasoning, writing, long-context text generation, and Korean-language enterprise applications
Current and accessible through CLOVA Studio
Fast, high-throughput text generation, classification, summarization, simple extraction, and function-calling workflows
Available
Visual encoding for HyperCLOVA X SEED 4B and edge-oriented multimodal applications
Current; proprietary vision encoder component for HyperCLOVA X SEED 4B
Lightweight Korean conversational interfaces, mobile and edge applications, smart-home devices, wearables, and customer-support chatbots
Current open-weight model; available for download
Korean-language applications, lightweight local inference, basic translation, education, business communication, specialized chatbots, and domain fine-tuning.
Current downloadable open-weight model; Hugging Face repository is gated and requires acceptance of access conditions.
Korean image and video understanding, visual question answering, chart and diagram interpretation, OCR-assisted analysis, tourism and cultural applications, and locally deployable fine-tuned systems.
Available open-weight model
Korean visual and document understanding, video-and-audio analysis, edge AI, public-sector systems, defense intelligence, and air-gapped deployments
Publicly announced; current model identity, with detailed commercial/API availability not publicly specified
Korean-first any-to-any multimodal assistants, speech and vision applications, multimodal research, and self-hosted deployments
Current open-weight model
Korean-language reasoning, mathematics, coding, instruction following, tool-connected agents, and self-hosted commercial applications
Current open-weight model; free for commercial use under the HyperCLOVA X SEED license
Korean-language reasoning, visual question answering, long-context multimodal analysis, document and chart understanding, coding assistance, and tool-using AI agents
Current; open-weight/open-source release
Real-time or batch speaker identification and active-speaker tagging in broadcast, video communication, dubbing, localization, conferencing, and media analytics workflows.
Current; downloadable NVIDIA NIM endpoint
Real-time video relighting, virtual production, media effects, HDR-based lighting changes, and foreground/background compositing.
Current and accessible through NVIDIA AI for Media and NVIDIA NIM documentation; version 1.1.0 is documented.
Autonomous-driving research, trajectory prediction, interpretable motion planning, navigation-conditioned driving, visual question answering and safety-oriented model evaluation
Current; open weights available; also available through NVIDIA Alpamayo 1.5 NIM
Multilingual offline speech transcription, speech translation, and NeMo-based ASR research or deployment
Active; open-weight checkpoint and NVIDIA deployment options available
Fast English speech-to-text transcription, NeMo experimentation, domain fine-tuning, and Riva-based ASR deployment
Available downloadable checkpoint; older NeMo ASR model
Controllable video world generation, robotics sim-to-real augmentation, autonomous-vehicle simulation, and Physical AI synthetic-data generation
Available; legacy relative to Cosmos 3; original repository under limited maintenance
Edge physical AI, robotics, visual reasoning, world simulation, video generation, and action-policy prototyping
Current; open model; gated Hugging Face access
Physical AI, robotics, autonomous-vehicle simulation, multimodal world generation, future-state prediction, action reasoning, and synthetic training data
Current; downloadable open-weight model and available through an NVIDIA NIM endpoint
High-quality Physical AI simulation, synthetic-data generation, robotics and autonomous-vehicle research, multimodal world modeling, and teacher-model distillation
Current; open-weight model; commercially and non-commercially usable
Physical-world video and image understanding, robotic perception, embodied-agent planning, spatial-temporal reasoning, and Physical AI research
Current; downloadable and hosted NIM endpoint
Rapid global weather forecasting, ensemble simulation, climate-risk analysis, renewable-energy forecasting, and scientific weather-model research
Current and downloadable; available through NVIDIA Earth-2 FourCastNet NIM and related deployment tooling
De novo molecular design, fragment-constrained generation, linker design, scaffold decoration, hit generation, and lead optimization
Current; downloadable model and available as an NVIDIA NIM
Humanoid robot manipulation, cross-embodiment policy learning, robot demonstration fine-tuning, physical AI research, and action-sequence deployment
Current; general availability
Quantum-computing calibration plot interpretation, QPU bring-up and retuning workflows, experiment diagnosis, parameter extraction, fit-quality assessment, and domain-specific calibration agents.
Current; available through NVIDIA Build, NVIDIA NIM, and downloadable checkpoints
Generative lip dubbing, multilingual video localization, broadcasting, conferencing, and digital-human facial animation
Current; downloadable model and NVIDIA LipSync NIM; private access may be required for some workflows
Multimodal semantic search, visual document retrieval, question-answer retrieval, vector databases, and retrieval-augmented generation
Current and downloadable
Reranking text, document images, and image-text candidates in visual search, multimodal RAG, and question-answering retrieval pipelines
Current; downloadable and available through NVIDIA NIM and retrieval APIs
Multilingual voice agents, accessibility, narration, audiobooks, dubbing, localization and interactive speech applications
Current
Low-latency multilingual text translation, speech-translation pipelines, and self-hosted NVIDIA GPU deployments
Current and downloadable; available through NVIDIA NIM and NVIDIA Riva
Small-molecule generation, molecular embeddings, chemical-space exploration, lead optimization, and oracle-guided drug-design workflows
Available through NVIDIA BioNeMo Framework and NVIDIA NIM; research and development model
Detecting table cells, rows, columns, and merged-cell structure in document images for OCR alignment, table reconstruction, document ingestion, and retrieval systems
Current; downloadable and available through NVIDIA NIM/Build
Agentic reasoning, long-context analysis, coding, tool-calling workflows, retrieval-augmented generation, collaborative agents, and high-volume enterprise inference.
Current; open-weight; available through NVIDIA NIM, NVIDIA's hosted API trial endpoint, downloadable checkpoints, and self-hosted deployments.
Frontier reasoning, complex coding, long-context analysis, enterprise RAG, tool-using agents, and multi-agent workflows
Current; open-weight model with BF16 and NVFP4 checkpoints
Multilingual semantic search, dense retrieval, RAG, agentic retrieval, code search, and vector-based document matching
Current; available through NVIDIA NIM and downloadable Hugging Face weights
Multimodal document intelligence, OCR, long-video and audio understanding, cross-modal reasoning, voice agents, and self-hosted enterprise inference
Current; generally available open-weight checkpoint with hosted NVIDIA NIM access
Real-time full-duplex voice agents, interruptible conversational interfaces, speech-to-speech research, and NVIDIA GPU-based enterprise voice applications
Early access; available for evaluation through NVIDIA NIM and qualified access programs
High-volume agent execution, long-running AI agents, reasoning, coding assistance, RAG, chat, low-latency text generation, and domain customization
Current open-weight model; NIM container available for early-access evaluation
Low-latency live speech-to-text, voice interfaces, live captions, and continuous audio streams
Current and accessible through NVIDIA Speech NIM; streaming-only deployment
Multilingual text-and-image moderation, LLM and VLM guardrails, response safety evaluation, and custom enterprise safety policies
Current; open-weight model and downloadable NVIDIA NIM; latest NGC container version bf16-v1.1 as of September 17, 2026
Detecting and localizing chart titles, axis labels, legends, mark labels, and value labels in document images
Current and downloadable
English OCR, document ingestion, layout-aware text extraction, multimodal retrieval, RAG preprocessing, and enterprise document intelligence
Available; English-only OCR model; newer Nemotron OCR v2 is available for updated English and multilingual OCR deployments
Multilingual OCR, scanned documents, forms, reports, charts, tables, image-based search, document ingestion, and retrieval-augmented generation preprocessing
Current; downloadable model and available through NVIDIA NIM and NVIDIA-hosted services
Document page-layout detection before OCR, table extraction, indexing, enterprise document processing, and multimodal RAG pipelines
Current; downloadable and available through NVIDIA NIM
Document OCR, layout analysis, table and chart extraction, retrieval pipelines, document indexing, and multimodal data curation
current
Quantum-computing calibration plot analysis, experiment interpretation, fit-quality assessment, parameter extraction, and technical research workflows
Current; downloadable open-weight model
Gaze correction in video conferencing, telepresence, digital-human applications, and video-processing pipelines
Current; downloadable NVIDIA NIM model
English speech-to-text transcription, local ASR, offline processing, RAG audio extraction, and customized NeMo deployments
Available open-weight checkpoint
Local English speech transcription, offline ASR, long-form audio processing, and custom ASR fine-tuning
Current and accessible open-weight model
Taiwanese Mandarin speech-to-text, Mandarin-English code-switching, live captions, voice interfaces, and NVIDIA GPU-hosted streaming or offline transcription.
Listed in current NVIDIA Speech NIM documentation and model catalog; the associated NGC container is marked no longer supported.
Mandarin-English automatic speech recognition, code-switched transcription, streaming transcription, and self-hosted enterprise speech-to-text
Current and downloadable; available through NVIDIA NIM and Riva interfaces
High-speed English speech transcription, captions, meeting transcription, audio search, and timestamped media workflows
Current open-weight model; publicly available through Hugging Face and NVIDIA NeMo
Multilingual sentence and document translation, localization, marketing content, and developer translation workflows
Current and accessible through NVIDIA NIM and downloadable model weights
Local or self-hosted multilingual sentence and document translation across English and 36 non-English languages
Current and publicly downloadable; supported by NVIDIA NIM
Camera-only multi-view 3D perception, autonomous-driving scene analysis, bird's-eye-view visualization, and object tracking
Current NVIDIA NIM endpoint; free endpoint access requires an NVIDIA API key
Real-time enhancement of speech captured with low-quality microphones in noisy or reverberant environments, including broadcast, conferencing, telecommunications, and media production.
Current and available through NVIDIA NIM, hosted preview services, and NVIDIA audio software; downloadable deployment may require applicable NVIDIA licensing or subscription access.
AI-generated video detection, media authentication, digital forensics, content moderation, and media-integrity monitoring
Current; downloadable NIM and managed trial endpoint, with self-hosted deployment requiring AI for Media Private Access
Professional video upscaling, broadcast enhancement, streaming pipelines, pre-encoding optimization, denoising, and deblurring
Current; downloadable NVIDIA NIM
ChatGPT-style instant responses, general-purpose writing and analysis, image-aware conversations, and tool-assisted workflows
Current rolling alias; underlying model snapshot is regularly updated
Zero-shot image classification, image-text similarity, semantic image retrieval, multimodal indexing, and computer-vision research
Public research release with downloadable weights; not verified as a current OpenAI hosted API model
Codex CLI coding workflows, code question answering, code editing, repository tasks, and low-latency software-engineering assistance
Retired; API access ended on 2026-02-12
Controlled browser automation, computer-use research, UI testing, and repetitive interface workflows
Deprecated
Historical research on text-to-image generation, legacy image workflows, and comparisons with newer OpenAI image models.
Retired; deprecated and removed from the OpenAI API on May 12, 2026.
Historical text-to-image generation, concept art, illustration, visual ideation, marketing imagery, and prompt-following research
Retired; deprecated and removed from the OpenAI API on May 12, 2026
Maintaining legacy text-completion applications, historical GPT-3 base-model behavior, and existing compatible fine-tuned workflows before shutdown
Deprecated; currently accessible but scheduled to shut down on 2026-09-28
Legacy text completion, code continuation, and inference from existing davinci-002 fine-tuned models before shutdown
Deprecated; API access scheduled to shut down on September 28, 2026
Low-cost, high-volume text generation, summarization, classification, extraction, simple chatbots, and legacy API integrations
Deprecated; still available through the OpenAI API
Maintaining established GPT-4 integrations, general-purpose text generation, analysis, writing, and coding workloads
Legacy; older high-intelligence GPT model
Legacy high-context text and image analysis, function calling, JSON-mode workflows, and existing GPT-4 Turbo integrations
Deprecated but currently accessible; scheduled for shutdown on October 23, 2026
Historical long-context text generation, document analysis, structured text generation, and general-purpose assistant applications.
Retired; the gpt-4-turbo-preview alias pointed to gpt-4-0125-preview, which was shut down on 2026-03-26.
Software engineering, long-context document analysis, precise instruction following, structured extraction, tool-enabled agents, and image understanding
Current; default GPT-4.1 alias with gpt-4.1-2025-04-14 snapshot
Fast, cost-efficient instruction following, coding assistance, image understanding, structured extraction, tool calling, and long-context API applications
current
High-volume, latency-sensitive classification, extraction, routing, summarization, lightweight assistants, image-assisted analysis, and simple tool-calling workflows
Deprecated; currently accessible as of September 23, 2026; scheduled for shutdown on October 23, 2026
Historical general-purpose writing, creative work, nuanced communication, image understanding, programming assistance, and applications needing function calling or structured outputs.
Retired from the API on 2025-07-14; retired from ChatGPT in June 2026
Fast general-purpose conversations, vision, voice interactions, coding, and everyday productivity
Active
General-purpose assistants, image understanding, coding help, structured extraction, multilingual generation, and latency-sensitive API workflows
Current in the OpenAI API; retired from ChatGPT on 2026-02-13. The gpt-4o-2024-05-13 snapshot is scheduled for API shutdown on 2026-10-23.
Voice assistants, spoken conversational agents, audio-enabled customer service, and applications requiring direct audio understanding and speech generation
Retired; API access ended May 7, 2026
Low-cost, high-volume text and image understanding, classification, extraction, translation, tagging, customer support, routing, and structured data generation
Current canonical model alias; dated snapshot gpt-4o-mini-2024-07-18 is available
Lower-cost audio understanding, conversational voice interfaces, and applications requiring text and spoken-audio input/output
Deprecated; scheduled for API shutdown on 2027-01-20
Low-cost realtime voice assistants, speech-to-speech interfaces, interactive audio applications, and conversational prototypes
Deprecated; scheduled for API removal on 2027-01-20
Low-latency voice assistants, speech-to-speech applications, live translation, language learning, and interactive customer support
Retired; API access ended 2026-05-07
Historical web-search applications built around OpenAI Chat Completions
Retired; shut down on 2026-07-23
Accurate speech-to-text conversion, meeting transcription, call transcription, voice-agent input, and prompted domain-specific transcription
Deprecated; currently accessible; scheduled for API shutdown on 2027-02-26
Lower-cost multilingual speech transcription, meeting notes, call-center transcripts, voice-note conversion, and audio-to-text pipelines
Current and available
Fast, controllable text-to-speech for narration, voice interfaces, customer service, accessibility, and realtime audio applications.
Current; the canonical alias currently points to the gpt-4o-mini-tts-2025-12-15 snapshot.
Legacy Chat Completions applications requiring low-cost, search-grounded text responses
Retired; access shut down on July 23, 2026
Multi-speaker meeting, interview, call, podcast, and research transcription with speaker labels.
Deprecated; currently accessible and scheduled for removal from the API on February 26, 2027.
Complex coding, reasoning, research, long-context analysis, visual understanding, tool-using agents, and structured professional workflows
Current canonical alias, but previous-generation model; dated snapshot gpt-5-2025-08-07 is deprecated and scheduled for API shutdown on 2026-12-11
ChatGPT-aligned conversational applications, text generation, image-aware question answering, structured outputs, and tool-enabled workflows requiring GPT-5 compatibility
Deprecated
Cost-sensitive reasoning, coding assistance, structured extraction, document processing, high-volume automation, and tool-enabled workflows
Current API alias; dated snapshot gpt-5-mini-2025-08-07 is deprecated
High-volume classification, summarization, extraction, ranking, routing, image-assisted analysis, and lightweight coding subagents
Deprecated dated snapshot; currently accessible until scheduled shutdown on 2026-12-11
Difficult research, mathematics, science, complex coding, high-stakes analysis, and tool-using workflows where maximum answer quality matters more than latency or cost.
Current canonical alias; dated snapshot gpt-5-pro-2025-10-06 is deprecated and scheduled for shutdown on 2026-12-11.
Agentic software engineering, repository-level coding, code review, debugging, refactoring, test generation, and frontend work using screenshots
Retired; API access shut down on 2026-07-23
Coding, long-context analysis, tool-using agents, structured outputs, and multi-step workflows
Current
Conversational assistants, instruction following, image-grounded chat, structured extraction, streaming responses, and tool-using API workflows.
Retired; the API alias gpt-5.1-chat-latest was shut down on 2026-07-23. GPT-5.1 models were retired from ChatGPT on 2026-03-11.
Agentic software engineering, code generation, debugging, refactoring, testing, code review, and long-running Codex workflows
Retired; API access shut down on July 23, 2026
Long-running agentic coding, repository-scale refactoring, multi-file implementation, debugging, code review, pull-request creation, and extended Codex workflows.
Retired; API access ended 2026-07-23
Cost-sensitive agentic coding, code editing, repository maintenance, and Codex-style workflows
Retired; API access shut down on 2026-07-23
Complex professional work, long-context analysis, coding, document and spreadsheet workflows, visual understanding, and multi-step agents
Currently available; previous flagship model
ChatGPT-aligned conversational applications, general writing, summarization, translation, vision-enabled assistants, and tool-calling workflows
Retired; API access ended on 2026-08-10
Complex professional reasoning, advanced analysis, scientific and mathematical work, high-quality coding, long-context document analysis, and tool-using workflows.
Previous Pro model; currently available through the Responses API
Long-horizon agentic coding, large refactors, code migrations, repository-scale changes, terminal workflows, Windows development and defensive cybersecurity
Retired; API access shut down on 2026-07-23
Fast general-purpose conversation, writing, summarization, text-and-image understanding, streaming responses, and function-calling applications
Retired; API access ended 2026-08-10
Long-running agentic software engineering, codebase maintenance, debugging, testing, web development, tool-driven development, and technical computer workflows
Current and available through OpenAI API and Codex surfaces
Complex professional work, advanced reasoning, software engineering, long-horizon agents, visual document analysis, computer use, research, and tool-heavy workflows
Current
High-volume coding assistants, computer-use agents, subagents, tool calling, image reasoning, document workflows, and latency-sensitive applications
Current
High-volume classification, data extraction, ranking, image understanding, routing, and lightweight coding subagents
Current; API-only model
High-stakes reasoning, professional knowledge work, long-context analysis, complex coding, web research and agentic workflows requiring maximum answer quality
Current; available in ChatGPT for Pro and Enterprise users and in the Responses API for developers
Authorized vulnerability research, defensive cybersecurity operations, malware analysis, security testing, and binary reverse engineering.
Deprecated; currently accessible through restricted Trusted Access for Cyber channels; scheduled for API shutdown on October 1, 2026.
Complex coding, long-context research, professional analysis, tool-heavy agents, computer use, and multi-step workflow execution
Current; available through the OpenAI API, ChatGPT, and Codex
High-accuracy reasoning, complex coding, long-context research, data analysis, and multi-step professional workflows
Current
Authorized vulnerability research, exploit validation, exploit-chain development, advanced security testing, vulnerability triage, and defensive cybersecurity agents
Current; restricted access through OpenAI Daybreak Red with separate approval and provisioning
High-volume classification, summarization, routing, extraction, document understanding, agent automation, routine coding assistance, and cost-sensitive tool-using applications.
current
Complex reasoning, coding, research, cybersecurity, science, long-context analysis, document-heavy workflows, and tool-using agents
Generally available
Cost-conscious reasoning, coding agents, long-context analysis, structured business automation, research workflows, and tool-enabled production applications
Generally available
Complex reasoning, agentic coding, computer use, web research, scientific and professional workflows, and long-context document tasks
Current; rolling out through the OpenAI API and selected ChatGPT, Azure, and Amazon Bedrock offerings
High-volume reasoning, document analysis, coding assistance, retrieval-augmented generation, and repeatable agent workflows
Current and available
Complex coding, long-context reasoning, software engineering, research, computer use, and agentic workflows with tools.
Current; generally available through the OpenAI API
Audio-enabled chat applications, voice interfaces, spoken assistants, and applications requiring direct audio understanding and generation through Chat Completions.
Deprecated; scheduled for shutdown on January 20, 2027
Audio-in, audio-out conversational applications using the Chat Completions API, including voice assistants and tool-enabled spoken interfaces.
Current; generally available
Cost-sensitive, turn-based audio conversations, voice assistants, and audio-enabled applications using function calling
Deprecated; currently accessible but scheduled for API removal on 2027-01-20
Precise image editing, detailed creative work, high-fidelity generation, infographics, layouts, and workflows where fewer retries matter more than minimum latency
Current
Natural low-latency voice agents, customer support, conversational workflows, live assistance, and applications requiring interruption-aware speech interaction
Current; available in the OpenAI API
Low-latency live captions, realtime call transcription, microphone streams, telephony audio, and voice-interface speech recognition
Current and generally available for realtime transcription
Local and private reasoning applications, coding assistants, agentic workflows, on-device or edge inference, fine-tuning, and cost-sensitive deployments with suitable hardware.
Current open-weight model; downloadable and usable through self-hosted or third-party inference infrastructure. Not served through the OpenAI API or ChatGPT.
Self-hosted reasoning, coding, agentic workflows, private deployments, research, and fine-tuning
Current open-weight model; downloadable and deployable locally or through third-party providers; not available through the OpenAI API
Policy-based safety classification, LLM input/output filtering, content labeling, trust and safety review, and self-hosted moderation workflows
Research preview; currently available as an open-weight model
Custom-policy safety classification, LLM input and output filtering, trust and safety labeling, nuanced moderation review, and offline safety analysis
Research preview; open-weight and downloadable
Low-latency speech-to-speech voice agents, realtime customer support, education, accessibility, and conversational applications with function calling
Deprecated; scheduled for API shutdown on January 20, 2027
Low-latency speech-to-speech voice agents, customer support, realtime assistants, and audio applications that need function calling.
Active and currently available
Reasoning voice agents, speech-to-speech applications, customer support, live assistants, tool-driven workflows, and long conversational sessions
Current
Low-latency speech-to-speech agents, customer-service voice workflows, realtime tool use, telephony, and multimodal assistants with image input
current
Low-latency spoken translation, multilingual calls, live interpretation, broadcasts, meetings, lessons, video rooms, captions, and translated audio experiences.
Current
Low-latency live transcription, captions, meeting notes, call analysis, voice-agent input, and continuous speech-to-text workflows
Current and available through the OpenAI Realtime API for realtime transcription
Cost-sensitive realtime voice agents, speech-to-speech applications, interactive assistants, and multimodal interfaces
Deprecated; currently accessible but scheduled for API shutdown on 2027-01-20
Lower-cost, low-latency realtime voice agents, speech-to-speech assistants, and tool-enabled conversational applications
Current
Governed biology, genomics, medicinal chemistry, protein analysis, drug discovery, literature synthesis, wet-lab troubleshooting, and scientific tool workflows
Generally available to eligible organizations through the trusted-access program; approved internal life sciences research only
High-accuracy transcription of recorded audio, streamed file transcripts, multilingual recordings, and domain-specific speech with keyword or language hints
Current
Existing ChatGPT image-generation and image-editing integrations
Deprecated; currently accessible; scheduled for shutdown on 2026-12-01
API-based image generation, image editing, reference-image workflows, inpainting, marketing assets, e-commerce imagery, and visual content production
Deprecated; currently accessible and scheduled to shut down on 2026-10-23
Production image generation, image editing, branded graphics, ecommerce product imagery, marketing assets, and workflows requiring preservation of important visual details
Deprecated; currently accessible with API shutdown scheduled for 2026-12-01
High-quality text-to-image generation, reference-based image editing, text-heavy visual assets, product imagery, marketing creatives, and production design workflows.
Active; GPT Image 2.5 models are available for newer workflows, but GPT-Image-2 remains accessible as a documented API model.
Cost-sensitive image generation and editing, high-volume variations, rapid ideation, previews, lightweight personalization, and draft creative assets.
Deprecated; currently accessible but scheduled for API shutdown on 2026-12-01.
Fast, high-quality image generation and editing, creator content, product experiences, visual search, rapid prototyping, and high-volume workflows
Current; available through the OpenAI API
Complex reasoning, advanced coding, mathematics, science, technical research, visual analysis and multi-step tool workflows
Current canonical alias; o3-2025-04-16 snapshot deprecated and scheduled for API shutdown on December 11, 2026
Complex reasoning, mathematics, science, coding analysis, visual reasoning, and high-accuracy multi-step tasks
Deprecated; still documented in the OpenAI API model catalog
Historically, difficult mathematics, science, coding, and other multi-step reasoning tasks requiring extended deliberation
Retired; API access shut down on 2025-07-28
Cost-sensitive mathematics, science, algorithmic programming, debugging, and text-only reasoning
Deprecated
Complex reasoning, difficult technical analysis, advanced programming, research workflows, and tasks where answer consistency matters more than latency or cost.
Deprecated in OpenAI's current model catalog; the dated snapshot o1-pro-2025-03-19 is also marked deprecated. No exact shutdown date for the canonical o1-pro alias was found in the reviewed official documentation.
Complex multi-step research, source synthesis, legal and scientific analysis, market research, and large-scale internal-data investigation
Deprecated
Coding, mathematics, science, technical analysis, structured extraction, text-to-SQL, and multi-step reasoning
Current canonical alias with deprecated snapshot; o3-mini-2025-01-31 is scheduled for API shutdown on 2026-10-23
High-reliability reasoning, advanced mathematics, scientific analysis, complex coding, research, and multi-step professional work.
Current canonical alias; the dated snapshot o3-pro-2025-06-10 is marked deprecated in the model documentation.
Fast, cost-sensitive reasoning; coding; mathematics; visual analysis; structured extraction; high-volume tool-using agents
Deprecated; currently available through the API; scheduled for shutdown on 2026-10-23
Complex multi-step research, source synthesis, market analysis, legal or scientific research, and long-form evidence-based reports.
Current canonical alias; the dated snapshot o4-mini-deep-research-2025-06-26 is deprecated.
Text and image safety classification, content filtering, AI-output screening, policy enforcement, and human-review routing
Current
Text and image safety classification, content filtering, moderation queues, policy enforcement, and generated-content screening
Current default moderation model
Rapid video concepting, social clips, image-to-video experiments, prototypes, rough cuts, and audiovisual creative iteration
Deprecated; currently accessible through the API as of September 23, 2026, with shutdown scheduled for September 24, 2026
Production-quality text-to-video and image-guided video generation, cinematic prototypes, marketing assets, and high-resolution short clips with synchronized audio.
Legacy; deprecated; currently accessible through September 23, 2026; scheduled for API shutdown on September 24, 2026
High-quality semantic search, multilingual retrieval, RAG, recommendations, clustering, classification and similarity matching
Current and available through the OpenAI API
Cost-efficient semantic search, retrieval-augmented generation, clustering, recommendations, anomaly detection, and text or code similarity
Current
Legacy semantic search, retrieval, clustering, recommendations, anomaly detection, and classification systems already built around ada-002 vectors
Older embedding model; currently listed and accessible through the embeddings API
Legacy text-only safety classification and historical moderation integrations
Retired; access ended October 27, 2025
Low-latency text-to-speech, realtime-oriented voice interfaces, narration, accessibility, and automated audio generation
Current and accessible; optimized for low-latency text-to-speech
High-quality text-to-speech generation, narration, accessibility audio, voice interfaces, and downloadable speech content
Current; available through the OpenAI Audio API speech endpoint
Multilingual audio transcription, English speech translation, language identification, subtitles, captions, and word- or segment-level timestamps.
Deprecated; currently accessible through the OpenAI API and scheduled for shutdown on February 26, 2027.
Cost-sensitive cross-modal retrieval, image and video search, multimedia catalog indexing, and vector search
Current and available through Alibaba Cloud Model Studio International deployment
Cross-modal retrieval, text-to-image search, image similarity, video search, semantic classification, clustering, and multimodal vector indexing.
Current and accessible through Alibaba Cloud Model Studio in the China (Beijing) region; free trial pricing is listed.
Agent-environment simulation, tool-interaction modeling, terminal and software-engineering trajectories, and research on language world models
Current open-weight model
Low-latency duplex voice assistants, real-time customer service, AI companions, and streamed speech-to-speech applications
Current and accessible; standard-edition real-time duplex speech model
Low-latency voice assistants, customer service, AI companions, full-duplex spoken interaction, and voice applications using tools or cloned voices.
Current and available
Low-latency voice assistants, real-time customer service, duplex speech conversations, interactive voice agents, and applications requiring streaming audio responses.
Available; current API access confirmed, but newer models are recommended for some new projects
Real-time multilingual speech transcription, live captions, meeting transcription, voice interfaces, streaming subtitles, and Chinese-dialect recognition.
Current and available through Alibaba Cloud Model Studio in the International/Singapore and China (Beijing) regions.
Expressive text-to-speech, audiobooks, film and video dubbing, content creation, premium voice services, multilingual speech, dialect synthesis, and voice cloning
Current and available
Long-form offline transcription of meetings, interviews, calls, media files, and multilingual or dialect-rich recordings
Current and publicly available
Autonomous-driving research, driving-scene VQA, 3D BEV perception, trajectory prediction, and embodied-AI experimentation
Current open-weight research release
Text-to-image generation, image editing, text rendering in images, photorealistic scenes, creative design, and producing multiple image variants.
Current; accelerated model; functionally equivalent to qwen-image-2.0-2026-03-03
Cost-sensitive text-to-image generation, posters, marketing graphics, illustrations, and images containing Chinese or English text
Current and accessible; currently equivalent to qwen-image
Natural-language single-image editing, bilingual text changes, object insertion or removal, style transfer, pose changes, and image fusion
Current and accessible
High-quality image editing, multi-image composition, industrial design concepts, geometric transformations, character-consistent edits, and controlled visual revisions
Current and available; canonical model ID is functionally equivalent to qwen-image-edit-max-2026-01-16
Professional text-to-image generation, image editing, posters, infographics, multilingual in-image text, photorealistic scenes, and reference-based creative production
Current rolling model; functionally equivalent to qwen-image-2.0-pro-2026-04-22
General-purpose text generation, long-context analysis, multilingual applications, structured business workflows, function-calling agents, and applications that need optional reasoning mode.
Current; qwen-plus currently resolves to the qwen-plus-2025-12-01 snapshot
Complex image and video understanding, document analysis, chart interpretation, visual question answering, and structured extraction
Currently accessible legacy visual language model; current qwen-vl-max endpoint is functionally equivalent to qwen-vl-max-2025-08-13
Self-hosted chat assistants, multilingual text generation, coding and mathematics assistance, long-context document work, structured text generation, and cost-sensitive private deployments
Current open-weight model; publicly available for download and self-hosted deployment
Self-hosted multilingual assistants, document processing, RAG, coding support, structured extraction, and cost-conscious production deployments.
Legacy open-weight model; downloadable and self-hostable, while Qwen2.5 API models marked deprecated are no longer callable through current Alibaba Cloud Model Studio pricing documentation.
Self-hosted assistants, multilingual text generation, long-document processing, coding support, structured extraction, RAG, and agent applications
Available; open-weight model
Self-hosted multilingual assistants, coding, mathematics, document analysis, structured text generation, and long-context workloads
Available as an open-weight model; legacy relative to newer Qwen generations
Multimodal assistants, audio and video understanding, visual question answering, voice interaction, speech instruction following, and local multimodal AI research
Current; open-weight model and available through Alibaba Cloud Model Studio
Local or self-hosted image and video understanding, OCR, document extraction, chart and diagram analysis, visual question answering, visual grounding, and multimodal research
Open-weight and downloadable; supported as a fine-tuning base model in Alibaba Cloud Model Studio; current hosted inference availability and standard pricing for this exact model are not clearly listed in the latest Model Studio inference-pricing catalog.
High-quality image, document, chart, screenshot, OCR, visual-grounding, and video analysis; multimodal agents and self-hosted experimentation
Available open-weight model; older Qwen2.5-VL generation and still listed by Alibaba Cloud Model Studio
Fast, high-volume text generation; long-context analysis; summarization; extraction; structured outputs; and applications needing optional reasoning.
Current and available; canonical qwen-flash identifier is functionally equivalent to qwen-flash-2025-07-28
Lightweight local assistants, offline prototypes, embedded experimentation, education, simple text generation, and resource-constrained deployments
Current open-weight model
Efficient local inference, edge applications, multilingual chat, lightweight reasoning, coding assistance, and tool-enabled agents
Current open-weight model
Local and self-hosted chat, compact reasoning, coding assistance, multilingual applications, retrieval-augmented generation, and lightweight tool-using agents
Current open-weight model
Cost-efficient reasoning, coding, multilingual assistants, local deployment, structured text generation, tool-enabled agents, and fine-tuned applications
Current and accessible; open-weight model available for self-hosting and Alibaba Cloud Model Studio API deployment
Local deployment, multilingual assistants, reasoning, mathematics, coding, structured text generation, research, and cost-sensitive agent workflows.
Current open-weight model; also available as qwen3-14b through Alibaba Cloud Model Studio, with regional capability and pricing differences.
Self-hosted assistants, coding, mathematical and logical reasoning, multilingual applications, long-context document processing, and agentic tool-use systems
Current and accessible; open-weight release with hosted Alibaba Cloud Model Studio availability
Self-hosted reasoning assistants, coding agents, mathematics, multilingual applications, structured text generation, and tool-calling workflows
Current; open-weight model and available through Alibaba Cloud Model Studio
Complex reasoning, mathematics, software development, multilingual applications, function calling, agentic workflows, research and self-hosted open-weight deployment
Current and accessible through Alibaba Cloud Model Studio; original Qwen3 open-weight release, with newer 2507 instruct and thinking variants available separately
Multilingual semantic search, RAG candidate reranking, enterprise document retrieval, knowledge-base search, and improving search-result relevance.
Current and available
Low-cost multilingual speech-to-text, language identification, offline transcription, real-time streaming ASR, and high-throughput deployments
Current, open-weight, Apache 2.0 licensed
Multilingual speech transcription, language identification, long-audio processing, and self-hosted or streaming ASR applications
Current; open-weight and downloadable
Repository-scale coding agents, code generation, code completion, debugging, refactoring, terminal workflows, and cost-sensitive self-hosted deployments.
Current; open-weight model and available through Alibaba Cloud Model Studio
Large-codebase analysis, code generation, refactoring, debugging, documentation and long-context coding-agent workflows
Current; canonical qwen3-coder-plus currently equivalent to qwen3-coder-plus-2025-09-23
Streaming translation of recorded or uploaded audio and video, multilingual subtitles, translated voice tracks, and applications requiring translated text or synthesized speech.
Current stable model
Real-time multilingual speech interpretation, live voice translation, conference translation, streaming media, and audiovisual translation with text or synthesized speech output
Legacy; still available; no longer recommended for new use
Complex reasoning, coding assistance, web-grounded agents, function calling, structured extraction, and long-context text analysis
Current
Local visual assistants, OCR, document and chart analysis, image question answering, lightweight video understanding, and multimodal prototyping
Current open-weight model
Local image and video understanding, OCR, document extraction, visual question answering, visual coding, and lightweight multimodal agents
Current open-weight model; available on Hugging Face and supported for supervised fine-tuning in Alibaba Cloud Model Studio
Local or hosted image and video understanding, OCR, document extraction, visual question answering, spatial reasoning, screenshot analysis, multimodal agents, and structured data extraction
Current; open-weight model with hosted inference available through Alibaba Cloud Model Studio
Image and video understanding, OCR, document analysis, spatial reasoning, visual coding, long-context multimodal tasks, and visual-agent applications
Current and available; open-weight release and Alibaba Cloud Model Studio API model
Document intelligence, OCR, image and video understanding, spatial reasoning, visual coding, and visual-agent applications
Current; open-weight checkpoint and available through Alibaba Cloud Model Studio managed inference
High-quality image and video understanding, OCR, document intelligence, visual coding, spatial reasoning, long-context multimodal analysis and visual-agent applications
Current and accessible; open-weight release and Alibaba Cloud Model Studio API availability
Long-context multimodal analysis, document understanding, video and image interpretation, general reasoning, coding, tool-enabled assistants, and self-hosted deployment
Accessible; no longer recommended for new projects
Efficient multimodal assistants, coding, reasoning, long-context analysis, local deployment, and tool-using agents
Available; open-weight Apache 2.0 model with hosted API access through Alibaba Cloud Model Studio
Advanced multimodal reasoning, image and video understanding, coding, long-context analysis, document and chart interpretation, function-calling agents, and web-grounded workflows.
Current and available through Alibaba Cloud Model Studio; released globally on February 24, 2026.
Advanced multimodal reasoning, coding, video and image understanding, long-context analysis, tool-using agents, and self-hosted open-weight deployments
Current; open-weight model and available through Alibaba Cloud Model Studio
Fast long-context text, image and video understanding; structured extraction; tool-enabled agents; web-grounded applications; and high-volume multimodal workloads.
Current and available
Long-context reasoning, multimodal document and video analysis, coding, structured enterprise automation, function-calling agents, and web-grounded research.
Current; canonical rolling model identifier currently equivalent to qwen3.5-plus-2026-02-15
Fast multimodal analysis, long audio understanding, audiovisual question answering, voice assistants and spoken-response applications
Current; canonical model ID functionally equivalent to qwen3.5-omni-flash-2026-03-15
Low-latency voice assistants, speech-to-speech applications, realtime multimedia analysis, interactive agents, and multimodal conversations.
Current and available; rolling model identity functionally equivalent to qwen3.5-omni-flash-realtime-2026-03-15
Multilingual voice assistants, speech-enabled multimodal applications, audio-visual analysis, spoken explanations, accessibility tools, and interactive media workflows.
current
Real-time voice assistants, speech-to-speech applications, multimodal customer service, visual conversational agents, live multimedia analysis, and interactive applications requiring controllable speech output.
Current and available; canonical model identifier functionally equivalent to snapshot qwen3.5-omni-plus-realtime-2026-03-15
Coding agents, repository-level software engineering, visual document analysis, video understanding, STEM reasoning, long-context assistants, and self-hosted multimodal applications
Current and available; open-weight release and Alibaba Cloud Model Studio API model
Agentic coding, long-context software engineering, multimodal analysis, tool-using agents, and self-hosted deployments
Current; open-weight and available through Alibaba Cloud Model Studio
Fast multimodal assistants, coding agents, visual document analysis, video understanding, tool-using workflows, object localization, and large-context applications
Current
Advanced coding agents, front-end development, long-context analysis, structured API workflows, and text generation with web search
Preview; currently accessible; scheduled for deprecation on October 10, 2026
Long-context multimodal analysis, agentic coding, OCR, object localization, frontend development, visual reasoning, and tool-enabled enterprise assistants
Current; rolling model ID qwen3.6-plus is currently equivalent to qwen3.6-plus-2026-04-02
Fast multimodal agents, visual coding, tool-use workflows, search agents, and long-context document or screen analysis
Current and available; rolling qwen3.7-flash identifier is currently equivalent to qwen3.7-flash-2026-07-15
Long-context reasoning, multimodal document and video analysis, coding, tool-using agents, structured extraction, and enterprise productivity workflows
Current; rolling identifier currently equivalent to qwen3.7-plus-2026-05-26
Multilingual semantic search, retrieval-augmented generation, code retrieval, recommendation, clustering, classification, and large-scale text vectorization
Current and available through Alibaba Cloud Model Studio internationally
Advanced reasoning, coding, scientific and professional research, long-context analysis, and long-horizon agent workflows
Current and available; open-weight release and hosted API model
Coding assistants, repository analysis, visual document workflows, long-context research, office automation, multimodal agents, and tool-using applications
Active and currently available; open-weight release and Alibaba Cloud Model Studio API model
Fast long-context reasoning, coding assistance, visual document and chart analysis, video understanding, function-calling agents, and high-concurrency applications
Current and accessible through Alibaba Cloud Model Studio
High-volume coding, agentic workflows, long-context document and codebase analysis, office automation, and multimodal text-image-video understanding
Current open-weight experimental preview
Complex coding, autonomous software engineering, long-horizon agent workflows, professional document analysis, visual reasoning, long videos, and demanding research tasks.
Stable official release; currently available through Alibaba Cloud Model Studio
Real-time speech translation, multilingual meetings, live interpretation, translated voice communication, and audiovisual translation with low latency.
Current stable model
Long-form audio and video understanding, multimedia analysis, audio-visual agents, content summarization, and tool-using workflows
Current and available
Real-time voice assistants, speech-to-speech applications, interactive video agents, live media analysis, multimodal customer service, meeting and collaboration interfaces, and applications requiring tool or MCP integration.
Current and available; international deployment in Singapore
Real-time moderation of streamed prompts and language-model responses
Available open-weight model
Low-latency character dialogue, virtual companions, game NPCs, role-playing applications, IP character replication, and conversational smart devices
Current; dynamically updated managed model
Complex text-to-image layouts, multilingual typography, posters, menus, storyboards, interface mockups, product visuals, and image editing with one to three reference images
Current and available
Instruction-based image editing, multi-image fusion, character-consistent compositions, object replacement, style transfer, poster and text editing, and product-image variation.
Current; canonical model currently equivalent to qwen-image-edit-plus-2025-10-30
Realistic text-to-image generation, creative concepts, marketing visuals, editorial artwork, product concepts, and general-purpose image creation.
Current and available for text-to-image generation; the undated qwen-image-max identifier is currently equivalent to qwen-image-max-2025-12-30.
Fast text-to-image generation, text-heavy layouts, image editing, reference-image compositing, marketing graphics, and high-volume image workflows.
Current and available through Alibaba Cloud Model Studio
Complex multimodal analysis, long-context document understanding, image and video question answering, audio-aware workflows, coding, and enterprise applications requiring broad input support.
Available/listed for API pricing, but not identified as a baseline model that is always publicly available; account-dependent availability should be verified.
Low-latency image and video understanding, object detection, physical AI, robotics, edge deployment, and on-device visual assistants
Current and publicly available
Fast multimodal applications, document and image analysis, short-video understanding, structured extraction, multilingual chat, coding, and tool-using agents
Current and publicly available baseline model
Low-latency reasoning, coding assistance, function calling, local deployment, on-device applications and cost-sensitive inference
Current and accessible; open-weight research-preview model with API availability through Reka Gateway
Coding, mathematical reasoning, local inference, research experimentation, and fine-tuning for agentic workflows
Available as an open-weight model; official API availability was announced but the current public baseline API catalog does not list the exact model identifier
Audio-driven digital humans, lip-sync video, singing avatars, multilingual character animation, multi-person dialogue, and long-duration talking-video generation
Current; ongoing project
Visual question answering, high-resolution image understanding, multimodal search, image-grounded research, and tool-assisted agentic reasoning
Current open-weight research model
Visual deep-search, fine-grained image understanding, multimodal agent research, and self-hosted tool-using vision-language applications
Current open-weight model
Spatial intelligence research, image-text question answering, visual reasoning, 3D scene analysis, and solid-geometry problems
Current open-weight release
Native image generation, high-resolution visual creation, image editing, infographic and layout generation, visual understanding, and multimodal research or creative workflows
Current open-weight flagship checkpoint
Unified computer-vision research, detection, OCR, segmentation, depth and normal estimation, visual grounding, and multi-view geometry
Current open-weight research model
Long-horizon multimodal agents, data analysis, deep research, complex information presentation, office automation, and tool-driven workflows
Preview; currently available through the SenseNova Token Plan
Character-consistent image generation for multi-episode videos, motion comics, short dramas, storyboards, and cross-shot visual production
Current; integrated into the SenseNova Seko series and Seko 2.0 platform
Professional image creation, infographics, advertising, e-commerce assets, presentations, educational diagrams, storyboards, and multi-step visual delivery workflows
Current production release; enterprise API available through whitelist access
Open-source visual understanding, image generation, image editing, infographic creation, visual reasoning, and continuous image-text workflows.
Open-source model series; U1-8B-MoT and U1-A3B-MoT variants are available. SenseNova U1.5 is the newer successor for visual creation and editing, while U1 checkpoints remain documented and downloadable.
Fast infographic generation, dense visual explanations, charts, diagrams, presentation graphics, and information-heavy layouts
Current and accessible through SenseNova Token Plan and dedicated image-generation integrations
Text-to-image generation, reference-image creation, visual design, posters, infographics, product imagery, and iterative image editing
Current and available through the SenseNova Token Plan
Coding agents, software engineering, long-context reasoning, tool-using agents, private local inference, and work-centric automation
Current; open-weight model available for local deployment and accessible through StepFun's API platform
High-throughput coding agents, visual document and UI understanding, search-heavy research workflows, long-context analysis, and tool-using autonomous agents
Current and available
Long-running agentic workflows, software engineering, coding, professional knowledge work, financial analysis, research, and vision-assisted tasks
Current preview; available through StepFun products and API. Open-weight release planned for October 15, 2026.
Multilingual transcription, meetings, subtitles, live content, customer-service recordings, specialized terminology, noisy audio, dialects, code-switching, and singing or music transcription
Current and available through the StepFun API
Zero-shot text-to-speech, natural-language voice design, singing and vocal generation, music, sound effects, ambience, and complete multi-element audio scenes
Current research and platform-listed model; public API and access details are limited
Song generation, instrumental music, lyric-to-song workflows, accompaniment, cover-style synthesis, and rapid music prototyping
Current; publicly showcased and available through StepFun's StepAudio 3 music experience
Natural realtime voice conversation, full-duplex interaction, interruption-aware assistants, emotional audio understanding and voice agents that use tools.
Current and publicly listed
Controllable multilingual text-to-speech, narration, voice interfaces, localization, and expressive spoken-audio generation
Current and publicly accessible
On-device visual assistants, GUI grounding, screen understanding, spatial reasoning, automotive interaction, and low-latency local agent workflows
Current; officially announced as an edge-deployment base model
On-device speech recognition, audio understanding, voice assistants, in-vehicle interaction, privacy-sensitive audio processing, and low-latency edge AI.
Current; edge-deployment model with public capability information but no verified public hosted API pricing or general model-download documentation located.
On-device text-to-image generation, local image editing, privacy-sensitive creative features, and low-latency edge-device applications.
Current research model; public technical information is available, but official public API availability and downloadable weights were not verified.
Low-latency desktop and mobile GUI automation, visual grounding, local computer-use agents, and privacy-sensitive edge workflows
Current; edge-deployment GUI agent model
Self-hosted text generation, research, quantization, and domain-specific fine-tuning
Available open-weight model; older Falcon generation with newer successors
Research, self-hosted text generation, model fine-tuning, summarization, and general language-model experimentation
Available as open weights; legacy-generation model
Research, self-hosted text generation, language-model evaluation, and domain-specific fine-tuning when substantial GPU infrastructure is available
Available open-weight model; legacy-generation checkpoint
Memory-efficient local text generation, edge-device experimentation, research, and downstream fine-tuning
Current open-weight model; downloadable from Hugging Face
Efficient local, edge, and resource-constrained English text generation; experimentation with BitNet models; full fine-tuning on the supplied prequantized revision.
Current open-weight/downloadable model
Memory-efficient local text generation, edge deployment, BitNet research, continued pretraining, and custom fine-tuning
Current open-weight model; publicly downloadable
Lightweight local text generation, continued pretraining, domain adaptation, language-model research, and edge-oriented deployments
Current open-weight model
Lightweight local chat, text generation, edge deployment, prototyping, and fine-tuning experiments
Current; open-weight; instruction-tuned
Local multilingual text generation, long-context experimentation, research, and downstream fine-tuning
Current open-weight model; available for download and self-hosted inference
Efficient local text generation, multilingual language modeling, long-context applications, model adaptation, and compact reasoning research.
Current open-weight model
Efficient local inference, compact conversational assistants, multilingual text generation, long-context applications, and resource-constrained deployments
Current; open-weight
Efficient local or private deployment, multilingual instruction following, compact conversational applications, coding assistance, mathematics, and long-context text processing
Current open-weight model; downloadable and usable for self-hosted inference
Local text generation, multilingual research, domain adaptation, fine-tuning, long-context experiments, and resource-conscious deployment.
Current open-weight model; downloadable from Hugging Face and usable with supported local inference frameworks.
Local multilingual chat, lightweight instruction following, retrieval-augmented generation, efficient text generation, and cost-sensitive self-hosted deployments
Current open-weight model
Private or local multilingual text generation, long-context experimentation, domain adaptation, and fine-tuning from a 7B-scale open-weight foundation
Current open-weight model; pretrained base checkpoint
Self-hosted multilingual assistants, long-context text processing, coding, research, private deployments, and quantized local inference
Current open-weight model
Self-hosted multilingual text generation, long-context research, coding, RAG, foundation-model experimentation, and downstream fine-tuning
Current open-weight model; available from the official Hugging Face repository
Long-context text generation, multilingual instruction following, document processing, retrieval-augmented generation, coding and self-hosted enterprise deployments
Current open-weight model; publicly downloadable and usable with Transformers, vLLM and llama.cpp
Arabic NLP, dialect-aware assistants, long-context document analysis, summarization, multilingual reasoning, and locally deployed applications
Current; open model family
Ultra-lightweight local text generation, edge deployment, offline instruction following, rewriting, extraction, and embedded AI experiments
Current open-weight model
Lightweight local Python generation, fill-in-the-middle completion, edge deployment, code education, and low-resource developer tools
Current and publicly accessible open-weight model
Lightweight local text generation, instruction-following, edge devices, embedded applications, experimentation, and privacy-sensitive deployments.
Current open-weight model
Lightweight local reasoning, edge deployment, offline assistants, experimentation, and compact text-generation applications
Current open-weight model; downloadable and available for self-hosted inference
Lightweight local function calling, edge automation, API argument generation, and resource-constrained assistants
Current downloadable open-weight model
Local reasoning experiments, edge deployment, offline assistants, privacy-sensitive text processing, and resource-constrained inference
Available open-weight checkpoint
Ultra-lightweight local reasoning, embedded applications, edge deployment, experimentation, and low-memory text generation
Current open-weight model
Mathematical reasoning, programming, long-context analysis, local deployment, self-hosted inference, and test-time scaling
Current open-weight model
Research, multilingual text generation, fine-tuning, quantization, and self-hosted inference
Available as an open-weight pretrained base model
Local inference, multilingual text completion, research, continued pretraining, domain adaptation, and fine-tuning on constrained hardware
Available open-weight pretrained base model
Lightweight local assistants, multilingual instruction following, extraction, classification, education, and resource-conscious deployments
Current open-weight model; downloadable from Hugging Face
Fine-tuning, multilingual text generation, research, compact local inference, and edge-oriented deployments
Available open-weight model
Self-hosted chat assistants, multilingual text generation, lightweight coding and mathematics, local inference, and resource-conscious deployments
Available as an open-weight model
Fine-tuning, multilingual text generation, code and mathematics experiments, research, and self-hosted inference
Current open-weight model; downloadable and self-hostable
Self-hosted multilingual chat, instruction following, reasoning, mathematics, coding, research, and long-context text generation
Available open-weight model
Fine-tuning, multilingual text generation, language-model research, code and mathematics experimentation, and self-hosted inference
Available open-weight model
Self-hosted multilingual assistants, STEM and mathematics tasks, coding, instruction following, research, and local function-calling systems.
Available open-weight model; official repository remains accessible. No official deprecation or shutdown date found.
Local English text generation, open-weight research, long-sequence experimentation, and memory-conscious inference
Available for download; legacy or superseded by Falcon-H1-7B-Base
Local or self-hosted document OCR, formula recognition, table extraction, receipts, invoices, papers, and layout-aware document parsing
Current open-weight release
Natural-language object grounding, open-vocabulary detection, promptable instance segmentation, crowded-scene perception, robotics and visual inspection pipelines
Current; open-weight; Apache-2.0 licensed
Open-source text-to-speech, reference-voice generation, speech and lyric editing, emotion and timbre transformation, speech enhancement, and source separation
Current open-source release
Lightweight local assistants, edge-oriented inference, long-context text processing, prototyping, quantized deployment, and low-resource instruction-following workloads.
Current open-weight release
Efficient local text generation, long-context analysis, lightweight reasoning, coding assistance, and agent-oriented applications
Current open-weight model; publicly available for download and local deployment
Local and self-hosted text generation, Chinese and multilingual instruction following, long-context analysis, mathematics, coding, reasoning, and lightweight agent workloads
Current open-weight model; publicly available for self-hosted deployment
Chinese and multilingual text generation, reasoning, coding, long-context workloads, private local inference, quantized deployment, and domain fine-tuning.
Current open-weight downloadable model
Open-weight reasoning, long-context analysis, mathematics, science, coding, agent workflows, and cost-conscious self-hosted inference
Current and accessible as an open-weight model; also listed as the Tencent Cloud API model hunyuan-a13b. Tencent Cloud documentation notes an ongoing migration of Hunyuan services toward TokenHub.
Large-scale Chinese and English text generation, reasoning, mathematics, coding, long-context analysis, research, and self-hosted experimentation
Available as an open-weight release; legacy for the historical Tencent Cloud hunyuan-large API identifier
Self-hosted multilingual translation, localization, translation research, and Chinese dialect or minority-language translation
Current open-weight model
Image-to-3D asset creation, game and virtual-world content, product visualization, design prototyping, and self-hosted 3D generation
Current open-weight release; self-hosted deployment
Local image-to-3D asset generation, textured mesh creation, game and design prototypes, Blender workflows, and research on open 3D generative models
Available open-weight model system; superseded by the newer Hunyuan3D-2.1 release but still publicly accessible
Text-to-3D, image-to-3D, sketch-to-3D, rapid game and e-commerce asset creation, 3D printing, and production-oriented asset prototyping.
Current and accessible through Tencent Cloud APIs; newer HY-3D-3.1 is also available as a separate model version.
Local English- and Chinese-language text-to-image generation, creative prototyping, ComfyUI workflows, LoRA customization, and ControlNet-based image conditioning
Open-weight and downloadable; no official deprecation or shutdown notice found
Text-to-image generation, reference-guided image creation, marketing graphics, e-commerce content, visual ideation, and image-generation applications requiring custom aspect ratios
Current and accessible through Tencent Cloud TokenHub; also available as an open-weight HunyuanImage-3.0 release
Multilingual OCR, document parsing, text spotting, table and formula extraction, structured information extraction, and local visual-document processing
Current open-weight release
Video description, video question answering, video summarization, content review, scene analysis, and video metadata generation.
Online and currently listed in Tencent Cloud TokenHub; the older Hunyuan platform entry was retired on 2026-06-22, while the TokenHub model remains available.
Research and production experimentation with high-quality local text-to-video generation
Open-source and currently accessible; original HunyuanVideo model, with HunyuanVideo-1.5 released later as a lighter successor
Local text-to-video and image-to-video generation, open-source video research, creative prototyping, and developers needing a comparatively lightweight high-quality video model.
Current and publicly accessible open-weight model
Image-to-video generation, reference-image animation, visual effects, creative video prototyping, and locally hosted open-weight video workflows
Available as an open-weight model and official open-source repository; no official shutdown date found
Image-grounded reasoning, OCR, chart and document analysis, visual localization, educational problem solving, and multilingual visual question answering
Online and currently available through Tencent Cloud TokenHub
Coding agents, long-context analysis, complex reasoning, productivity automation, structured workflows, and multi-step tool use
Current; open-weight and available through Tencent Cloud TokenHub
Automated decomposition of FBX 3D models into separate model components.
Current and API-accessible
Fast text-to-3D and image-to-3D asset generation, prototyping, and automated 3D content pipelines
Current and available through Tencent Cloud TokenHub
Automated conversion of existing 3D assets between common interchange and delivery formats
Current TokenHub 3D format-conversion service
Generating short human character animations from natural-language action descriptions
Current and accessible through Tencent TokenHub
Automated retopology, polygon reduction, and preparation of existing 3D meshes for games, rendering, animation, and downstream asset workflows
Current and API-accessible
Automated rigging and skinning of human or animal 3D characters for animation, games, virtual characters, and asset prototyping
Current and accessible through Tencent Cloud TokenHub
Reference-guided texturing of existing OBJ or GLB meshes, including PBR material generation and automated asset preparation
Current and available through Tencent Cloud TokenHub API
Automated UV unwrapping and preparation of 3D assets for texturing, rendering, game development, and digital-content workflows.
Current and accessible through Tencent Cloud TokenHub
Real-time and asynchronous Mandarin, English, mixed-language, and Chinese-dialect transcription, captions, subtitles, and short voice-command recognition
Preview; internal-test availability
High-resolution text-to-image generation, reference-image creation, multi-turn image editing, posters, UI concepts, marketing assets, and product-image workflows
Current preview model
Low-latency multilingual translation, localized content, structured translation instructions, and cost-sensitive production workflows
Current and available through Tencent Cloud TokenHub
Professional multilingual translation, localization, terminology-sensitive workflows, and context-aware business translation
Current and available through Tencent Cloud TokenHub
Professional, domain-specific and high-quality multilingual translation with contextual disambiguation and instruction following
Online and currently available through Tencent Cloud TokenHub
Chinese role-play, character simulation, fictional dialogue, AI avatars, and emotionally oriented conversational experiences
Current and available through Tencent Cloud TokenHub
Image understanding, OCR, chart and diagram analysis, STEM visual reasoning, visual question answering, and multi-image comparison
Current and available through Tencent Cloud TokenHub
Generating explorable 3D environments, Gaussian-splat scenes, point clouds, and collision meshes from text or reference images
Current and accessible through Tencent Cloud TokenHub
Text-to-3D, image-to-3D, multi-view reconstruction, game assets, product visualization, digital-human content, e-commerce assets and 3D printing workflows
Current and available through Tencent Cloud HY-3D APIs and TokenHub
Long-context coding agents, complex tool-use workflows, productivity automation, document analysis, game development, and scientific reasoning
Preview; currently available as an open-weight model and through Tencent products, Tencent Cloud TokenHub, and OpenRouter
Text-to-panorama and image-to-panorama generation for immersive environments, virtual tours, games, simulations, visualization, and 3D-world pipelines.
Current and accessible through Tencent Cloud TokenHub
High-volume semantic retrieval, vector search, FAQ matching, text clustering, classification, and cost- or latency-sensitive knowledge-base applications
Current and available through Tencent Cloud TokenHub
High-quality multilingual semantic search, retrieval-augmented generation, enterprise knowledge bases, similarity matching, and text classification.
Current and available through Tencent Cloud TokenHub and the Tencent Cloud embeddings API.
Fast cross-modal image-text retrieval, multimodal semantic matching, and video search
Current and available through Tencent Cloud TokenHub
High-precision multimodal retrieval, cross-modal image-text search, video search, and semantic matching across text, image, and video collections
Current and available through Tencent Cloud TokenHub
Video, image, audio, and text understanding; multimedia summarization; content tagging; structural video analysis; object localization; enterprise media workflows
Currently accessible; scheduled for shutdown on October 15, 2026 at 00:00 Beijing time on Tencent Cloud TokenHub and ADP
Multilingual video translation, voice-preserving dubbing, subtitle translation, online courses, films, and short-form video localization.
Available
Multilingual video localization, translated online courses, short-form video dubbing, film and media localization, and long-video voice-preserving translation
Current and available through Tencent Cloud TokenHub as an asynchronous AI dubbing model
Fast everyday image creation, social-media graphics, marketing materials, short-video covers, and reference-guided image generation
Current and available through Tencent Cloud TokenHub
Low-cost, high-volume text-to-image and reference-to-image generation, including e-commerce product imagery, batch visual assets, and marketing content
Currently available managed image-generation model
Professional text-to-image and reference-to-image generation, brand visuals, refined product imagery, high-quality design assets, and 1K-to-4K creative production.
Current and available through Tencent Cloud TokenHub
Fast, cost-conscious generation of short e-commerce, advertising, social media, and reference-guided video assets
Current and accessible through Tencent Cloud TokenHub
Audio and video transcription, subtitle generation, sentence-level timestamps, and short-form speech-to-text workflows
Current and available through Tencent Cloud TokenHub
Multi-view and video-to-3D reconstruction, camera and depth estimation, point-map prediction, surface-normal estimation, novel-view synthesis, and 3D Gaussian Splatting workflows
Legacy but publicly available open-weight model; superseded by WorldMirror-2.0 in Tencent's HY-World 2.0 catalog
Long-context analysis, enterprise agents, research, coding assistance, structured extraction and tool-enabled workflows
current
Software engineering, codebase analysis, technical reasoning, long-context work, tool-using agents, document analysis, and structured workflow automation
current
Advanced software engineering, long-context reasoning, agentic tool use, research with web or X search, and professional knowledge work.
Current; available through the xAI API, Grok Build, Cursor, and selected model gateways.
Agentic coding, long-context software engineering, research, knowledge work, visual analysis, structured extraction, and tool-using workflows.
Current and available; superseded as xAI's flagship by Grok 4.7 but still supported on the xAI API.
Low-latency coding, interactive development, and agentic workflows in Cursor or Grok Build
Available as a faster-serving variant in Cursor and Grok Build; not available as a separate public xAI API model
Deep research, parallel investigation, long-context analysis, and tool-assisted synthesis
Beta; currently available through the xAI API
Fast general-purpose text generation, image-aware analysis, coding assistance, structured extraction, tool-calling agents and large-context workflows
current
Complex reasoning, coding, technical research, long-context document analysis, image understanding, structured responses, and tool-enabled agentic workflows
current
Agentic coding, web development, debugging, software engineering workflows, MCP integrations, and fast tool-calling applications.
Current; public beta
API-based text-to-image generation, image editing, visual prototyping, creative applications, and per-image billing
Current and available through the xAI API as of September 24, 2026
High-quality text-to-image generation, image editing, multi-reference compositing, marketing graphics, product imagery, typography, and design assets
Generally available
High-quality image generation and editing, realistic product and marketing imagery, detailed scenes, creative assets, and images requiring stronger text rendering or prompt adherence.
Currently available; scheduled for API retirement on November 2, 2026. After retirement, requests using the slug are served by Grok Imagine Image 2.0 with quality set to low.
Short-form text-to-video, image-to-video, reference-guided video, cinematic prototyping, marketing clips, and audiovisual creative workflows
Generally available
Low-cost text-to-video and image-to-video drafts, social clips, rapid creative iteration, and high-volume generation
Current and available through the xAI Imagine API
Realtime voice agents, customer support, telephony, sales, multilingual conversations, and tool-enabled spoken workflows
Current and available through the xAI Speech to Speech API
Batch and real-time multilingual speech transcription, dictation, voice assistants, accessibility, meetings, and customer-support audio
Currently accessible; original speech-to-text model; scheduled for deprecation in the coming weeks
Batch and real-time transcription of multilingual, noisy, conversational, telephony, and voice-agent audio
Current
Few-shot audio-language research, speech continuation, voice and style conversion, speech translation, speech editing, and audio-text experimentation
Current open-weight downloadable model
Local audio understanding, speech-to-text dialogue, spoken conversational agents, and controllable text-to-speech research
Current open-weight model; locally downloadable and usable
Embodied AI research, autonomous-driving perception and planning, spatial reasoning, affordance prediction, robot navigation, and multimodal visual analysis
Active open-weight model
Research, custom post-training, foundation-model experimentation, and self-hosted long-context inference
Current open-weight base model
Speech-to-text transcription for Chinese, English, regional Chinese dialects, code-switched speech, lyrics, noisy recordings, and multi-speaker conversations
Current and accessible through the Xiaomi MiMo API; open-source code and weights available
Expressive text-to-speech, audiobooks, podcasts, dubbing, character dialogue, voice interfaces, narrated content, and stylized speech or singing.
Available
Zero-shot voice cloning, expressive narration, character dialogue, personalized speech, and custom-voice audio production
Current; limited-time free access
Custom synthetic voices for narration, characters, podcasts, ASMR, games, assistants, and creative audio production
Current; limited-time free access
Local coding agents, terminal automation, cybersecurity experimentation, visual coding and agentic reinforcement-learning research
Current open-weight SFT checkpoint
High-volume multimodal API workloads, coding assistants, agent automation, long-context document and repository analysis, tool-using workflows, and cost-sensitive professional applications.
Current; open-weight and available through the Xiaomi MiMo API, MiMo Studio, MiMo Desktop, MiMo Code, and other supported integrations.
Long-horizon agents, coding, cybersecurity, research, computer use, multimodal analysis, and complex multi-step workflows
Current; open-source weights and API access available
Latency-sensitive multimodal reasoning, coding, tool-use, long-context research, and interactive agent workflows
Current; high-speed inference variant of MiMo-V2.6-Pro
Local image and video understanding, multimodal reasoning, visual question answering, OCR-oriented tasks, visual grounding, and GUI analysis
Current open-weight model; publicly available for download and local deployment
Self-hosted image and video understanding, multimodal reasoning research, supervised fine-tuning, and reinforcement-learning experimentation
Current open-weight checkpoint; available for download
Long-horizon agent workflows, repository-scale coding, complex software engineering, tool-driven automation, and very large documents
Deprecated; API model name scheduled to expire on 2026-10-21 at 10:00 Beijing Time
Russian-language conversational assistants, long-context dialogue, RAG, document analysis, information extraction, reporting, and complex business text generation
Current and available in Yandex AI Studio
High-volume business text processing, customer-support dialogs, moderation, classification, summarization, document extraction, and knowledge-base search.
Current and available in Yandex AI Studio
Text-to-image generation, Russian-language visual content, illustrations, graphic design, advertising assets, presentations, landing pages, and product cards
Current
Russian-language enterprise text generation, document analysis, RAG, structured extraction, rewriting, classification, reporting, and tool-augmented business assistants
Current and available through Yandex Cloud AI Studio; explicit model URI recommended. The RC alias remains available until the end of support.
Fast, cost-efficient Russian-language text generation and everyday conversational tasks
Unverified
Short-form text-to-video, image animation, start-and-end-frame transitions, advertising, marketing, realistic scenes, and 3D-style video generation
Current and available through the Z.ai video-generation API
Long-horizon software engineering, repository-scale coding, code agents, complex refactoring, large-document analysis, research implementation, and tool-using workflows.
Active; still listed in Z.AI API pricing and available as downloadable open weights, although GLM-5.3 is the newer flagship successor.
Complex software engineering, large codebases, long-horizon coding agents, terminal workflows, technical research, and authorized cybersecurity analysis
Current and available
Multimodal coding agents, screenshot and interface understanding, long-context document work, visual verification, tool-assisted research, and cost-sensitive agent workflows
Current and available through the Z.ai API, GLM Coding Plan, and publicly released open weights
Hosted multilingual speech-to-text, short real-time transcriptions, captions, meeting snippets, voice input, customer-service audio, and terminology-aware transcription
Current and publicly available through the Z.AI audio transcription API
Text-heavy posters, presentation graphics, educational diagrams, social-media layouts, image editing, and open-weight image-generation research
Current; available through the Z.ai API and as downloadable open weights
High-volume OCR, PDF parsing, table and formula recognition, handwriting, receipts, forms, information extraction, and local document-processing pipelines
Current; open-weight; available through hosted API and self-hosted deployment