≡
Models by output type

Text Output in AI Models: Meaning, Uses, Limits, and Evaluation

Text output is one of the most common capabilities in AI models: the system returns words, symbols, code, or other text-encoded content in response to an instruction or input. That output might be a conversational answer, summary, translation, classification, document draft, JSON object, or function-call request. The category describes what the model produces, not necessarily what it accepts. This distinction matters because a model can analyze an image, audio recording, video, or document and still produce only text. This guide explains what text-output models return, how they differ from speech, image, video, embedding, and tool-use systems, where they are useful, and what to evaluate before choosing one.
What this means

Text output means the model can generate language or other text-based responses. It does not necessarily imply support for images, audio or other output formats.

Text models

720 models currently match this capability.

View all models →
◎
AI21

AI21-Jamba-Large-1.6

Jamba 1.6

Long-context retrieval-augmented generation, enterprise document analysis, grounded question answering, structured extraction, classification, and private deployment

Text General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
AI21 Labs

AI21-Jamba-Mini-1.5

Jamba 1.5

Long-context document analysis, RAG, structured generation, function calling, multilingual enterprise assistants, and private deployment

Text General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
AI21

AI21-Jamba-Mini-1.6

Jamba 1.6

Long-context RAG, grounded question answering, enterprise document processing, classification, structured text generation, function calling and privacy-sensitive private deployments.

Text General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
AI21

AI21-Jamba-Mini-1.7

Jamba 1.7

Long-document analysis, enterprise RAG, grounded question answering, structured text generation, private deployment, and cost-sensitive text workflows

Text General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
AI21 Labs

AI21-Jamba-Reasoning-3B

Jamba Reasoning

Local reasoning, long-context document analysis, private RAG, extraction, coding assistance, and lightweight agent controllers

Text Reasoning 256,000 ctx Tool use Streaming
View model →
◎
AI21

AI21-Jamba2-3B

Jamba2

Long-context RAG, grounded enterprise question answering, document extraction, local inference, on-device assistants, and lightweight agent workflows

Text Lightweight 262,144 ctx Tool use Streaming
View model →
◎
AI21 Labs

AI21-Jamba2-Mini

Jamba2

Long-context enterprise question answering, grounded generation, document analysis, instruction-heavy workflows, RAG systems and self-hosted deployments

Text General Purpose 256,000 ctx Streaming
View model →
◎
Aleph Alpha

Aleph-Alpha-GermanWeb-Grammar-Classifier-BERT

GermanWeb Grammar Classifier

German grammar-related text-quality filtering, dataset curation, and binary classification of web text

Text Other 512 ctx
View model →
◎
Aleph Alpha

Aleph-Alpha-GermanWeb-Grammar-Classifier-fastText

Aleph-Alpha GermanWeb

Fast local classification of German documents for grammar-oriented corpus filtering and data curation

Text Other
View model →
◎
Aleph Alpha

Aleph-Alpha-GermanWeb-Quality-Classifier-BERT

GermanWeb Quality Classifier

German web-document quality scoring, dataset filtering, ranking, and pretraining-data curation

Text Other 512 ctx
View model →
◎
Aleph Alpha

Aleph-Alpha-GermanWeb-Quality-Classifier-fastText

Aleph-Alpha-GermanWeb Quality Classifier

Fast German-language document quality classification and large-scale web-data filtering

Text Other
View model →
◎
Yandex

Alice AI LLM

Alice AI

Russian-language conversational assistants, long-context dialogue, RAG, document analysis, information extraction, reporting, and complex business text generation

Text General Purpose 131,072 ctx Streaming
View model →
◎
Yandex

Alice AI LLM Flash

Alice AI

High-volume business text processing, customer-support dialogs, moderation, classification, summarization, document extraction, and knowledge-base search.

Text Lightweight 65,536 ctx
View model →
◎
NVIDIA

Alpamayo 1.5 Nano

Alpamayo

Autonomous-driving research, trajectory prediction, interpretable motion planning, navigation-conditioned driving, visual question answering and safety-oriented model evaluation

Text Reasoning Image input Video input
View model →
◎
Amazon

Amazon Nova 2 Lite

Amazon Nova 2

High-volume multimodal applications, document and video analysis, customer service, business automation, software engineering, long-context workflows, and cost-sensitive AI agents.

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Amazon

Amazon Nova 2 Sonic

Amazon Nova 2

Real-time voice assistants, customer-service automation, telephony, interactive learning, multilingual conversations, and tool-enabled speech agents.

Text Multimodal 1,000,000 ctx Audio input Tool use Structured output
View model →
◎
Amazon

Amazon Nova Lite

Amazon Nova

Low-cost multimodal document analysis, image and video understanding, visual question answering, summarization, RAG, and tool-enabled agents.

Text Multimodal 300,000 ctx Image input Video input Tool use
View model →
◎
Amazon

Amazon Nova Micro

Amazon Nova

High-volume, low-latency text classification, summarization, translation, extraction, routing, FAQs and narrowly defined fine-tuned tasks

Text Lightweight 128,000 ctx Tool use Streaming
View model →
◎
Amazon

Amazon Nova Premier

Amazon Nova

Long-context multimodal analysis, enterprise document workflows, complex tool calling, agentic orchestration, codebase analysis, and teacher-model distillation before retirement.

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Amazon

Amazon Nova Pro

Amazon Nova

Enterprise multimodal applications, document analysis, visual question answering, video understanding, long-context summarization, RAG, and tool-using assistants

Text Multimodal 300,000 ctx Image input Video input Tool use
View model →
◎
Amazon

Amazon Nova Sonic

Amazon Nova

Real-time voice assistants, customer-service automation, interactive education, language learning, and speech-enabled enterprise workflows

Text Multimodal 300,000 ctx Audio input Tool use Streaming
View model →
◎
MiniMax

ASR 1.0

MiniMax ASR

Multilingual audio transcription, meeting transcription, speaker-labeled transcripts, live captions, call analysis, and subtitle generation

Text Other Audio input Streaming
View model →
◎
Cohere

Aya Expanse 32B

Aya Expanse

Multilingual generation, translation, summarization, customer support, global communication, and multilingual research

Text General Purpose 128,000 ctx
View model →
◎
Cohere Labs

Aya Vision 32B

Aya Vision

Multilingual image understanding, OCR, image captioning, visual question answering, image-based translation, visual reasoning and research deployments using open weights.

Text Multimodal 16,000 ctx Image input Streaming
View model →
◎
OpenAI

babbage-002

GPT-3

Maintaining legacy text-completion applications, historical GPT-3 base-model behavior, and existing compatible fine-tuned workflows before shutdown

Text Lightweight
View model →
◎
NVIDIA

Canary-1B

Canary

Multilingual offline speech transcription, speech translation, and NeMo-based ASR research or deployment

Text Other Audio input
View model →
◎
Cerebras

Cerebras-GPT-1.3B

Cerebras-GPT

Research, local text generation, fine-tuning experiments, benchmarking, and lightweight English NLP prototypes

Text General Purpose 2,048 ctx
View model →
◎
Cerebras Systems

Cerebras-GPT-111M

Cerebras-GPT

Research, small-scale language-model experimentation, causal text generation, fine-tuning studies, and local deployment

Text Lightweight 2,048 ctx
View model →
◎
Cerebras

Cerebras-GPT-13B

Cerebras-GPT

Open LLM research, local text generation, reproducible training experiments, and downstream fine-tuning

Text General Purpose 2,048 ctx
View model →
◎
Cerebras

Cerebras-GPT-2.7B

Cerebras-GPT

Open-weight language-model research, local text generation, fine-tuning experiments, scaling-law studies, and educational or reference implementations.

Text General Purpose 2,048 ctx
View model →
◎
Cerebras Systems

Cerebras-GPT-256M

Cerebras-GPT

Compact local language-model experiments, text completion, education, benchmarking, and fine-tuning research

Text Lightweight 2,048 ctx
View model →
◎
Cerebras

Cerebras-GPT-590M

Cerebras-GPT

Local text generation, language-model research, benchmarking, fine-tuning experiments, and studying compute-efficient scaling at small model sizes

Text Lightweight 2,048 ctx
View model →
◎
Cerebras

Cerebras-GPT-6.7B

Cerebras-GPT

Open LLM research, scaling-law experiments, language-model evaluation, fine-tuning, and self-hosted English text generation

Text General Purpose 2,048 ctx
View model →
◎
OpenAI

Chat Latest

Chat Latest

ChatGPT-style instant responses, general-purpose writing and analysis, image-aware conversations, and tool-assisted workflows

Text General Purpose 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

ChatGPT-4o

GPT-4o

Fast general-purpose conversations, vision, voice interactions, coding, and everyday productivity

Text Multimodal 128,000 ctx Image input Audio input Video input
View model →
◎
OpenAI

chatgpt-image-latest

GPT Image

Existing ChatGPT image-generation and image-editing integrations

Text Multimodal Image input
View model →
◎
Anthropic

Claude Fable 5

Claude Fable 5

Complex reasoning, long-running autonomous agents, advanced coding, multi-stage research, document-heavy analysis, and high-value enterprise workflows

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Fable 5.1

Claude Fable 5

Long-running agentic coding, demanding reasoning, multistep research, complex document analysis, spreadsheets, presentations, and high-stakes knowledge work

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Haiku 4.5

Claude 4.5

Low-latency assistants, customer support, high-volume text processing, coding assistance, image understanding, and parallel subagent workloads

Text Lightweight 200,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Mythos 5

Claude Mythos

Advanced cybersecurity, vulnerability research, biology, healthcare, life-sciences research, and long-running technical workflows requiring extensive context and reasoning

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Mythos 5.1

Claude Mythos

Vetted cybersecurity defense, advanced biology and life-sciences research, complex coding, long-running agents, technical investigation, and high-value research workflows

Text Reasoning 1,000,000 ctx Image input Tool use Structured output
View model →
◎
Anthropic

Claude Mythos Preview

Claude Mythos

Defensive cybersecurity research, vulnerability discovery, attack-surface auditing, autonomous coding, long-running agents, and large-context analysis.

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 4.5

Claude Opus 4

Complex software engineering, coding agents, multi-step research, computer-use workflows, enterprise analysis, visual document understanding, and high-value tool-using applications.

Text Reasoning 200,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 4.6

Claude Opus

Complex reasoning, agentic coding, repository-scale software work, long-context analysis, research, document and spreadsheet workflows, and enterprise automation

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 4.7

Claude Opus 4

Complex reasoning, agentic coding, repository-scale software engineering, research, document analysis, computer-use workflows, and high-stakes knowledge work

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 4.8

Claude Opus 4

Complex reasoning, agentic coding, long-horizon workflows, enterprise knowledge work, document analysis, vision tasks, and computer-use agents

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 5

Claude Opus

Complex agentic coding, code review, enterprise analysis, long-context research, document workflows, financial and legal reasoning, and multi-step tool-using applications.

Text General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 5.5

Claude Opus

Long-running agentic coding, repository-scale software engineering, code review, knowledge work, multimodal document analysis, and tool-using applications

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Sonnet 4.5

Claude Sonnet

Complex coding, software agents, computer-use workflows, visual document analysis, research, and long-running multi-step tasks

Text Coding 200,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Sonnet 4.6

Claude Sonnet 4.6

Agentic coding, computer use, long-context document analysis, enterprise knowledge work, structured extraction, and high-volume applications needing strong reasoning at Sonnet-tier pricing

Text General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Sonnet 5

Claude Sonnet

Agentic coding, software engineering, browser and computer-use workflows, long-context analysis, tool-driven automation, and high-volume assistants

Text General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Sonnet 5.5

Claude Sonnet 5.5

Fast coding agents, long-context analysis, image-aware workflows, document creation, and tool-enabled business applications

Text General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
Mistral AI

Codestral 25.08

Codestral

Low-latency IDE autocomplete, fill-in-the-middle completion, code generation, code editing, test generation and developer assistants

Text Coding 128,000 ctx Tool use Structured output Streaming
View model →
◎
OpenAI

codex-mini-latest

Codex

Codex CLI coding workflows, code question answering, code editing, repository tasks, and low-latency software-engineering assistance

Text Coding 200,000 ctx Image input Tool use Structured output
View model →
◎
Cohere

Cohere Transcribe

Cohere Transcribe

Multilingual audio transcription, enterprise speech archives, meeting and interview transcription, call-center audio, and low-latency ASR workflows

Text Speech Recognition Audio input
View model →
◎
Cohere

Cohere Transcribe Arabic

Cohere Transcribe

Arabic speech transcription, multilingual Arabic-English audio, regional dialects, code-switched speech, call-center audio and high-throughput ASR workloads.

Text Other Audio input
View model →
◎
Cohere

Command A

Command A

Enterprise RAG, long-context document analysis, multilingual applications, tool use, agentic workflows, financial text processing and structured text generation

Text General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
Cohere

Command A Reasoning

Command A

Complex enterprise agents, tool use, retrieval-augmented generation, multilingual reasoning, long-context analysis, and workflow automation.

Text Reasoning 256,000 ctx Tool use Web search Structured output
View model →
◎
Cohere

Command A Translate

Command A

High-quality multilingual text translation, enterprise document translation, and privacy-sensitive translation workflows.

Text Specialized 8,000 ctx Tool use Structured output Streaming
View model →
◎
Cohere

Command A Vision

Command A

Enterprise document intelligence, OCR, chart and table analysis, visual question answering, and multilingual image understanding

Text Multimodal 128,000 ctx Image input Structured output Streaming
View model →
◎
Cohere

Command A+

Command A

Enterprise agents, multimodal document and image analysis, multilingual workflows, reasoning-intensive automation, retrieval-augmented generation, and tool-using applications.

Text Multimodal 128,000 ctx Image input Tool use Structured output
View model →
◎
Cohere

Command R 08-2024

Command R

Low-cost enterprise RAG, multilingual document workflows, long-context chat, structured extraction, and tool-using agents

Text General Purpose 128,000 ctx Tool use Structured output Streaming
View model →
◎
Cohere

Command R+ 08-2024

Command R+

Complex enterprise RAG, long-context document analysis, multilingual assistants, citations, structured data tasks and multi-step tool-use agents.

Text General Purpose 128,000 ctx Tool use Structured output Streaming
View model →
◎
Cohere

Command R7B

Command R

Cost-sensitive RAG, enterprise chat, tool use, coding assistance, and fast multi-step agents

Text Lightweight 128,000 ctx Tool use Web search Structured output
View model →
◎
OpenAI

computer-use-preview

Computer-Using Agent

Controlled browser automation, computer-use research, UI testing, and repetitive interface workflows

Text Other 8,192 ctx Image input Tool use
View model →
◎
NVIDIA

Cosmos3-Edge

Cosmos 3

Edge physical AI, robotics, visual reasoning, world simulation, video generation, and action-policy prototyping

Text Multimodal 131,072 ctx Image input Video input
View model →
◎
NVIDIA

Cosmos3-Nano

Cosmos 3

Physical AI, robotics, autonomous-vehicle simulation, multimodal world generation, future-state prediction, action reasoning, and synthetic training data

Text Multimodal Image input Audio input Video input
View model →
◎
NVIDIA

Cosmos3-Super

Cosmos 3

High-quality Physical AI simulation, synthetic-data generation, robotics and autonomous-vehicle research, multimodal world modeling, and teacher-model distillation

Text Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Microsoft

CxrReportGen Premium

CxrReportGen

First-pass chest X-ray report drafting, structured findings extraction, radiologist workflow assistance, research evaluation, and institution-specific fine-tuning

Text Multimodal Image input
View model →
◎
OpenAI

davinci-002

GPT-3

Legacy text completion, code continuation, and inference from existing davinci-002 fine-tuned models before shutdown

Text General Purpose
View model →
◎
Databricks

DBRX Base

DBRX

Self-hosted English text completion, code completion, model research, and fine-tuning experiments using an open-weight mixture-of-experts model.

Text General Purpose 32,768 ctx Streaming
View model →
◎
Databricks

DBRX Instruct

DBRX

Self-hosted general English-language instruction following, coding assistance, text generation, and domain-specific fine-tuning

Text General Purpose 32,768 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-Coder-V2

DeepSeek-Coder-V2

Self-hosted code generation, completion, repository analysis, debugging, code translation, and programming-focused research.

Text Coding 128,000 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-Coder-V2-Lite-Base

DeepSeek-Coder-V2

Local code completion, fill-in-the-middle generation, IDE integrations, repository-level coding, and private or self-hosted inference

Text Coding 131,072 ctx
View model →
◎
DeepSeek

DeepSeek-Coder-V2-Lite-Instruct

DeepSeek-Coder-V2

Local code generation, code completion, debugging, code explanation, repository-scale prompts, and developers needing an open-weight coding model

Text Coding 128,000 ctx
View model →
◎
DeepSeek

DeepSeek-LLM 67B

DeepSeek-LLM

Local deployment, bilingual English-Chinese generation, language-model research, mathematics, coding experiments, and custom fine-tuning workflows

Text General Purpose 4,096 ctx
View model →
◎
DeepSeek

DeepSeek-LLM 7B

DeepSeek-LLM

Local bilingual text generation, research, experimentation, instruction tuning, and lightweight self-hosted assistants

Text General Purpose 4,096 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-Math-V2

DeepSeek-Math

Advanced mathematical reasoning, natural-language theorem proving, proof generation, proof verification, and research on self-correcting reasoning systems

Text Reasoning 163,840 ctx
View model →
◎
DeepSeek

DeepSeek-OCR

DeepSeek-OCR

Local OCR, document digitization, PDF and image parsing, layout-aware markdown conversion, table extraction, figure parsing, and visual-text compression research

Text Multimodal 8,192 ctx Image input Streaming
View model →
◎
DeepSeek

DeepSeek-OCR 2

DeepSeek-OCR

Local OCR, scanned-document transcription, layout-aware document parsing, table extraction, and document-to-Markdown workflows

Text Multimodal 8,192 ctx Image input Streaming
View model →
◎
DeepSeek

DeepSeek-Prover-V1

DeepSeek-Prover

Lean 4 theorem proving, formal mathematics, automated proof generation, and theorem-proving research

Text Reasoning
View model →
◎
DeepSeek

DeepSeek-Prover-V1.5-Base

DeepSeek-Prover V1.5

Lean 4 theorem proving, formal mathematics research, proof completion, proof-search experiments, and open-weight model fine-tuning

Text Reasoning 4,096 ctx
View model →
◎
DeepSeek

DeepSeek-Prover-V1.5-SFT

DeepSeek-Prover

Lean 4 proof completion, formal mathematics research, automated theorem proving, and verifier-guided proof search

Text Reasoning 4,096 ctx
View model →
◎
DeepSeek

DeepSeek-Prover-V2-671B

DeepSeek-Prover

Lean 4 theorem proving, formal mathematics, automated proof synthesis, proof-search research, and verifier-guided reasoning

Text Reasoning 163,840 ctx
View model →
◎
DeepSeek

DeepSeek-R1

DeepSeek-R1

Mathematical reasoning, coding, technical analysis, research, and self-hosted reasoning applications

Text Reasoning 128,000 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-R1-Distill-Llama-70B

DeepSeek-R1

Self-hosted mathematics, coding, research, complex reasoning, and long-form analysis

Text Reasoning 131,072 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-R1-Distill-Qwen-1.5B

DeepSeek-R1-Distill

Local mathematical reasoning, compact reasoning experiments, educational applications, lightweight coding assistance, and self-hosted inference on limited hardware

Text Reasoning 131,072 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-R1-Zero

DeepSeek-R1

Reasoning research, mathematics, coding experiments, reinforcement-learning studies, open-weight evaluation, and model distillation

Text Reasoning 131,072 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-V2

DeepSeek-V2

Local or third-party deployment for general text generation, translation, mathematics, research, and code generation

Text General Purpose 128,000 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-V2-Lite

DeepSeek-V2

Local text generation, Chinese and English language tasks, MoE research, fine-tuning, and efficient self-hosted inference

Text Lightweight 32,768 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-V2.5

DeepSeek-V2

Open-weight general language generation, coding assistance, code completion, and self-hosted experimentation

Text General Purpose 128,000 ctx Tool use Streaming
View model →
◎
DeepSeek

DeepSeek-V3

DeepSeek-V3

Open-weight general language generation, coding, long-context text tasks, research, and cost-sensitive third-party inference.

Text General Purpose 128,000 ctx Tool use Streaming
View model →
◎
DeepSeek

DeepSeek-V3.1

DeepSeek-V3

Open-weight reasoning, coding, tool-calling, long-context analysis, and self-hosted agent systems

Text General Purpose 128,000 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V3.1-Terminus

DeepSeek-V3.1

Open-weight deployment, coding assistance, long-context text processing, reasoning workflows, search agents, and terminal-oriented automation

Text General Purpose 128,000 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V3.2

DeepSeek V3

Open-weight reasoning, coding, long-context analysis, tool-using agents, research workflows, and cost-sensitive deployments

Text Reasoning 131,072 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V3.2-Speciale

DeepSeek-V3.2

Difficult mathematics, advanced coding, scientific reasoning, long-form analysis, benchmark evaluation, and research deployment

Text Reasoning 163,840 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-V4-Flash

DeepSeek V4

Long-context reasoning, coding, agent workflows, and cost-sensitive API applications

Text Lightweight 1,000,000 ctx Tool use Streaming
View model →
◎
DeepSeek

DeepSeek-V4-Flash-Base

DeepSeek V4

Self-hosted language-model research, custom post-training, domain adaptation, and large-context text generation

Text General Purpose 1,048,576 ctx
View model →
◎
DeepSeek

DeepSeek-V4-Flash-Vision-Exp

DeepSeek V4

Image understanding, screenshot and chart analysis, multimodal coding agents, visual tool-use workflows, and text-plus-image reasoning

Text Multimodal Image input Tool use
View model →
◎
DeepSeek

DeepSeek-V4-Pro

DeepSeek V4

Complex reasoning, coding agents, long-context analysis, tool-using workflows, and large document or codebase processing

Text General Purpose 1,000,000 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V4-Pro-Base

DeepSeek V4

Research, continued pretraining, fine-tuning, foundation-model evaluation, and custom large-scale inference

Text General Purpose 1,048,576 ctx
View model →
◎
DeepSeek

DeepSeek-V4.1-Flash

DeepSeek V4.1

Low-cost, high-throughput reasoning and coding, long-context analysis, agentic workflows, tool calling, and text-plus-image understanding.

Text Lightweight 1,048,576 ctx Image input Tool use Structured output
View model →
◎
DeepSeek

DeepSeek-VL-1.3B-Base

DeepSeek-VL

Local image understanding, visual question answering, diagram and document analysis, multimodal research, and compact deployments

Text Multimodal 4,096 ctx Image input
View model →
◎
Databricks

Dolly v2 12B

Dolly v2

Local experimentation, research, instruction-tuning studies, and organizations needing an openly downloadable model with commercial-use licensing

Text General Purpose 2,048 ctx
View model →
◎
Baidu

ERNIE 4.5 Turbo

ERNIE 4.5

Fast, low-cost Chinese text generation, long-context applications, enterprise agents, content creation, reasoning, and code assistance.

Text General Purpose 138,240 ctx Tool use Web search
View model →
◎
Baidu

ERNIE 4.5 Turbo VL

ERNIE 4.5 Turbo

Cost-efficient multimodal understanding, image and video analysis, OCR, document comprehension, translation, visual question answering, and code-related tasks

Text Multimodal 128,000 ctx Image input Video input Tool use
View model →
◎
Baidu

ERNIE 5.0

ERNIE

Multimodal understanding, long-context Chinese and English applications, complex reasoning, coding, tool-enabled agents, and enterprise workloads

Text Multimodal 248,832 ctx Image input Audio input Video input
View model →
◎
Baidu

ERNIE 5.1

ERNIE 5

Agentic workflows, web-search-assisted tasks, reasoning, Chinese-language knowledge work, creative writing, and general-purpose assistant applications.

Text General Purpose Tool use Web search
View model →
◎
Baidu

ERNIE X1 Turbo

ERNIE X1

Chinese-language reasoning, long-form analysis, complex calculations, literary and document writing, agent workflows, and function calling

Text Reasoning 32,768 ctx Tool use Web search Structured output
View model →
◎
Baidu

ERNIE X1.1

ERNIE X1

Deep reasoning, Chinese and English question answering, mathematics, coding, factual responses, tool calling, web-grounded applications, and agent workflows

Text Reasoning 65,536 ctx Tool use Web search Streaming
View model →
◎
Baidu

ERNIE-4.5-0.3B

ERNIE 4.5

Lightweight local text generation, Chinese and English language experimentation, compact conversational systems, and domain adaptation

Text Lightweight 131,072 ctx Streaming
View model →
◎
Baidu

ERNIE-4.5-VL-28B-A3B

ERNIE 4.5 VL

Local or self-hosted multimodal applications, visual question answering, document and chart understanding, image analysis, video understanding, and efficient vision-language inference

Text Multimodal 131,072 ctx Image input Video input Streaming
View model →
◎
Baidu

ERNIE-Code3-128K

ERNIE Code

Code completion, code generation, unit-test generation, code optimization, code explanation, and fine-tuning for software-development workflows.

Text Coding 128,000 ctx
View model →
◎
LG AI Research

EXAONE-3.0-7.8B-Instruct

EXAONE 3.0

Bilingual English-Korean chat, instruction following, Korean-language applications, local inference, and open-weight LLM research

Text General Purpose 4,096 ctx Streaming
View model →
◎
LG AI Research

EXAONE-3.5-2.4B-Instruct

EXAONE 3.5

Lightweight local text generation, English-Korean assistants, summarization, rewriting, classification, and deployment on resource-constrained devices

Text Lightweight 32,768 ctx
View model →
◎
LG AI Research

EXAONE-3.5-32B-Instruct

EXAONE 3.5

Local or self-hosted English-Korean text generation, instruction following, long-context processing, coding, mathematics, and organizations requiring downloadable model weights.

Text General Purpose 32,768 ctx
View model →
◎
LG AI Research

EXAONE-3.5-7.8B-Instruct

EXAONE 3.5

Bilingual English-Korean assistants, long-document processing, local deployment, coding support, summarization, translation, and research

Text General Purpose 32,768 ctx
View model →
◎
LG AI Research

EXAONE-4.0-1.2B

EXAONE 4.0

On-device assistants, local text generation, Korean and multilingual conversational applications, lightweight reasoning, and resource-constrained deployments

Text Lightweight 65,536 ctx Tool use Streaming
View model →
◎
LG AI Research

EXAONE-4.0-32B

EXAONE 4.0

Self-hosted Korean, English, and Spanish language tasks requiring a combination of general-purpose generation, reasoning, coding, long-context processing, or agentic tool use

Text Reasoning 131,072 ctx Tool use Streaming
View model →
◎
LG AI Research

EXAONE-4.5-33B

EXAONE 4.5

Open-weight multimodal reasoning, Korean-language tasks, document understanding, OCR, visual question answering, and self-hosted deployments

Text Multimodal 262,144 ctx Image input Tool use Streaming
View model →
◎
LG AI Research

EXAONE-Deep-2.4B

EXAONE Deep

Local mathematics, coding, reasoning experiments, compact bilingual English-Korean applications, and resource-conscious deployment

Text Reasoning 32,768 ctx Streaming
View model →
◎
LG AI Research

EXAONE-Deep-32B

EXAONE Deep

Mathematical reasoning, scientific problem solving, coding evaluation, and local research deployments

Text Reasoning 32,768 ctx Streaming
View model →
◎
LG AI Research

EXAONE-Deep-7.8B

EXAONE Deep

Local research, mathematical reasoning, science problem solving, coding evaluation, Korean and English text generation, and self-hosted experimentation

Text Reasoning 32,768 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon Mamba 7B

Falcon Mamba

Local English text generation, open-weight research, long-sequence experimentation, and memory-conscious inference

Text General Purpose 8,192 ctx
View model →
◎
Technology Innovation Institute

Falcon OCR

Falcon Perception

Local or self-hosted document OCR, formula recognition, table extraction, receipts, invoices, papers, and layout-aware document parsing

Text Multimodal 16,384 ctx Image input Streaming
View model →
◎
Technology Innovation Institute

Falcon-180B

Falcon

Research, self-hosted text generation, language-model evaluation, and domain-specific fine-tuning when substantial GPU infrastructure is available

Text General Purpose 2,048 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-40B

Falcon

Research, self-hosted text generation, model fine-tuning, summarization, and general language-model experimentation

Text General Purpose 2,048 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-7B

Falcon

Self-hosted text generation, research, quantization, and domain-specific fine-tuning

Text General Purpose 2,048 ctx
View model →
◎
Technology Innovation Institute

Falcon-E-1B-Base

Falcon-E

Memory-efficient local text generation, edge-device experimentation, research, and downstream fine-tuning

Text Lightweight 32,768 ctx
View model →
◎
Technology Innovation Institute

Falcon-E-1B-Instruct

Falcon-E

Efficient local, edge, and resource-constrained English text generation; experimentation with BitNet models; full fine-tuning on the supplied prequantized revision.

Text Lightweight 32,768 ctx
View model →
◎
Technology Innovation Institute

Falcon-E-3B-Base

Falcon-E

Memory-efficient local text generation, edge deployment, BitNet research, continued pretraining, and custom fine-tuning

Text Lightweight 32,768 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-0.5B-Base

Falcon-H1

Lightweight local text generation, continued pretraining, domain adaptation, language-model research, and edge-oriented deployments

Text Lightweight 16,384 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-0.5B-Instruct

Falcon-H1

Lightweight local chat, text generation, edge deployment, prototyping, and fine-tuning experiments

Text Lightweight 16,384 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-1.5B-Base

Falcon-H1

Local multilingual text generation, long-context experimentation, research, and downstream fine-tuning

Text General Purpose 131,072 ctx
View model →
◎
Technology Innovation Institute

Falcon-H1-1.5B-Deep-Base

Falcon-H1

Efficient local text generation, multilingual language modeling, long-context applications, model adaptation, and compact reasoning research.

Text General Purpose 131,072 ctx
View model →
◎
Technology Innovation Institute

Falcon-H1-1.5B-Deep-Instruct

Falcon-H1

Efficient local inference, compact conversational assistants, multilingual text generation, long-context applications, and resource-constrained deployments

Text Lightweight 131,072 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-1.5B-Instruct

Falcon-H1

Efficient local or private deployment, multilingual instruction following, compact conversational applications, coding assistance, mathematics, and long-context text processing

Text Lightweight 131,072 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-34B-Base

Falcon-H1

Self-hosted multilingual text generation, long-context research, coding, RAG, foundation-model experimentation, and downstream fine-tuning

Text General Purpose 262,144 ctx
View model →
◎
Technology Innovation Institute

Falcon-H1-34B-Instruct

Falcon-H1

Long-context text generation, multilingual instruction following, document processing, retrieval-augmented generation, coding and self-hosted enterprise deployments

Text General Purpose 262,144 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-3B-Base

Falcon-H1

Local text generation, multilingual research, domain adaptation, fine-tuning, long-context experiments, and resource-conscious deployment.

Text General Purpose 131,072 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-3B-Instruct

Falcon-H1

Local multilingual chat, lightweight instruction following, retrieval-augmented generation, efficient text generation, and cost-sensitive self-hosted deployments

Text General Purpose 131,072 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-7B-Base

Falcon-H1

Private or local multilingual text generation, long-context experimentation, domain adaptation, and fine-tuning from a 7B-scale open-weight foundation

Text General Purpose 262,144 ctx
View model →
◎
Technology Innovation Institute

Falcon-H1-7B-Instruct

Falcon-H1

Self-hosted multilingual assistants, long-context text processing, coding, research, private deployments, and quantized local inference

Text General Purpose 262,144 ctx Tool use
View model →
◎
Technology Innovation Institute

Falcon-H1-Arabic

Falcon-H1-Arabic

Arabic NLP, dialect-aware assistants, long-context document analysis, summarization, multilingual reasoning, and locally deployed applications

Text General Purpose
View model →
◎
Technology Innovation Institute

Falcon-H1-Tiny-90M-Instruct

Falcon-H1-Tiny

Ultra-lightweight local text generation, edge deployment, offline instruction following, rewriting, extraction, and embedded AI experiments

Text Lightweight 262,144 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-Tiny-Coder-90M

Falcon-H1-Tiny

Lightweight local Python generation, fill-in-the-middle completion, edge deployment, code education, and low-resource developer tools

Text Coding 262,144 ctx
View model →
◎
Technology Innovation Institute

Falcon-H1-Tiny-Multilingual-100M-Instruct

Falcon-H1-Tiny

Lightweight local text generation, instruction-following, edge devices, embedded applications, experimentation, and privacy-sensitive deployments.

Text Lightweight 262,144 ctx
View model →
◎
Technology Innovation Institute

Falcon-H1-Tiny-R-0.6B

Falcon-H1-Tiny

Lightweight local reasoning, edge deployment, offline assistants, experimentation, and compact text-generation applications

Text Reasoning 262,144 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-Tiny-R-0.6B-pre-GRPO

Falcon-H1-Tiny-R

Local reasoning experiments, edge deployment, offline assistants, privacy-sensitive text processing, and resource-constrained inference

Text Reasoning 262,144 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-Tiny-R-90M

Falcon-H1-Tiny-R

Ultra-lightweight local reasoning, embedded applications, edge deployment, experimentation, and low-memory text generation

Text Reasoning 262,144 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1-Tiny-Tool-Calling-90M

Falcon-H1-Tiny

Lightweight local function calling, edge automation, API argument generation, and resource-constrained assistants

Text Lightweight 262,144 ctx Tool use Streaming
View model →
◎
Technology Innovation Institute

Falcon-H1R-7B

Falcon-H1R

Mathematical reasoning, programming, long-context analysis, local deployment, self-hosted inference, and test-time scaling

Text Reasoning 262,144 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon2-11B

Falcon 2

Research, multilingual text generation, fine-tuning, quantization, and self-hosted inference

Text General Purpose 8,192 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon3-10B-Base

Falcon3

Fine-tuning, multilingual text generation, language-model research, code and mathematics experimentation, and self-hosted inference

Text General Purpose 32,768 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon3-10B-Instruct

Falcon3

Self-hosted multilingual assistants, STEM and mathematics tasks, coding, instruction following, research, and local function-calling systems.

Text General Purpose 32,768 ctx Tool use Streaming
View model →
◎
Technology Innovation Institute

Falcon3-1B-Base

Falcon3

Local inference, multilingual text completion, research, continued pretraining, domain adaptation, and fine-tuning on constrained hardware

Text Lightweight 4,096 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon3-1B-Instruct

Falcon3

Lightweight local assistants, multilingual instruction following, extraction, classification, education, and resource-conscious deployments

Text Lightweight 8,192 ctx Tool use Streaming
View model →
◎
Technology Innovation Institute

Falcon3-3B-Base

Falcon3

Fine-tuning, multilingual text generation, research, compact local inference, and edge-oriented deployments

Text General Purpose 8,192 ctx Streaming
View model →
◎
Technology Innovation Institute

Falcon3-3B-Instruct

Falcon3

Self-hosted chat assistants, multilingual text generation, lightweight coding and mathematics, local inference, and resource-conscious deployments

Text General Purpose 32,000 ctx Tool use Streaming
View model →
◎
Technology Innovation Institute

Falcon3-7B-Base

Falcon3

Fine-tuning, multilingual text generation, code and mathematics experiments, research, and self-hosted inference

Text General Purpose 32,768 ctx
View model →
◎
Technology Innovation Institute

Falcon3-7B-Instruct

Falcon3

Self-hosted multilingual chat, instruction following, reasoning, mathematics, coding, research, and long-context text generation

Text General Purpose 32,768 ctx Tool use Streaming
View model →
◎
Microsoft

Florence-2-base

Florence-2

Local image captioning, object detection, phrase grounding, region description, OCR and lightweight computer-vision pipelines

Text Multimodal 1,024 ctx Image input
View model →
◎
Microsoft

Florence-2-large

Florence-2

Local image captioning, object detection, visual grounding, OCR, region annotation and multi-task computer-vision pipelines.

Text Multimodal 1,024 ctx Image input
View model →
◎
Google DeepMind

Gemini 2.5 Computer Use

Gemini 2.5

Browser automation, visual UI interaction, repetitive web workflows, form filling, and user-interface testing

Text Multimodal 128,000 ctx Image input Tool use Structured output
View model →
◎
Google DeepMind

Gemini 2.5 Flash

Gemini 2.5

Large-scale multimodal processing, low-latency reasoning, coding, data extraction, tool-using agents, and applications requiring a very large context window.

Text General Purpose 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 2.5 Flash Live

Gemini 2.5

Real-time voice and video agents, speech-to-speech assistants, interactive customer support, tutoring, coaching, and multimodal Live API applications

Text Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 2.5 Flash-Lite

Gemini 2.5

High-volume classification, simple extraction, lightweight multimodal analysis, routing, tagging, summarization, and extremely latency-sensitive applications.

Text Lightweight 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 2.5 Pro

Gemini 2.5

Advanced coding, complex reasoning, mathematics, STEM analysis, long documents, large codebases, multimodal analysis, and tool-using agents.

Text Reasoning 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3 Flash

Gemini 3

Agentic workflows, everyday coding, reasoning and planning, multimodal analysis, long-context document work, and cost-sensitive tool-using applications

Text Multimodal 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.1 Flash Live Preview

Gemini 3.1

Low-latency voice agents, real-time dialogue, multimodal live sessions, and interactive audio applications

Text Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.1 Flash-Lite

Gemini 3.1

High-volume translation, classification, extraction, summarization, document processing, and lightweight tool-using agent workflows

Text Lightweight 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.1 Pro

Gemini 3

Complex reasoning, advanced software engineering, long-context multimodal analysis, and agentic workflows requiring reliable tool use

Text Reasoning 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.5 Flash

Gemini 3.5

Agentic workflows, coding agents, long-context multimodal analysis, tool use, and scaled production applications

Text General Purpose 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google

Gemini 3.5 Flash-Lite

Gemini 3.5

High-volume, latency-sensitive agentic workflows, document parsing, translation, classification, data extraction, search-backed applications, and multimodal sub-agents.

Text Lightweight 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.5 Live Translate

Gemini 3.5 Audio

Low-latency, real-time speech-to-speech translation for calls, meetings, travel, customer support, and multilingual voice applications

Text Multimodal 131,072 ctx Audio input Streaming
View model →
◎
Google DeepMind

Gemini 3.5 Transcribe

Gemini 3.5 Audio

Pre-recorded audio transcription, multilingual speech recognition, speaker-labeled transcripts, timestamped transcripts, smart dictation, and domain-specific vocabulary

Text Other 96,000 ctx Audio input
View model →
◎
Google DeepMind

Gemini 3.6 Flash

Gemini 3

Fast multimodal applications, coding assistants, long-context document and video analysis, tool-using agents, enterprise workflows, and rapid agentic execution

Text General Purpose 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.7 Flash

Gemini 3

Coding, software engineering, tool-using agents, long-context multimodal analysis, structured extraction and high-throughput enterprise workflows

Text General Purpose 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.8 Flash

Gemini 3

Long-horizon software engineering, autonomous agents, multimodal document workflows, enterprise knowledge work, and tool-using applications.

Text Multimodal 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.8 Live

Gemini 3.8

Low-latency voice agents, real-time audio-to-audio dialogue, multimodal assistants, and interactive tool-using applications

Text Multimodal 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.8 Live Extended Thinking

Gemini 3.8 Audio

Complex real-time voice agents, multi-step problem solving, asynchronous tool workflows, technical support, travel coordination, and spoken STEM or coding tutoring

Text Reasoning 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Deep Research

Gemini Deep Research

Autonomous market research, due diligence, literature reviews, competitive analysis, source-heavy investigations, and cited research reports

Text Other 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Deep Research Max

Gemini Deep Research

Comprehensive market research, competitive analysis, due diligence, literature reviews, and source-rich investigative reports

Text Reasoning 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Robotics ER 2

Gemini Robotics ER

High-level robot planning, spatial reasoning, video progress tracking, tool orchestration, and multi-robot collaboration

Text Reasoning 131,072 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Robotics ER 2 Streaming

Gemini Robotics ER 2

Low-latency robotic agents, continuous audio/video monitoring, function-based robot orchestration, warehouse workflows, and multi-robot coordination

Text Reasoning 131,072 ctx Image input Audio input Video input
View model →
◎
NVIDIA

GenMol

GenMol

De novo molecular design, fragment-constrained generation, linker design, scaffold decoration, hit generation, and lead optimization

Text Other 512 ctx
View model →
◎
Z.AI

GLM-5.2

GLM-5

Long-horizon software engineering, repository-scale coding, code agents, complex refactoring, large-document analysis, research implementation, and tool-using workflows.

Text Reasoning 1,000,000 ctx Tool use Structured output Streaming
View model →
◎
Z.ai

GLM-5.3

GLM-5

Complex software engineering, large codebases, long-horizon coding agents, terminal workflows, technical research, and authorized cybersecurity analysis

Text Reasoning 1,000,000 ctx Tool use Structured output Streaming
View model →
◎
Z.ai

GLM-5.3-Flash

GLM-5

Multimodal coding agents, screenshot and interface understanding, long-context document work, visual verification, tool-assisted research, and cost-sensitive agent workflows

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Z.AI

GLM-ASR-2512

GLM-ASR

Hosted multilingual speech-to-text, short real-time transcriptions, captions, meeting snippets, voice input, customer-service audio, and terminology-aware transcription

Text Other Audio input Streaming
View model →
◎
Z.ai

GLM-OCR

GLM-OCR

High-volume OCR, PDF parsing, table and formula recognition, handwriting, receipts, forms, information extraction, and local document-processing pipelines

Text Other Image input Structured output
View model →
◎
OpenAI

GPT-3.5 Turbo

GPT-3.5

Low-cost, high-volume text generation, summarization, classification, extraction, simple chatbots, and legacy API integrations

Text General Purpose 16,385 ctx
View model →
◎
OpenAI

GPT-4

GPT-4

Maintaining established GPT-4 integrations, general-purpose text generation, analysis, writing, and coding workloads

Text General Purpose 8,192 ctx Image input Tool use Streaming
View model →
◎
OpenAI

GPT-4 Turbo

GPT-4

Legacy high-context text and image analysis, function calling, JSON-mode workflows, and existing GPT-4 Turbo integrations

Text Multimodal 128,000 ctx Image input Tool use Streaming
View model →
◎
OpenAI

GPT-4 Turbo Preview

GPT-4 Turbo

Historical long-context text generation, document analysis, structured text generation, and general-purpose assistant applications.

Text General Purpose 128,000 ctx
View model →
◎
OpenAI

GPT-4.1

GPT-4.1

Software engineering, long-context document analysis, precise instruction following, structured extraction, tool-enabled agents, and image understanding

Text General Purpose 1,047,576 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-4.1 Mini

GPT-4.1

Fast, cost-efficient instruction following, coding assistance, image understanding, structured extraction, tool calling, and long-context API applications

Text Lightweight 1,047,576 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-4.1 nano

GPT-4.1

High-volume, latency-sensitive classification, extraction, routing, summarization, lightweight assistants, image-assisted analysis, and simple tool-calling workflows

Text Lightweight 1,047,576 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-4.5 Preview

GPT-4.5

Historical general-purpose writing, creative work, nuanced communication, image understanding, programming assistance, and applications needing function calling or structured outputs.

Text General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-4o

GPT-4o

General-purpose assistants, image understanding, coding help, structured extraction, multilingual generation, and latency-sensitive API workflows

Text Multimodal 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-4o Audio

GPT-4o

Voice assistants, spoken conversational agents, audio-enabled customer service, and applications requiring direct audio understanding and speech generation

Text Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-4o Mini

GPT-4o

Low-cost, high-volume text and image understanding, classification, extraction, translation, tagging, customer support, routing, and structured data generation

Text Lightweight 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-4o Mini Audio

GPT-4o

Lower-cost audio understanding, conversational voice interfaces, and applications requiring text and spoken-audio input/output

Text Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-4o Mini Realtime

GPT-4o

Low-cost realtime voice assistants, speech-to-speech interfaces, interactive audio applications, and conversational prototypes

Text Lightweight 16,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-4o Mini Search Preview

GPT-4o Mini Search Preview

Legacy Chat Completions applications requiring low-cost, search-grounded text responses

Text Lightweight 128,000 ctx Tool use Web search Structured output
View model →
◎
OpenAI

GPT-4o Mini Transcribe

GPT-4o Mini

Lower-cost multilingual speech transcription, meeting notes, call-center transcripts, voice-note conversion, and audio-to-text pipelines

Text Other 16,000 ctx Audio input Streaming
View model →
◎
OpenAI

GPT-4o Realtime

GPT-4o

Low-latency voice assistants, speech-to-speech applications, live translation, language learning, and interactive customer support

Text Multimodal 32,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-4o Search Preview

GPT-4o

Historical web-search applications built around OpenAI Chat Completions

Text Other 128,000 ctx Tool use Structured output Streaming
View model →
◎
OpenAI

GPT-4o Transcribe

GPT-4o

Accurate speech-to-text conversion, meeting transcription, call transcription, voice-agent input, and prompted domain-specific transcription

Text Other 16,000 ctx Audio input Streaming
View model →
◎
OpenAI

GPT-4o Transcribe Diarize

GPT-4o Transcribe

Multi-speaker meeting, interview, call, podcast, and research transcription with speaker labels.

Text Other 16,000 ctx Audio input Streaming
View model →
◎
OpenAI

GPT-5

GPT-5

Complex coding, reasoning, research, long-context analysis, visual understanding, tool-using agents, and structured professional workflows

Text Reasoning 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5 Chat

GPT-5

ChatGPT-aligned conversational applications, text generation, image-aware question answering, structured outputs, and tool-enabled workflows requiring GPT-5 compatibility

Text General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5 Mini

GPT-5

Cost-sensitive reasoning, coding assistance, structured extraction, document processing, high-volume automation, and tool-enabled workflows

Text Lightweight 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5 nano

GPT-5

High-volume classification, summarization, extraction, ranking, routing, image-assisted analysis, and lightweight coding subagents

Text Lightweight 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5 Pro

GPT-5

Difficult research, mathematics, science, complex coding, high-stakes analysis, and tool-using workflows where maximum answer quality matters more than latency or cost.

Text Reasoning 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5-Codex

GPT-5

Agentic software engineering, repository-level coding, code review, debugging, refactoring, test generation, and frontend work using screenshots

Text Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.1

GPT-5

Coding, long-context analysis, tool-using agents, structured outputs, and multi-step workflows

Text General Purpose 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.1 Chat

GPT-5.1

Conversational assistants, instruction following, image-grounded chat, structured extraction, streaming responses, and tool-using API workflows.

Text General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.1-Codex

GPT-5.1

Agentic software engineering, code generation, debugging, refactoring, testing, code review, and long-running Codex workflows

Text Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.1-Codex Mini

GPT-5.1-Codex

Cost-sensitive agentic coding, code editing, repository maintenance, and Codex-style workflows

Text Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.1-Codex-Max

GPT-5.1-Codex

Long-running agentic coding, repository-scale refactoring, multi-file implementation, debugging, code review, pull-request creation, and extended Codex workflows.

Text Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.2

GPT-5.2

Complex professional work, long-context analysis, coding, document and spreadsheet workflows, visual understanding, and multi-step agents

Text Reasoning 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.2 Chat

GPT-5.2

ChatGPT-aligned conversational applications, general writing, summarization, translation, vision-enabled assistants, and tool-calling workflows

Text General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.2 Pro

GPT-5.2

Complex professional reasoning, advanced analysis, scientific and mathematical work, high-quality coding, long-context document analysis, and tool-using workflows.

Text Reasoning 400,000 ctx Image input Tool use Streaming
View model →
◎
OpenAI

GPT-5.2-Codex

GPT-5.2

Long-horizon agentic coding, large refactors, code migrations, repository-scale changes, terminal workflows, Windows development and defensive cybersecurity

Text Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.3 Chat

GPT-5.3

Fast general-purpose conversation, writing, summarization, text-and-image understanding, streaming responses, and function-calling applications

Text General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.3-Codex

GPT-5.3

Long-running agentic software engineering, codebase maintenance, debugging, testing, web development, tool-driven development, and technical computer workflows

Text Coding 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.4

GPT-5.4

Complex professional work, advanced reasoning, software engineering, long-horizon agents, visual document analysis, computer use, research, and tool-heavy workflows

Text Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.4 Mini

GPT-5.4

High-volume coding assistants, computer-use agents, subagents, tool calling, image reasoning, document workflows, and latency-sensitive applications

Text Lightweight 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.4 nano

GPT-5.4

High-volume classification, data extraction, ranking, image understanding, routing, and lightweight coding subagents

Text Lightweight 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.4 Pro

GPT-5.4

High-stakes reasoning, professional knowledge work, long-context analysis, complex coding, web research and agentic workflows requiring maximum answer quality

Text Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.4-Cyber

GPT-5.4

Authorized vulnerability research, defensive cybersecurity operations, malware analysis, security testing, and binary reverse engineering.

Text Coding
View model →
◎
OpenAI

GPT-5.5

GPT-5.5

Complex coding, long-context research, professional analysis, tool-heavy agents, computer use, and multi-step workflow execution

Text General Purpose 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.5 Pro

GPT-5.5

High-accuracy reasoning, complex coding, long-context research, data analysis, and multi-step professional workflows

Text Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.6 Cyber

GPT-5.6

Authorized vulnerability research, exploit validation, exploit-chain development, advanced security testing, vulnerability triage, and defensive cybersecurity agents

Text Coding 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.6 Luna

GPT-5.6

High-volume classification, summarization, routing, extraction, document understanding, agent automation, routine coding assistance, and cost-sensitive tool-using applications.

Text Lightweight 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.6 Sol

GPT-5.6

Complex reasoning, coding, research, cybersecurity, science, long-context analysis, document-heavy workflows, and tool-using agents

Text Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.6 Terra

GPT-5.6

Cost-conscious reasoning, coding agents, long-context analysis, structured business automation, research workflows, and tool-enabled production applications

Text General Purpose 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-6 Astra

GPT-6

Complex reasoning, agentic coding, computer use, web research, scientific and professional workflows, and long-context document tasks

Text Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-6 Luna

GPT-6

High-volume reasoning, document analysis, coding assistance, retrieval-augmented generation, and repeatable agent workflows

Text Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-6 Sol

GPT-6

Complex coding, long-context reasoning, software engineering, research, computer use, and agentic workflows with tools.

Text Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-Audio

GPT-Audio

Audio-enabled chat applications, voice interfaces, spoken assistants, and applications requiring direct audio understanding and generation through Chat Completions.

Text Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Audio Mini

GPT-Audio

Cost-sensitive, turn-based audio conversations, voice assistants, and audio-enabled applications using function calling

Text Multimodal 128,000 ctx Audio input Tool use
View model →
◎
OpenAI

GPT-Audio-1.5

GPT-Audio

Audio-in, audio-out conversational applications using the Chat Completions API, including voice assistants and tool-enabled spoken interfaces.

Text Multimodal 128,000 ctx Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Image-1.5

GPT Image

Production image generation, image editing, branded graphics, ecommerce product imagery, marketing assets, and workflows requiring preservation of important visual details

Text Other Image input
View model →
◎
OpenAI

GPT-Live 1

GPT-Live

Natural low-latency voice agents, customer support, conversational workflows, live assistance, and applications requiring interruption-aware speech interaction

Text Multimodal Audio input Tool use Streaming
View model →
◎
OpenAI

GPT-Live-Transcribe

GPT-Live-Transcribe

Low-latency live captions, realtime call transcription, microphone streams, telephony audio, and voice-interface speech recognition

Text Other Audio input Streaming
View model →
◎
OpenAI

gpt-oss-120b

GPT-OSS

Self-hosted reasoning, coding, agentic workflows, private deployments, research, and fine-tuning

Text Reasoning 131,072 ctx Tool use Web search Structured output
View model →
◎
OpenAI

gpt-oss-20b

gpt-oss

Local and private reasoning applications, coding assistants, agentic workflows, on-device or edge inference, fine-tuning, and cost-sensitive deployments with suitable hardware.

Text Reasoning 131,072 ctx Tool use Web search Structured output
View model →
◎
OpenAI

gpt-oss-safeguard-120b

gpt-oss-safeguard

Custom-policy safety classification, LLM input and output filtering, trust and safety labeling, nuanced moderation review, and offline safety analysis

Text Safety 131,072 ctx Structured output
View model →
◎
OpenAI

gpt-oss-safeguard-20b

GPT-OSS-Safeguard

Policy-based safety classification, LLM input/output filtering, content labeling, trust and safety review, and self-hosted moderation workflows

Text Other 131,072 ctx Structured output
View model →
◎
OpenAI

GPT-Realtime

GPT-Realtime

Low-latency speech-to-speech voice agents, realtime customer support, education, accessibility, and conversational applications with function calling

Text Realtime 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime Mini

GPT-Realtime

Cost-sensitive realtime voice agents, speech-to-speech applications, interactive assistants, and multimodal interfaces

Text Realtime 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-1.5

GPT-Realtime

Low-latency speech-to-speech voice agents, customer support, realtime assistants, and audio applications that need function calling.

Text Realtime Audio 32,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2

GPT-Realtime

Reasoning voice agents, speech-to-speech applications, customer support, live assistants, tool-driven workflows, and long conversational sessions

Text Multimodal 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2.1

GPT-Realtime

Low-latency speech-to-speech agents, customer-service voice workflows, realtime tool use, telephony, and multimodal assistants with image input

Text Realtime 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-2.1 Mini

GPT-Realtime-2.1

Lower-cost, low-latency realtime voice agents, speech-to-speech assistants, and tool-enabled conversational applications

Text Lightweight 128,000 ctx Image input Audio input Tool use
View model →
◎
OpenAI

GPT-Realtime-Translate

GPT-Realtime

Low-latency spoken translation, multilingual calls, live interpretation, broadcasts, meetings, lessons, video rooms, captions, and translated audio experiences.

Text Other 16,000 ctx Audio input Streaming
View model →
◎
OpenAI

GPT-Realtime-Whisper

GPT-Realtime

Low-latency live transcription, captions, meeting notes, call analysis, voice-agent input, and continuous speech-to-text workflows

Text Other 16,000 ctx Audio input Streaming
View model →
◎
OpenAI

GPT-Rosalind

GPT-Rosalind

Governed biology, genomics, medicinal chemistry, protein analysis, drug discovery, literature synthesis, wet-lab troubleshooting, and scientific tool workflows

Text Reasoning Image input Tool use
View model →
◎
OpenAI

GPT-Transcribe

GPT-Transcribe

High-accuracy transcription of recorded audio, streamed file transcripts, multilingual recordings, and domain-specific speech with keyword or language hints

Text Other Audio input Streaming
View model →
◎
IBM

Granite 4.0 3B Vision

Granite Vision

Chart extraction, table parsing, semantic key-value extraction, visual document processing, and local enterprise RAG pipelines

Text Multimodal Image input Structured output
View model →
◎
IBM

Granite 4.1 30B

Granite 4.1

Self-hosted enterprise assistants, long-context RAG, multilingual applications, coding, structured extraction, and tool-calling agents

Text General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite 4.1 3B

Granite 4.1

Efficient local or private deployment, multilingual enterprise text processing, RAG, summarization, extraction, coding assistance, function calling, and lightweight AI assistants.

Text Lightweight 131,072 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite 4.1 8B

Granite 4.1

Self-hosted enterprise assistants, multilingual text generation, RAG, coding assistance, structured extraction, and tool-calling agents

Text General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite 4.2 30B

Granite 4.2

Self-hosted enterprise reasoning, coding agents, multilingual applications, tool calling, long-context workflows, and organizations requiring Apache 2.0 licensing

Text Reasoning 131,072 ctx Tool use
View model →
◎
IBM

Granite 4.2 3B

Granite 4.2

Efficient reasoning, coding assistance, tool calling, multilingual dialogue, local deployment, and lightweight enterprise agents

Text Reasoning 131,072 ctx Tool use
View model →
◎
IBM

Granite 4.2 8B

Granite 4.2

Local or self-hosted reasoning, coding assistants, tool calling, multilingual dialogue, retrieval-augmented generation, and agentic workflows

Text Reasoning 131,072 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite Guardian 3.8B

Granite Guardian

Prompt and response safety classification, jailbreak detection, RAG groundedness and relevance checks, hallucination evaluation, and enterprise AI guardrails

Text Other 131,072 ctx
View model →
◎
IBM

Granite Guardian 4.1 8B

Granite Guardian

AI safety guardrails, jailbreak detection, RAG groundedness and relevance checks, function-call hallucination detection, custom criteria evaluation, and best-of-N response ranking

Text Other 8,192 ctx
View model →
◎
IBM

Granite Speech 4.1 2B

Granite Speech 4.1

Multilingual offline transcription, speech-to-text, speech translation, subtitle generation, and domain-specific recognition with keyword biasing

Text Other 128,000 ctx Audio input
View model →
◎
IBM

Granite Speech 4.1 2B NAR

Granite Speech 4.1

High-throughput multilingual speech transcription where low inference latency is more important than maximum recognition accuracy.

Text Speech Recognition 4,096 ctx Audio input
View model →
◎
IBM

Granite Speech 5.0 470M TurboCTC

Granite Speech 5.0

High-throughput, low-latency English speech-to-text transcription on local, edge, and enterprise systems

Text Speech Recognition Audio input Streaming
View model →
◎
IBM

Granite Speech 5.0 470M TurboCTC NC

Granite Speech 5.0

Fast local English speech-to-text transcription, edge-device ASR, research, browser demos, and high-throughput noncommercial batch processing

Text Other Audio input Streaming
View model →
◎
IBM

Granite Vision 3.3 2B

Granite Vision

Enterprise document understanding, chart and table extraction, OCR-oriented image analysis, visual question answering, and multimodal RAG

Text Multimodal 131,072 ctx Image input
View model →
◎
IBM

Granite Vision 4.1 4B

Granite Vision 4.1

Structured extraction from charts, tables, invoices, forms, and enterprise document images

Text Multimodal Image input Structured output Streaming
View model →
◎
IBM

granite-13b-chat-v2

Granite 13B V2

English enterprise chat, retrieval-augmented generation, question answering, summarization, extraction, and classification

Text General Purpose 8,192 ctx
View model →
◎
IBM

granite-20b-code-base-schema-linking

Granite Code

Schema linking and relevant-table or relevant-column selection in text-to-SQL pipelines

Text Coding 8,192 ctx
View model →
◎
IBM

Granite-20B-Code-Base-SQL-Gen

Granite Code

Natural-language-to-SQL generation over structured databases, analytics assistants, and the SQL-generation stage of text-to-SQL pipelines.

Text Coding 8,192 ctx
View model →
◎
IBM

Granite-20B-Code-Instruct

Granite Code

Self-hosted coding assistants, code generation, code conversion, code explanation, and programming experiments where an Apache 2.0 open-weight model is preferred.

Text Coding 8,192 ctx Streaming
View model →
◎
IBM

granite-20b-multilingual

Granite

Multilingual enterprise question answering, retrieval-augmented generation, summarization, extraction, classification, and text generation in English, German, Spanish, French, and Portuguese

Text General Purpose 8,192 ctx
View model →
◎
IBM

Granite-3.0-8B-Base

Granite 3.0

Fine-tuning, domain adaptation, text generation, summarization, extraction, classification, and question answering

Text General Purpose 4,096 ctx
View model →
◎
IBM

Granite-3.1-8B-Base

Granite 3.1

Fine-tuning, long-context document processing, enterprise text classification, extraction, summarization, question answering, and self-managed deployment

Text General Purpose 131,072 ctx
View model →
◎
IBM

Granite-3.1-8B-Instruct

Granite 3.1

Self-hosted enterprise assistants, long-context document analysis, RAG, summarization, extraction, multilingual text workflows, and function calling

Text General Purpose 131,072 ctx Tool use Streaming
View model →
◎
IBM

Granite-3.2-8B-Instruct

Granite 3.2

Long-context enterprise RAG, document and meeting summarization, information extraction, multilingual dialogue, code-related tasks, function calling, and self-hosted deployments

Text Reasoning 131,072 ctx Tool use Streaming
View model →
◎
IBM

Granite-3.3-2B-Instruct

Granite 3.3

Local or dedicated enterprise assistants, RAG, long-document summarization, multilingual text workflows, coding assistance, reasoning, and resource-conscious inference

Text General Purpose 131,072 ctx Tool use Streaming
View model →
◎
IBM

Granite-3.3-8B-Instruct

Granite 3.3

Self-hosted enterprise assistants, long-context RAG, summarization, multilingual text tasks, coding assistance, function calling, and cost-sensitive deployments.

Text General Purpose 131,072 ctx Tool use Streaming
View model →
◎
IBM

Granite-34B-Code-Instruct

Granite Code

Self-hosted code generation, code explanation, code repair, code conversion, coding assistants, and legacy Granite Code compatibility

Text Coding 8,192 ctx Streaming
View model →
◎
IBM

Granite-3B-Code-Instruct

Granite Code

Self-hosted coding assistants, code generation, code explanation, code conversion, repository-scale prompts, and historical or reproducible research

Text Coding 128,000 ctx
View model →
◎
IBM

Granite-4.0-H-Micro

Granite 4.0

Low-latency local inference, long-context text generation, RAG, multilingual assistants, coding assistance, and lightweight tool-calling agents

Text Lightweight 128,000 ctx Tool use
View model →
◎
IBM

Granite-4.0-H-Small

Granite 4.0

Enterprise RAG, multi-tool agents, function calling, customer-support automation, multilingual instruction following, and long-context workloads

Text General Purpose 131,072 ctx Tool use Structured output
View model →
◎
IBM

Granite-4.0-H-Tiny

Granite 4.0

Efficient enterprise assistants, multilingual text generation, RAG, classification, extraction, summarization, coding assistance, fill-in-the-middle completion, structured JSON, and tool-calling workflows

Text Lightweight 128,000 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite-7B-Lab

Granite 7B

Self-hosted English text generation, summarization, extraction, classification, experimentation with LAB-aligned open-weight models, and resource-conscious deployments

Text General Purpose 8,192 ctx
View model →
◎
IBM

Granite-8B-Code-Instruct

Granite Code

Local code generation, explanation, repair, translation, and research with an open Apache 2.0 model

Text Coding 4,096 ctx
View model →
◎
IBM

granite-8b-japanese

Granite

Japanese text generation, summarization, classification, extraction, question answering, and Japanese-English translation in legacy IBM enterprise deployments

Text General Purpose 4,096 ctx
View model →
◎
IBM

Granite-Speech-4.1-2B-Plus

Granite Speech 4.1

Multilingual speech-to-text with speaker labels, word-level timestamps, keyword biasing, and self-hosted enterprise transcription

Text Other 4,096 ctx Audio input
View model →
◎
xAI

Grok 4.20 Multi-Agent Beta

Grok 4.20

Deep research, parallel investigation, long-context analysis, and tool-assisted synthesis

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.20-0309-non-reasoning

Grok 4.20

Fast general-purpose text generation, image-aware analysis, coding assistance, structured extraction, tool-calling agents and large-context workflows

Text General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.20-0309-reasoning

Grok 4.20

Complex reasoning, coding, technical research, long-context document analysis, image understanding, structured responses, and tool-enabled agentic workflows

Text Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.3

Grok 4

Long-context analysis, enterprise agents, research, coding assistance, structured extraction and tool-enabled workflows

Text Multimodal 1,000,000 ctx Image input Tool use Web search
View model →
◎
SpaceXAI

Grok 4.5

Grok 4

Software engineering, codebase analysis, technical reasoning, long-context work, tool-using agents, document analysis, and structured workflow automation

Text Coding 500,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.6

Grok 4.6

Agentic coding, long-context software engineering, research, knowledge work, visual analysis, structured extraction, and tool-using workflows.

Text Reasoning 500,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.7

Grok 4

Advanced software engineering, long-context reasoning, agentic tool use, research with web or X search, and professional knowledge work.

Text Reasoning 500,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.7 Fast

Grok 4.7

Low-latency coding, interactive development, and agentic workflows in Cursor or Grok Build

Text General Purpose 500,000 ctx Image input Tool use Structured output
View model →
◎
SpaceXAI

Grok Build 0.1

Grok Build

Agentic coding, web development, debugging, software engineering workflows, MCP integrations, and fast tool-calling applications.

Text Coding 256,000 ctx Image input Tool use Structured output
View model →
◎
xAI

Grok Voice Think Fast 2.0

Grok Voice Think Fast

Realtime voice agents, customer support, telephony, sales, multilingual conversations, and tool-enabled spoken workflows

Text Multimodal Audio input Tool use Web search
View model →
◎
xAI

Grok Voice Transcribe 1.0

Grok Voice Transcribe

Batch and real-time multilingual speech transcription, dictation, voice assistants, accessibility, meetings, and customer-support audio

Text Speech Recognition Audio input Streaming
View model →
◎
xAI

Grok Voice Transcribe 2.0

Grok Voice Transcribe

Batch and real-time transcription of multilingual, noisy, conversational, telephony, and voice-agent audio

Text Other Audio input Streaming
View model →
◎
NAVER

HCX-003

HyperCLOVA X

Korean-focused text generation, summarization, extraction, classification, tuning, and batch text workflows

Text General Purpose 8,192 ctx Streaming
View model →
◎
NAVER

HCX-005

HyperCLOVA X

Korean-language business applications, image understanding, document and visual analysis, instruction following, and API-based assistants.

Text Multimodal 128,000 ctx Image input Tool use Streaming
View model →
◎
NAVER

HCX-007

HyperCLOVA X

Complex reasoning, mathematics, science, language reasoning, writing, long-context text generation, and Korean-language enterprise applications

Text Reasoning 128,000 ctx Tool use Structured output Streaming
View model →
◎
NAVER Cloud

HCX-DASH-001

HCX-DASH

Fast, cost-sensitive Korean text generation, classification, summarization, report drafting, data expansion and customized enterprise chatbots

Text Lightweight 4,096 ctx Streaming
View model →
◎
NAVER

HCX-DASH-002

HyperCLOVA X

Fast, high-throughput text generation, classification, summarization, simple extraction, and function-calling workflows

Text Lightweight 32,000 ctx Tool use Streaming
View model →
◎
Tencent

Hunyuan-0.5B-Instruct

Hunyuan

Lightweight local assistants, edge-oriented inference, long-context text processing, prototyping, quantized deployment, and low-resource instruction-following workloads.

Text Lightweight 262,144 ctx
View model →
◎
Tencent

Hunyuan-1.8B-Instruct

Hunyuan

Efficient local text generation, long-context analysis, lightweight reasoning, coding assistance, and agent-oriented applications

Text Lightweight 256,000 ctx Tool use Streaming
View model →
◎
Tencent

Hunyuan-4B-Instruct

Hunyuan

Local and self-hosted text generation, Chinese and multilingual instruction following, long-context analysis, mathematics, coding, reasoning, and lightweight agent workloads

Text General Purpose 262,144 ctx Tool use Streaming
View model →
◎
Tencent

Hunyuan-7B-Instruct

Hunyuan

Chinese and multilingual text generation, reasoning, coding, long-context workloads, private local inference, quantized deployment, and domain fine-tuning.

Text General Purpose 262,144 ctx
View model →
◎
Tencent

Hunyuan-A13B

Hunyuan-A13B

Open-weight reasoning, long-context analysis, mathematics, science, coding, agent workflows, and cost-conscious self-hosted inference

Text Reasoning 262,144 ctx Tool use Streaming
View model →
◎
Tencent

Hunyuan-Large

Hunyuan-Large

Large-scale Chinese and English text generation, reasoning, mathematics, coding, long-context analysis, research, and self-hosted experimentation

Text General Purpose 256,000 ctx Streaming
View model →
◎
Tencent

Hunyuan-MT-7B

Hunyuan-MT

Self-hosted multilingual translation, localization, translation research, and Chinese dialect or minority-language translation

Text Other 32,768 ctx
View model →
◎
Tencent

HunyuanOCR-1.5

HunyuanOCR

Multilingual OCR, document parsing, text spotting, table and formula extraction, structured information extraction, and local visual-document processing

Text Multimodal 131,072 ctx Image input Structured output Streaming
View model →
◎
Tencent

Hy-ASR-3.0-Preview

Hy-ASR

Real-time and asynchronous Mandarin, English, mixed-language, and Chinese-dialect transcription, captions, subtitles, and short voice-command recognition

Text Other Audio input Streaming
View model →
◎
Tencent

Hy-MT2-Lite

Hy-MT2

Low-latency multilingual translation, localized content, structured translation instructions, and cost-sensitive production workflows

Text Lightweight 8,192 ctx Streaming
View model →
◎
Tencent

Hy-MT2-Pro

Hy-MT2

Professional, domain-specific and high-quality multilingual translation with contextual disambiguation and instruction following

Text Specialized 8,192 ctx Streaming
View model →
◎
Tencent

Hy-Role

Hy-Role

Chinese role-play, character simulation, fictional dialogue, AI avatars, and emotionally oriented conversational experiences

Text Other 32,000 ctx Streaming
View model →
◎
Tencent Hunyuan

HY-Vision-1.5-Thinking

Hunyuan Vision 1.5

Image-grounded reasoning, OCR, chart and document analysis, visual localization, educational problem solving, and multilingual visual question answering

Text Multimodal 40,000 ctx Image input Video input Tool use
View model →
◎
Tencent

HY-Vision-2.0-Instruct

HY-Vision

Image understanding, OCR, chart and diagram analysis, STEM visual reasoning, visual question answering, and multi-image comparison

Text Multimodal 44,000 ctx Image input
View model →
◎
Tencent

HY-Vision-Video

Hunyuan turbos-vision

Video description, video question answering, video summarization, content review, scene analysis, and video metadata generation.

Text Multimodal 32,000 ctx Video input
View model →
◎
Tencent

Hy3

Hy

Coding agents, long-context analysis, complex reasoning, productivity automation, structured workflows, and multi-step tool use

Text Reasoning 256,000 ctx Tool use Web search Structured output
View model →
◎
Tencent

Hy4 preview

Hy4

Long-context coding agents, complex tool-use workflows, productivity automation, document analysis, game development, and scientific reasoning

Text General Purpose 1,000,000 ctx Tool use Structured output Streaming
View model →
◎
NAVER

HyperCLOVA X SEED 0.5B

HyperCLOVA X SEED

Lightweight Korean conversational interfaces, mobile and edge applications, smart-home devices, wearables, and customer-support chatbots

Text Lightweight 4,096 ctx
View model →
◎
NAVER

HyperCLOVA X SEED 1.5B

HyperCLOVA X SEED

Korean-language applications, lightweight local inference, basic translation, education, business communication, specialized chatbots, and domain fine-tuning.

Text Lightweight 16,000 ctx
View model →
◎
NAVER Cloud

HyperCLOVA X SEED 14B Think

HyperCLOVA X SEED

Korean-language reasoning, mathematics, coding, instruction following, tool-connected agents, and self-hosted commercial applications

Text Reasoning 32,768 ctx Tool use Streaming
View model →
◎
NAVER

HyperCLOVA X SEED 32B Think

HyperCLOVA X SEED Think

Korean-language reasoning, visual question answering, long-context multimodal analysis, document and chart understanding, coding assistance, and tool-using AI agents

Text Reasoning 128,000 ctx Image input Video input Tool use
View model →
◎
NAVER Cloud

HyperCLOVA X SEED 3B

HyperCLOVA X SEED

Korean image and video understanding, visual question answering, chart and diagram interpretation, OCR-assisted analysis, tourism and cultural applications, and locally deployable fine-tuned systems.

Text Multimodal 16,000 ctx Image input Video input Streaming
View model →
◎
NAVER Cloud

HyperCLOVA X SEED 4B

HyperCLOVA X SEED

Korean visual and document understanding, video-and-audio analysis, edge AI, public-sector systems, defense intelligence, and air-gapped deployments

Text Multimodal Image input Audio input Video input
View model →
◎
NAVER

HyperCLOVA X SEED 8B Omni

HyperCLOVA X SEED

Korean-first any-to-any multimodal assistants, speech and vision applications, multimodal research, and self-hosted deployments

Text Multimodal 32,768 ctx Image input Audio input Video input
View model →
◎
AI21

Jamba 1.5 Large

Jamba 1.5

Long-context document analysis, retrieval-augmented generation, enterprise assistants, structured text generation, multilingual workflows, and self-hosted deployments

Text General Purpose 262,144 ctx Tool use Structured output Streaming
View model →
◎
AI21 Labs

Jamba Large 1.7

Jamba

Long-document analysis, grounded generation, enterprise RAG, document summarization, information extraction, private deployment, and multilingual text workflows.

Text General Purpose 256,000 ctx
View model →
◎
AI21 Labs

Jamba-v0.1

Jamba

Long-context text generation, open-weight research, experimentation, and private or self-hosted deployment

Text General Purpose 256,000 ctx
View model →
◎
DeepSeek

Janus-1.3B

Janus

Local research, image understanding, visual question answering, multimodal prototyping, and lightweight text-to-image experimentation

Text Multimodal 4,096 ctx Image input
View model →
◎
DeepSeek

JanusFlow-1.3B

JanusFlow

Local research, visual question answering, image interpretation, and compact text-to-image experimentation

Text Multimodal 4,096 ctx Image input
View model →
◎
LG AI Research

K-EXAONE-2.0-750B-A37B

K-EXAONE

Long-context reasoning, multilingual generation, Korean-language workloads, coding, tool-using agents, research and distributed self-hosted inference

Text Reasoning 262,144 ctx Tool use
View model →
◎
LG AI Research

K-EXAONE-236B-A23B

K-EXAONE

Self-hosted multilingual reasoning, Korean-language applications, coding, long-context document analysis, and tool-enabled agents

Text Reasoning 256,000 ctx Tool use
View model →
◎
Moonshot AI

Kimi K2 Base

Kimi K2

Fine-tuning, foundation-model research, custom language systems, coding experiments, and self-hosted inference

Text General Purpose 131,072 ctx Streaming
View model →
◎
Moonshot AI

Kimi K2 Instruct

Kimi K2

Open-weight coding assistants, tool-using agents, general-purpose chat, and self-hosted research deployments

Text General Purpose 131,072 ctx Tool use Streaming
View model →
◎
Moonshot AI

Kimi K2 Thinking

Kimi K2

Self-hosted reasoning agents, autonomous research, long-horizon tool workflows, coding, and complex multi-step analysis

Text Reasoning 262,144 ctx Tool use Streaming
View model →
◎
Moonshot AI

Kimi K2.5

Kimi K2.5

Multimodal coding, visual debugging, long-context analysis, tool-using agents, and complex research or office workflows

Text Multimodal 256,000 ctx Image input Video input Tool use
View model →
◎
Moonshot AI

Kimi K2.6

Kimi K2

Long-horizon software engineering, agentic coding, visual document understanding, tool-using workflows, and multi-agent orchestration

Text Multimodal 262,144 ctx Image input Tool use Web search
View model →
◎
Moonshot AI

Kimi K2.7 Code

Kimi K2

Long-horizon software engineering, repository-level coding, multi-file refactoring, debugging, coding agents, and tool-driven development workflows

Text Coding 262,144 ctx Image input Video input Tool use
View model →
◎
Moonshot AI

Kimi K2.8 Preview

Kimi K2

Long-context software development, repository analysis, code completion, multi-file refactoring, and agentic coding workflows

Text Coding 1,048,576 ctx Image input Video input Tool use
View model →
◎
Moonshot AI

Kimi K3

Kimi K3

Long-context coding, software engineering, multimodal document and video understanding, agentic workflows, technical research, and complex reasoning

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Moonshot AI

Kimi-Audio-7B

Kimi-Audio

Fine-tuning and research on speech recognition, audio understanding, audio classification, audio question answering, and speech-audio generation

Text Multimodal 8,192 ctx Audio input Streaming
View model →
◎
Moonshot AI

Kimi-Audio-7B-Instruct

Kimi-Audio

Self-hosted speech recognition, audio understanding, audio question answering, audio captioning, and spoken conversational agents

Text Multimodal Audio input Streaming
View model →
◎
Moonshot AI

Kimi-Dev-72B

Kimi-Dev

Repository-level software issue resolution, code repair, test generation, coding agents, and self-hosted software engineering workflows

Text Coding 131,072 ctx
View model →
◎
Moonshot AI

Kimi-Linear-48B-A3B-Base

Kimi Linear

Long-context text generation, local inference, continued pretraining, research, and fine-tuning

Text Lightweight 1,048,576 ctx Streaming
View model →
◎
Moonshot AI

Kimi-Linear-48B-A3B-Instruct

Kimi Linear

Long-context text generation, self-hosted assistants, document analysis, research, and efficient local inference

Text General Purpose 1,048,576 ctx Streaming
View model →
◎
Moonshot AI

Kimi-VL-A3B-Instruct

Kimi-VL

Self-hosted image and video understanding, OCR, long-document analysis, visual question answering, screenshot perception, and efficient multimodal applications

Text Multimodal 131,072 ctx Image input Video input Streaming
View model →
◎
Moonshot AI

Kimi-VL-A3B-Thinking

Kimi-VL

Local multimodal reasoning, mathematical visual question answering, OCR, document understanding, image and video analysis, and research on open-weight vision-language models

Text Reasoning 131,072 ctx Image input Video input Streaming
View model →
◎
Moonshot AI

Kimi-VL-A3B-Thinking-2506

Kimi-VL

Open-weight image and video reasoning, OCR, chart interpretation, visual mathematics, long PDFs, high-resolution screenshots and GUI-agent grounding

Text Reasoning 131,072 ctx Image input Video input Streaming
View model →
◎
Mistral AI

Leanstral 1.5

Leanstral

Lean 4 theorem proving, formal verification, autoformalization, proof debugging, and agentic proof engineering

Text Other 256,000 ctx Tool use Structured output
View model →
◎
Meta

Llama 3.1 405B

Llama 3.1

High-quality open-weight research, multilingual applications, coding, reasoning, synthetic-data generation, model distillation and self-hosted deployments with substantial infrastructure

Text General Purpose 131,072 ctx Tool use Streaming
View model →
◎
Meta

Llama 3.1 8B

Llama 3.1

Local inference, fine-tuning, private deployment, text generation, retrieval-augmented generation, research, and cost-sensitive applications

Text General Purpose 131,072 ctx Streaming
View model →
◎
Meta

Llama 3.2 11B Vision Instruct

Llama 3.2 Vision

Self-hosted visual question answering, image captioning, document analysis, visual reasoning, and multimodal assistants

Text Multimodal 128,000 ctx Image input Tool use Web search
View model →
◎
Meta

Llama 3.2 1B

Llama 3.2

Private local inference, mobile and edge assistants, summarization, rewriting, retrieval-supported generation, and lightweight multilingual applications

Text Lightweight 128,000 ctx Tool use Streaming
View model →
◎
Meta

Llama 3.2 3B

Llama 3.2

Private local inference, multilingual text generation, edge applications, model adaptation, and fine-tuning

Text Lightweight 128,000 ctx Tool use
View model →
◎
Meta

Llama 3.2 90B Vision Instruct

Llama 3.2

Visual question answering, image reasoning, chart and document understanding, image captioning, multimodal research, and self-hosted or partner-hosted AI applications.

Text Multimodal 128,000 ctx Image input Tool use Web search
View model →
◎
Meta

Llama 3.3 70B Instruct

Llama 3.3

Self-hosted or hosted multilingual chat, coding assistance, long-context text generation, tool calling, synthetic data, and applications requiring open model weights

Text General Purpose 128,000 ctx Tool use
View model →
◎
Meta

Llama 4 Maverick

Llama 4

Open-weight multimodal assistants, image understanding, visual question answering, coding, multilingual applications, creative writing and long-context text processing.

Text Multimodal 1,000,000 ctx Image input Tool use Streaming
View model →
◎
Meta

Llama 4 Scout

Llama 4

Long-context document and code analysis, visual question answering, multimodal assistants, multilingual applications, self-hosted inference, and customized deployments

Text Multimodal 10,000,000 ctx Image input Tool use Structured output
View model →
◎
Meta

Llama Guard 3-11B-Vision

Llama Guard 3

Safety classification of mixed text-and-image prompts and text responses in multimodal LLM systems

Text Other 128,000 ctx Image input
View model →
◎
Meta

Llama Guard 3-1B

Llama Guard 3

Low-cost prompt and response safety classification, local moderation, mobile and edge deployments, and customizable LLM guardrails

Text Other 131,072 ctx
View model →
◎
Meta

Llama Guard 3-8B

Llama Guard 3

Self-hosted multilingual input and output moderation for LLM applications

Text Other 131,072 ctx
View model →
◎
Meta

Llama Guard 4

Llama Guard

Text and image moderation for prompts and generated responses in generative AI systems

Text Other 8,192 ctx Image input
View model →
◎
Meta

Llama Prompt Guard 2 22M

Llama Prompt Guard 2

Low-latency detection of prompt injections and jailbreak attempts in LLM applications, agents, retrieved documents, and other untrusted text

Text Other 512 ctx
View model →
◎
Meta Llama

Llama Prompt Guard 2 86M

Llama Prompt Guard 2

Multilingual prompt-injection detection, jailbreak screening, agent security, and filtering untrusted text before it reaches an LLM

Text Other 512 ctx
View model →
◎
Aleph Alpha

Llama-3.1-70B-TFree-HAT-SFT

Llama-3.1-TFree-HAT

English and German instruction following, tokenizer-free language-model research, text-compression experiments, and self-hosted open-weight deployments

Text General Purpose 98,304 ctx
View model →
◎
Aleph Alpha

Llama-3.1-8B TFree HAT Base

Llama-3.1-8B TFree HAT

Research on tokenizer-free language modeling, English-German text generation, multilingual NLP, open-weight deployment, and custom model adaptation

Text General Purpose 262,144 ctx
View model →
◎
Aleph Alpha

Llama-3.1-8B TFree HAT DPO

TFree-HAT

English and German instruction following, multilingual research, tokenizer-free language-model experimentation, and self-hosted deployment

Text General Purpose 262,144 ctx
View model →
◎
Ai2

Llama-3.1-Tulu-3-405B

Tülu 3

Large-scale research, instruction-following evaluation, open-weight post-training research, mathematical reasoning, coding benchmarks, and self-hosted experimentation.

Text General Purpose 8,192 ctx Streaming
View model →
◎
Allen Institute for AI

Llama-3.1-Tulu-3-70B

Tülu 3

Self-hosted instruction following, reasoning, mathematics, coding, open-model research, and reproducible post-training experiments

Text General Purpose 131,072 ctx
View model →
◎
Allen Institute for AI

Llama-3.1-Tulu-3-8B

Tülu 3

Self-hosted instruction following, open post-training research, mathematics, coding, evaluation and domain adaptation

Text General Purpose 131,072 ctx
View model →
◎
Aleph Alpha

llama-3_1-8b-tfree-hat-sft

Llama-3.1-8B TFree HAT

English and German text generation, instruction-following research, tokenizer-free language-model research, and self-hosted non-commercial experimentation

Text General Purpose 262,144 ctx
View model →
◎
Aleph Alpha

Llama-TFree-HAT-Pretrained-7B-DPO

TFree-HAT

English and German instruction following, multilingual text generation, German-language applications, and research into tokenizer-free language models

Text General Purpose 163,840 ctx
View model →
◎
Cerebras

Llama3-DocChat-1.0-8B

Llama 3

Document-grounded conversational question answering, local retrieval-augmented generation, and open-weight research deployments

Text General Purpose 8,192 ctx Streaming
View model →
◎
Aleph Alpha

luminous-base

Luminous

Historical multilingual text completion, language-model research, and semantic representation workflows

Text General Purpose
View model →
◎
Aleph Alpha

luminous-base-control

Luminous

Legacy multilingual text generation, conversational applications, and explainability-oriented workflows using Aleph Alpha's first Luminous generation

Text General Purpose
View model →
◎
Aleph Alpha

luminous-extended

Luminous

Historical multilingual text-completion research and general language-task experiments

Text General Purpose
View model →
◎
Aleph Alpha

luminous-extended-control

Luminous

Steerable multilingual text generation, summarization, classification, question answering, and explainability-oriented enterprise workflows

Text General Purpose Streaming
View model →
◎
Aleph Alpha

luminous-supreme

Luminous

Historical multilingual text completion, language understanding research, and compatibility work involving Aleph Alpha’s original Luminous API

Text General Purpose
View model →
◎
Aleph Alpha

luminous-supreme-control

Luminous

Zero-shot multilingual text generation, instruction following, classification, conversational prototypes, and explainability-oriented enterprise workflows.

Text General Purpose 2,048 ctx Streaming
View model →
◎
Google DeepMind

Lyria 3 Pro

Lyria 3

Full-length AI music, soundtrack creation, songwriting, structured compositions, advertising audio, games, and creative production workflows

Text Other Image input
View model →
◎
Google DeepMind

Lyria 3.5

Lyria

Full-length AI-generated songs, vocal music, instrumental arrangements, songwriting experiments, soundtracks, and image-inspired music creation

Text Other 131,072 ctx Image input
View model →
◎
Aleph Alpha

MAGMA

MAGMA

Vision-language research, image captioning, visual question answering, and experiments with adapter-based multimodal fine-tuning

Text Multimodal Image input
View model →
◎
Microsoft

MAI-DS-R1

MAI-DS

Open-weight reasoning, mathematics, coding, research, and general text-generation applications requiring DeepSeek-R1-style reasoning with Microsoft post-training

Text Reasoning 163,840 ctx
View model →
◎
Microsoft AI

MAI-Thinking-1

MAI-Thinking

Complex mathematical reasoning, software engineering, quantitative enterprise analysis, long-context document work, and application-managed agent workflows.

Text Reasoning 256,000 ctx Tool use Streaming
View model →
◎
Microsoft

MAI-Transcribe-1.5

MAI-Transcribe

Multilingual speech-to-text, captions, meeting transcription, accessibility, call analysis, content workflows, voice-agent audio understanding, and domain-specific terminology

Text Other Audio input
View model →
◎
Microsoft AI

MAI-Transcribe-2

MAI-Transcribe

Multilingual audio transcription, meeting and contact-center records, captions, clinical notes, accessibility, media search, voice-agent evaluation, and domain-specific transcription with speaker labels and timestamps

Text Other Audio input
View model →
◎
Microsoft

MedImageParse

BiomedParse

Text-guided biomedical image segmentation, annotation assistance, organ and tumor delineation, pathology-cell analysis, and research-oriented medical imaging pipelines

Text Other Image input Structured output
View model →
◎
Xiaomi MiMo

MiMo-Audio-7B-Base

MiMo-Audio

Few-shot audio-language research, speech continuation, voice and style conversion, speech translation, speech editing, and audio-text experimentation

Text Multimodal 8,192 ctx Audio input
View model →
◎
Xiaomi

MiMo-Audio-7B-Instruct

MiMo-Audio

Local audio understanding, speech-to-text dialogue, spoken conversational agents, and controllable text-to-speech research

Text Multimodal 8,192 ctx Audio input
View model →
◎
Xiaomi

MiMo-Embodied-7B

MiMo-Embodied

Embodied AI research, autonomous-driving perception and planning, spatial reasoning, affordance prediction, robot navigation, and multimodal visual analysis

Text Multimodal 128,000 ctx Image input Video input
View model →
◎
Xiaomi

MiMo-V2-Flash-Base

MiMo-V2-Flash

Research, custom post-training, foundation-model experimentation, and self-hosted long-context inference

Text General Purpose 256,000 ctx
View model →
◎
Xiaomi

MiMo-V2.5-ASR

MiMo-V2.5

Speech-to-text transcription for Chinese, English, regional Chinese dialects, code-switched speech, lyrics, noisy recordings, and multi-speaker conversations

Text Other Audio input Streaming
View model →
◎
Xiaomi

MiMo-V2.5-Pro

MiMo V2.5

Long-horizon agent workflows, repository-scale coding, complex software engineering, tool-driven automation, and very large documents

Text Reasoning 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Xiaomi MiMo

MiMo-V2.6-Distill-Qwen-9B

MiMo-V2.6

Local coding agents, terminal automation, cybersecurity experimentation, visual coding and agentic reinforcement-learning research

Text Multimodal 262,144 ctx Image input Tool use Streaming
View model →
◎
Xiaomi

MiMo-V2.6-Flash

MiMo-V2.6

High-volume multimodal API workloads, coding assistants, agent automation, long-context document and repository analysis, tool-using workflows, and cost-sensitive professional applications.

Text Multimodal 1,000,000 ctx Image input Audio input Video input
View model →
◎
Xiaomi MiMo

MiMo-V2.6-Pro

MiMo-V2.6

Long-horizon agents, coding, cybersecurity, research, computer use, multimodal analysis, and complex multi-step workflows

Text Multimodal 1,000,000 ctx Image input Audio input Video input
View model →
◎
Xiaomi

MiMo-V2.6-Pro UltraSpeed

MiMo-V2.6

Latency-sensitive multimodal reasoning, coding, tool-use, long-context research, and interactive agent workflows

Text Reasoning 1,000,000 ctx Image input Audio input Video input
View model →
◎
Xiaomi MiMo

MiMo-VL-7B-RL-2508

MiMo-VL

Local image and video understanding, multimodal reasoning, visual question answering, OCR-oriented tasks, visual grounding, and GUI analysis

Text Multimodal 128,000 ctx Image input Video input
View model →
◎
Xiaomi MiMo

MiMo-VL-7B-SFT-2508

MiMo-VL

Self-hosted image and video understanding, multimodal reasoning research, supervised fine-tuning, and reinforcement-learning experimentation

Text Multimodal 128,000 ctx Image input Video input
View model →
◎
MiniMax

MiniMax M2

MiniMax M2

Coding agents, multi-step tool workflows, long-context codebase analysis, research automation, and self-hosted experimentation

Text Coding 196,608 ctx Tool use Streaming
View model →
◎
MiniMax

MiniMax M2.1

MiniMax M2

Multilingual software engineering, coding agents, tool-using workflows, application development, and cost-sensitive automation

Text Coding 204,800 ctx Tool use Streaming
View model →
◎
MiniMax

MiniMax M2.5

MiniMax M2

Coding agents, software engineering, search and browser agents, tool-calling workflows, office automation, and long-context technical work

Text Reasoning 204,800 ctx Tool use Web search Streaming
View model →
◎
MiniMax

MiniMax M2.5-highspeed

MiniMax M2.5

Low-latency coding assistants, software-engineering agents, tool-using workflows, search tasks, and long-context productivity automation

Text Coding 204,800 ctx Tool use Web search Structured output
View model →
◎
MiniMax

MiniMax M2.7

MiniMax M2

Agentic software engineering, repository-level coding, production debugging, complex tool workflows, office document automation, and long-context professional tasks

Text Coding 196,608 ctx Tool use
View model →
◎
MiniMax

MiniMax M2.7-highspeed

MiniMax M2.7

Low-latency coding assistants, software-engineering agents, tool-calling workflows, and interactive developer applications

Text Coding 204,800 ctx Tool use Streaming
View model →
◎
MiniMax

MiniMax M3

MiniMax M3

Long-context coding agents, autonomous tool-using workflows, multimodal document and video analysis, and private deployment

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
MiniMax

MiniMax-M2.1-highspeed

MiniMax M2.1

Low-latency coding assistants, multilingual software development, tool-using agents, long-horizon workflows, and interactive office automation

Text Coding 204,800 ctx Tool use Streaming
View model →
◎
Mistral AI

Ministral 3 14B

Ministral 3

Private assistants, local vision-language applications, multilingual workloads, document and image analysis, and cost-efficient agentic systems

Text Multimodal 262,144 ctx Image input Tool use Structured output
View model →
◎
Mistral AI

Ministral 3 3B

Ministral 3

Low-cost edge and local inference, image-aware assistants, document analysis, structured extraction, lightweight agents, task routing, and privacy-sensitive deployments.

Text Lightweight 256,000 ctx Image input Tool use Structured output
View model →
◎
Mistral AI

Ministral 3 8B

Ministral 3

Efficient edge and local inference, image understanding, document workflows, structured extraction, lightweight agents, and high-volume text generation.

Text Lightweight 256,000 ctx Image input Tool use Structured output
View model →
◎
Mistral AI

Mistral Large 3

Mistral Large

Long-context enterprise assistants, multilingual applications, image-aware document analysis, agentic workflows, coding, RAG and self-hosted sovereign deployments

Text Multimodal 256,000 ctx Image input Tool use Web search
View model →
◎
Mistral AI

Mistral Medium 3.5

Mistral Medium

Agentic coding, software engineering, long-context analysis, multimodal document workflows, structured outputs and multi-step tool use

Text Multimodal 256,000 ctx Image input Tool use Web search
View model →
◎
Mistral AI

Mistral Small 4

Mistral Small

Cost-efficient general chat, multimodal document analysis, coding, agentic workflows, and configurable reasoning

Text Multimodal 256,000 ctx Image input Tool use Web search
View model →
◎
Microsoft

model-router

Microsoft Foundry model-router

Enterprise applications with mixed-complexity workloads, model selection automation, agent workflows, cost optimization, and configurable quality-versus-latency trade-offs.

Text General Purpose 200,000 ctx Image input Tool use Streaming
View model →
◎
NVIDIA

MolMIM

MolMIM

Small-molecule generation, molecular embeddings, chemical-space exploration, lead optimization, and oracle-guided drug-design workflows

Text Other 128 ctx
View model →
◎
Ai2

Molmo 2-O 7B

Molmo 2

Open multimodal research, local image and video understanding, visual grounding, pointing, counting, captioning, and tracking

Text Multimodal 65,536 ctx Image input Video input
View model →
◎
Allen Institute for AI

Molmo-72B

Molmo

Self-hosted image understanding, visual question answering, captioning, image grounding, pointing, counting, and multimodal research.

Text Multimodal 4,096 ctx Image input Streaming
View model →
◎
Allen Institute for AI

Molmo-7B-D-0924

Molmo

Local image understanding, visual question answering, captioning, document and chart analysis, counting, pointing, and multimodal research

Text Multimodal 4,096 ctx Image input
View model →
◎
Allen Institute for AI

Molmo-7B-O-0924

Molmo

Local image-and-text understanding, visual question answering, document and chart analysis, image captioning, and multimodal research

Text Multimodal 4,096 ctx Image input
View model →
◎
Allen Institute for AI

Molmo2-4B

Molmo 2

Efficient local image and video understanding, visual grounding, pointing, captioning, counting, tracking, and multimodal research

Text Multimodal Image input Video input
View model →
◎
Allen Institute for AI

Molmo2-8B

Molmo 2

Open multimodal research, video question answering, visual grounding, pointing, counting, captioning, object tracking, document understanding, and robotics perception

Text Multimodal 36,864 ctx Image input Video input
View model →
◎
Allen Institute for AI

MolmoAct 7B-D Pretrain

MolmoAct

Robotic manipulation research, action reasoning, downstream mid-training, and reproducing zero-shot SimplerEnv experiments

Text Multimodal Image input
View model →
◎
Allen Institute for AI

MolmoAct 7B-O

MolmoAct

Open robotics research, visual action reasoning, robot-manipulation experiments, and fine-tuning on custom robot datasets

Text Multimodal Image input
View model →
◎
Allen Institute for AI

MolmoAct-7B-D

MolmoAct

Open research and downstream fine-tuning for vision-guided robotic manipulation, spatial reasoning, trajectory planning, and robot action prediction

Text Multimodal 4,096 ctx Image input
View model →
◎
Allen Institute for AI

MolmoAct2

MolmoAct2

Open robotics research, embodied visual reasoning, robot-policy fine-tuning, and manipulation tasks on supported or closely related hardware

Text Multimodal 16,384 ctx Image input
View model →
◎
Allen Institute for AI

MolmoAct2-Think

MolmoAct2

Depth-aware robot manipulation research, embodied reasoning, and fine-tuning vision-language-action policies for target robot embodiments.

Text Robotics Image input
View model →
◎
Allen Institute for AI

MolmoE-1B-0924

Molmo

Local image understanding, visual question answering, image captioning, document and chart analysis, counting, and experimentation with open multimodal models

Text Multimodal 4,096 ctx Image input
View model →
◎
Allen Institute for AI

MolmoPoint-8B

MolmoPoint

Open research and applications requiring image or video grounding, visual pointing, object localization, counting, tracking, spatial reasoning, and multimodal analysis.

Text Multimodal Image input Video input
View model →
◎
Allen Institute for AI

MolmoPoint-GUI-8B

MolmoPoint

GUI screenshot grounding, screen-element localization, computer-use perception, and research on pointing models

Text Multimodal 36,864 ctx Image input
View model →
◎
Allen Institute for AI

MolmoPoint-Vid-4B

MolmoPoint

Video object pointing, temporal grounding, object counting, tracking, and research on language-guided video understanding

Text Multimodal 35,200 ctx Video input
View model →
◎
Moonshot AI

Moonlight-16B-A3B

Moonlight

Self-hosted text generation, language-model research, efficient MoE inference, code and mathematics experimentation, and fine-tuning research

Text General Purpose 8,192 ctx
View model →
◎
Moonshot AI

Moonlight-16B-A3B-Instruct

Moonlight

Self-hosted instruction following, general text generation, research, experimentation, and cost-conscious local inference

Text General Purpose 8,192 ctx
View model →
◎
Databricks

MPT-7B

MPT

Self-hosted general-purpose text generation, experimentation, and fine-tuning with open model weights.

Text General Purpose 65,536 ctx Streaming
View model →
◎
Databricks

MPT-7B-Instruct

MPT

Short-form instruction following, local inference, open-model research, and domain fine-tuning

Text General Purpose 2,048 ctx
View model →
◎
Meta

Muse Glimmer

Muse Glimmer

Local agents, long-running tool workflows, coding assistants, multimodal document and screenshot understanding, private on-device inference, and model customization

Text Reasoning 131,072 ctx Image input Tool use Streaming
View model →
◎
Meta

Muse Spark 1.1

Muse Spark

Agentic workflows, coding agents, computer-use automation, multimodal document and media analysis, long-context reasoning, tool orchestration, and web-grounded applications.

Text Reasoning 1,048,576 ctx Image input Audio input Video input
View model →
◎
Meta

Muse Spark 1.2

Muse Spark

Long-horizon coding agents, repository-scale software engineering, multimodal code generation, debugging, refactoring, and tool-driven workflows

Text Coding 1,048,576 ctx Image input Audio input Video input
View model →
◎
Meta

Muse Spark 1.3

Muse Spark

Long-horizon coding agents, software engineering, browser and computer-use workflows, tool orchestration, large repositories, document analysis, and multimodal reasoning

Text Multimodal 1,048,576 ctx Image input Audio input Video input
View model →
◎
Meta AI

Muse Voice Transcribe 1.0

Muse Spark

Real-time speech-to-text, live captions, voice agents, meeting transcription, call intelligence, dictation, and speaker-aware transcription

Text Other Audio input Streaming
View model →
◎
Google DeepMind

Nano Banana

Gemini 2.5 Flash Image

Fast image generation, conversational image editing, image transformation, and high-volume visual workflows

Text Multimodal 65,536 ctx Image input
View model →
◎
Google DeepMind

Nano Banana 2

Gemini Image

Fast, high-volume image generation and editing, visual iteration, marketing assets, diagrams, infographics, localization, and applications requiring image-search grounding

Text Multimodal 131,072 ctx Image input Video input Web search
View model →
◎
Google DeepMind

Nano Banana 2 Lite

Gemini 3.1 Flash Lite Image

Fast, low-cost 1K image generation and editing, rapid visual prototyping, interactive applications, high-volume image variations, storyboarding, and lightweight creative workflows

Text Multimodal 65,536 ctx Image input Video input
View model →
◎
Google DeepMind

Nano Banana Pro

Gemini Image

Professional image generation and editing, complex compositions, product mockups, infographics, branded creative, multilingual localization, and high-fidelity visual prototyping

Text Multimodal 65,536 ctx Image input Web search
View model →
◎
NVIDIA

Nemotron 3.5 Content Safety

Nemotron Content Safety

Multilingual text-and-image moderation, LLM and VLM guardrails, response safety evaluation, and custom enterprise safety policies

Text Other 128,000 ctx Image input Streaming
View model →
◎
NVIDIA

Nemotron ASR Streaming

Nemotron ASR

Low-latency live speech-to-text, voice interfaces, live captions, and continuous audio streams

Text Other Audio input Streaming
View model →
◎
NVIDIA

Nemotron OCR v1

Nemotron OCR

English OCR, document ingestion, layout-aware text extraction, multimodal retrieval, RAG preprocessing, and enterprise document intelligence

Text Other Image input Structured output
View model →
◎
Cohere

North Micro Vision Instruct

North Micro Vision

Compact multilingual image understanding, OCR, documents, charts, visual question answering, and specialized fine-tuning

Text Multimodal 128,000 ctx Image input
View model →
◎
Cohere

North Mini Code

North

Agentic software engineering, repository-level code changes, terminal-based coding agents, code review, local inference, and private deployment.

Text Coding 256,000 ctx Image input Tool use Structured output
View model →
◎
Cohere

North Small Translate

North

Enterprise machine translation, multilingual documentation, localization, internal communications, safety procedures, and private or self-hosted translation workflows

Text Other 16,000 ctx Tool use Structured output Streaming
View model →
◎
NVIDIA

NVIDIA Conformer-CTC Large (en-US)

Conformer-CTC

Fast English speech-to-text transcription, NeMo experimentation, domain fine-tuning, and Riva-based ASR deployment

Text Other Audio input Streaming
View model →
◎
NVIDIA

NVIDIA Cosmos3-Nano-Reasoner

Cosmos 3

Physical-world video and image understanding, robotic perception, embodied-agent planning, spatial-temporal reasoning, and Physical AI research

Text Reasoning 256,000 ctx Image input Video input Structured output
View model →
◎
NVIDIA

NVIDIA Ising Calibration 1.5 31B

Ising Calibration

Quantum-computing calibration plot interpretation, QPU bring-up and retuning workflows, experiment diagnosis, parameter extraction, fit-quality assessment, and domain-specific calibration agents.

Text Multimodal 262,144 ctx Image input Streaming
View model →
◎
NVIDIA

NVIDIA Nemotron 3 Nano Omni 30B-A3B Reasoning

Nemotron 3 Nano Omni

Multimodal document intelligence, OCR, long-video and audio understanding, cross-modal reasoning, voice agents, and self-hosted enterprise inference

Text Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
NVIDIA

NVIDIA Nemotron 3 Super 120B-A12B

Nemotron 3

Agentic reasoning, long-context analysis, coding, tool-calling workflows, retrieval-augmented generation, collaborative agents, and high-volume enterprise inference.

Text Reasoning 1,000,000 ctx Tool use Streaming
View model →
◎
NVIDIA

NVIDIA Nemotron 3 Ultra 550B-A55B

Nemotron 3

Frontier reasoning, complex coding, long-context analysis, enterprise RAG, tool-using agents, and multi-agent workflows

Text Reasoning 1,000,000 ctx Tool use Streaming
View model →
◎
NVIDIA

NVIDIA Nemotron 3 VoiceChat

Nemotron 3 VoiceChat

Real-time full-duplex voice agents, interruptible conversational interfaces, speech-to-speech research, and NVIDIA GPU-based enterprise voice applications

Text Multimodal Audio input Streaming
View model →
◎
NVIDIA

NVIDIA Nemotron 3.5 Lightning 30B-A3B

Nemotron 3.5

High-volume agent execution, long-running AI agents, reasoning, coding assistance, RAG, chat, low-latency text generation, and domain customization

Text Reasoning 1,000,000 ctx Tool use Streaming
View model →
◎
NVIDIA

NVIDIA Nemotron OCR v2

Nemotron OCR

Multilingual OCR, scanned documents, forms, reports, charts, tables, image-based search, document ingestion, and retrieval-augmented generation preprocessing

Text Other Image input Structured output
View model →
◎
NVIDIA

NVIDIA Nemotron Parse 2.0

Nemotron Parse

Document OCR, layout analysis, table and chart extraction, retrieval pipelines, document indexing, and multimodal data curation

Text Multimodal Image input
View model →
◎
NVIDIA

NVIDIA Synthetic Video Detector

Synthetic Video Detector

AI-generated video detection, media authentication, digital forensics, content moderation, and media-integrity monitoring

Text Other Video input Streaming
View model →
◎
NVIDIA

NVIDIA-Ising-Calibration-1-35B-A3B

NVIDIA Ising Calibration

Quantum-computing calibration plot analysis, experiment interpretation, fit-quality assessment, parameter extraction, and technical research workflows

Text Multimodal 262,144 ctx Image input Streaming
View model →
◎
OpenAI

o1

o1

Complex reasoning, mathematics, science, coding analysis, visual reasoning, and high-accuracy multi-step tasks

Text Reasoning 200,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

o1 Preview

o1

Historically, difficult mathematics, science, coding, and other multi-step reasoning tasks requiring extended deliberation

Text Reasoning 128,000 ctx Tool use Structured output Streaming
View model →
◎
OpenAI

o1-mini

o1

Cost-sensitive mathematics, science, algorithmic programming, debugging, and text-only reasoning

Text Reasoning 128,000 ctx Streaming
View model →
◎
OpenAI

o1-pro

o1

Complex reasoning, difficult technical analysis, advanced programming, research workflows, and tasks where answer consistency matters more than latency or cost.

Text Reasoning 200,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

o3

o-series

Complex reasoning, advanced coding, mathematics, science, technical research, visual analysis and multi-step tool workflows

Text Reasoning 200,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

o3-deep-research

o3

Complex multi-step research, source synthesis, legal and scientific analysis, market research, and large-scale internal-data investigation

Text Reasoning 200,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

o3-mini

o3

Coding, mathematics, science, technical analysis, structured extraction, text-to-SQL, and multi-step reasoning

Text Reasoning 200,000 ctx Tool use Structured output Streaming
View model →
◎
OpenAI

o3-pro

o3

High-reliability reasoning, advanced mathematics, scientific analysis, complex coding, research, and multi-step professional work.

Text Reasoning 200,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

o4-mini

o4

Fast, cost-sensitive reasoning; coding; mathematics; visual analysis; structured extraction; high-volume tool-using agents

Text Reasoning 200,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

o4-mini-deep-research

o4

Complex multi-step research, source synthesis, market analysis, legal or scientific research, and long-form evidence-based reports.

Text Reasoning 200,000 ctx Image input Tool use Web search
View model →
◎
Mistral AI

OCR 3

OCR

High-volume document extraction, scanned forms, handwriting, invoices, complex tables, archival digitization, and document-to-knowledge pipelines.

Text Other Image input Structured output
View model →
◎
Mistral AI

OCR 4.0

Mistral OCR

High-volume OCR, structured document extraction, enterprise search, RAG ingestion, invoice processing, compliance workflows, and document automation

Text Other Image input Structured output
View model →
◎
Mistral AI

OCR 4.1

OCR

OCR, document parsing, structured extraction, enterprise search, RAG ingestion, invoice processing, and document AI workflows

Text Other Image input Structured output
View model →
◎
Allen Institute for AI

Olmo 3 32B Base

Olmo 3

Open-weight research, continued pretraining, fine-tuning, programming, mathematics, reading comprehension, and long-context language-model experiments

Text General Purpose 65,536 ctx Streaming
View model →
◎
Ai2

Olmo 3 32B Instruct

Olmo 3

Open-weight chat, tool-using assistants, multi-turn dialogue, synthetic data generation, self-hosting, and model research

Text General Purpose 65,536 ctx Tool use Streaming
View model →
◎
Allen Institute for AI

Olmo 3 32B Think

Olmo 3

Open-model reasoning research, mathematics, coding, long-context analysis, and self-hosted deployment

Text Reasoning 65,536 ctx Streaming
View model →
◎
Allen Institute for AI

Olmo 3 7B Base

Olmo 3

Continued pretraining, supervised fine-tuning, reinforcement-learning research, language-model research, custom deployment, and applications requiring a transparent open training pipeline.

Text General Purpose 65,536 ctx Streaming
View model →
◎
Allen Institute for AI

Olmo 3 7B Instruct

Olmo 3

Open, self-hosted chat assistants, instruction following, local coding help, tool-calling agents, and model research

Text General Purpose 65,536 ctx Tool use Streaming
View model →
◎
Allen Institute for AI

Olmo 3 7B Think

Olmo 3

Open research, mathematical reasoning, code generation, logic, long-context experimentation, and local self-hosted inference

Text Reasoning 65,536 ctx Streaming
View model →
◎
Allen Institute for AI

Olmo 3.1 7B RL-Zero Code

Olmo 3.1

Open coding-model research, reinforcement learning from verifiable rewards, code-generation experiments, automated evaluation, and fine-tuning

Text Coding 65,536 ctx
View model →
◎
Allen Institute for AI

Olmo 3.1 7B RL-Zero Math

Olmo 3

Math-focused RLVR research, reinforcement-learning experiments, verifiable reward design, open-model evaluation, and further fine-tuning

Text Reasoning 65,536 ctx
View model →
◎
Allen Institute for AI

Olmo 3.1 Instruct 32B

Olmo 3.1

Self-hosted chat assistants, instruction following, tool-use agents, synthetic data, and open-model research

Text General Purpose 65,536 ctx Tool use
View model →
◎
Allen Institute for AI

Olmo 3.1 Think 32B

Olmo 3

Mathematical reasoning, coding, complex multi-step tasks, open-model research, local deployment, and reinforcement-learning research

Text Reasoning 65,536 ctx Streaming
View model →
◎
Allen Institute for AI

Olmo Hybrid 7B

Olmo Hybrid

Open-weight language-model research, long-context text generation, local deployment, continued pretraining, and fine-tuning

Text General Purpose 65,536 ctx Streaming
View model →
◎
Allen Institute for AI

OLMo-2-0325-32B

OLMo 2

Open research, reproducible language-model experiments, local deployment, English text generation, and fine-tuning

Text General Purpose 4,096 ctx Streaming
View model →
◎
Allen Institute for AI

OLMo-2-1124-13B

OLMo 2

Fully open language-model research, local inference, reproducible training experiments, evaluation, and downstream fine-tuning

Text General Purpose 4,096 ctx Streaming
View model →
◎
Ai2

OLMo-2-1124-7B

OLMo 2

Local inference, open-model research, benchmarking, continued pretraining, and task-specific fine-tuning.

Text General Purpose 4,096 ctx
View model →
◎
Allen Institute for AI

OLMoASR-base.en

OLMoASR

English audio transcription, captions, meetings, lectures, podcasts, accessibility applications, and open ASR research

Text Other Audio input Streaming
View model →
◎
Allen Institute for AI

OLMoASR-large.en-v1

OLMoASR

Open English speech-to-text transcription, captioning, meetings, lectures, calls, podcasts, and local ASR research

Text Other Audio input Streaming
View model →
◎
Ai2

OLMoASR-large.en-v2

OLMoASR

Open, self-hosted English transcription for meetings, lectures, calls, podcasts, accessibility, and speech-recognition research

Text Other Audio input Streaming
View model →
◎
Ai2

OLMoASR-medium.en

OLMoASR

English short- and long-form speech transcription, meeting and call transcription, lecture captioning, podcast processing, broadcast transcription, and speech analytics

Text Other Audio input
View model →
◎
Allen Institute for AI

OLMoASR-small.en

OLMoASR

Open English speech transcription, meeting and podcast transcription, captioning, timestamped audio indexing, and ASR research

Text Other Audio input
View model →
◎
Allen Institute for AI

OLMoASR-tiny.en

OLMoASR

Efficient self-hosted English transcription, speech research, edge-oriented experiments, and applications where a small open ASR checkpoint is preferred.

Text Lightweight Audio input
View model →
◎
Allen Institute for AI

OLMoE-1B-7B-0924

OLMoE

Open-model research, efficient local text generation, language-model evaluation, and fine-tuning

Text Lightweight 4,096 ctx Streaming
View model →
◎
OpenAI

omni-moderation

omni-moderation

Text and image safety classification, content filtering, AI-output screening, policy enforcement, and human-review routing

Text Moderation Image input
View model →
◎
OpenAI

omni-moderation-latest

omni-moderation

Text and image safety classification, content filtering, moderation queues, policy enforcement, and generated-content screening

Text Other Image input
View model →
◎
PaddlePaddle

PaddleOCR-VL-1.5

PaddleOCR-VL

Multilingual OCR and structured parsing of complex documents, including tables, formulas, charts, seals, scanned pages, warped documents, and screen photographs

Text Multimodal 131,072 ctx Image input
View model →
◎
NVIDIA

Parakeet CTC 0.6B

Parakeet CTC

Local English speech transcription, offline ASR, long-form audio processing, and custom ASR fine-tuning

Text Other Audio input
View model →
◎
NVIDIA

Parakeet CTC 0.6B Taiwanese Mandarin-English ASR

Parakeet CTC

Taiwanese Mandarin speech-to-text, Mandarin-English code-switching, live captions, voice interfaces, and NVIDIA GPU-hosted streaming or offline transcription.

Text Other Audio input Streaming
View model →
◎
NVIDIA

Parakeet CTC 0.6B zh-CN

Parakeet CTC

Mandarin-English automatic speech recognition, code-switched transcription, streaming transcription, and self-hosted enterprise speech-to-text

Text Other Audio input Streaming
View model →
◎
NVIDIA

Parakeet CTC 1.1B

Parakeet

English speech-to-text transcription, local ASR, offline processing, RAG audio extraction, and customized NeMo deployments

Text Other Audio input Streaming
View model →
◎
NVIDIA

Parakeet TDT 0.6B v2

Parakeet TDT

High-speed English speech transcription, captions, meeting transcription, audio search, and timestamped media workflows

Text Other Audio input Streaming
View model →
◎
Cohere

Parse v5.0

Parse

High-volume enterprise document parsing, table and form extraction, search indexing, RAG ingestion, and document context for AI agents

Text Multimodal 8,192 ctx Image input
View model →
◎
Aleph Alpha

Pharia-1-LLM-7B-control

Pharia-1-LLM-7B

Concise multilingual generation, summarization, extraction, domain-specific text workflows, and self-hosted deployment

Text General Purpose 8,192 ctx Streaming
View model →
◎
Aleph Alpha

Pharia-1-LLM-7B-control-aligned

Pharia-1-LLM-7B

Multilingual text generation, classification, summarization, question answering, engineering and automotive applications, and safety-conscious research deployments

Text General Purpose 8,192 ctx
View model →
◎
Microsoft

Phi-3-medium-128k-instruct

Phi-3

Long-context chat, retrieval-augmented generation, document summarization, coding, mathematics, reasoning, and self-hosted text-generation applications

Text General Purpose 131,072 ctx Streaming
View model →
◎
Microsoft

Phi-3-medium-4k-instruct

Phi-3

Local and private text generation, coding assistance, mathematics, reasoning, summarization, and latency-sensitive applications

Text General Purpose 4,096 ctx
View model →
◎
Microsoft

Phi-3-mini-128k-instruct

Phi-3

Long-document analysis, local assistants, code and math tasks, retrieval-augmented generation, private or offline inference, and resource-constrained deployments.

Text Lightweight 131,072 ctx Streaming
View model →
◎
Microsoft

Phi-3-mini-4k-instruct

Phi-3

Local or low-latency text generation, lightweight assistants, mathematics, coding, summarization, and resource-constrained deployments

Text Lightweight 4,096 ctx Streaming
View model →
◎
Microsoft

Phi-3-small-128k-instruct

Phi-3

Long-context text generation, document analysis, summarization, local assistants, coding support, mathematics, and cost-sensitive self-hosted applications

Text Lightweight 131,072 ctx Streaming
View model →
◎
Microsoft

Phi-3-small-8k-instruct

Phi-3

Local and private text generation, instruction following, code and mathematics assistance, summarization, extraction, and resource-constrained deployments

Text Lightweight 8,192 ctx Streaming
View model →
◎
Microsoft

Phi-3.5-mini-instruct

Phi-3

Local multilingual chat, summarization, document analysis, coding assistance, long-context retrieval, and resource-constrained deployments

Text Lightweight 131,072 ctx Streaming
View model →
◎
Microsoft

Phi-3.5-MoE-instruct

Phi-3.5

Long-context multilingual assistants, coding, mathematics, reasoning, retrieval-augmented generation, and controlled local deployment.

Text General Purpose 131,072 ctx Streaming
View model →
◎
Microsoft

Phi-3.5-vision-instruct

Phi-3.5

Lightweight image and text reasoning, OCR, charts, tables, diagrams, documents, screenshots, and multi-image comparison

Text Multimodal 131,072 ctx Image input Streaming
View model →
◎
Microsoft

Phi-4

Phi-4

Efficient local or hosted text generation, STEM reasoning, mathematics, coding assistance, technical question answering, and research on small language models

Text General Purpose 16,384 ctx
View model →
◎
Microsoft

Phi-4-mini-flash-reasoning

Phi-4

Efficient mathematical reasoning, math tutoring, automated assessment, lightweight reasoning agents, edge deployment, mobile applications, and latency-sensitive local inference

Text Reasoning 65,536 ctx Streaming
View model →
◎
Microsoft

Phi-4-mini-instruct

Phi-4

Efficient local or cloud text generation, multilingual applications, mathematics, coding, reasoning, retrieval-augmented generation, and edge deployment

Text Lightweight 131,072 ctx Tool use Streaming
View model →
◎
Microsoft

Phi-4-mini-reasoning

Phi-4

Mathematical reasoning, educational tutoring, formal proof assistance, symbolic computation, local inference, edge deployment, and latency-sensitive applications

Text Reasoning 128,000 ctx
View model →
◎
Microsoft

Phi-4-multimodal-instruct

Phi-4

Compact multimodal assistants, OCR, document and chart analysis, image question answering, speech recognition, speech translation, audio summarization, and private or local deployment

Text Multimodal 131,072 ctx Image input Audio input Streaming
View model →
◎
Microsoft

Phi-4-reasoning

Phi-4

Mathematical reasoning, coding, science, logic, algorithmic problem solving, research, and resource-constrained local deployments

Text Reasoning 32,768 ctx Streaming
View model →
◎
Microsoft

Phi-4-reasoning-plus

Phi-4

Mathematical reasoning, scientific problem solving, coding assistance, algorithmic tasks, local deployment, and reasoning-model research

Text Reasoning 32,768 ctx
View model →
◎
Baidu

PP-StructureV3

PP-Structure

Document parsing, OCR, layout analysis, table extraction, and structured document understanding

Text Multimodal Image input Structured output
View model →
◎
Baidu

Qianfan-Agent-Intent-32K

Qianfan Agent

Enterprise agent workflows, intent recognition, instruction routing, and tool-calling tasks

Text Other 32,768 ctx Tool use
View model →
◎
Baidu

Qianfan-Agent-Lite-128K

Qianfan-Agent-Lite

Long-context Agent planning, task decomposition, component selection, and function-calling workflows.

Text Lightweight 128,000 ctx Tool use
View model →
◎
Baidu

Qianfan-Agent-Lite-8K

Qianfan Agent

Fast enterprise question answering, lightweight agent planning, component selection, and streaming text applications

Text Lightweight 8,192 ctx Tool use Streaming
View model →
◎
Baidu

Qianfan-Agent-Speed-32K

Qianfan-Agent

Fast enterprise question answering, summarization, workflow nodes, agent response generation, and text processing with a 32K context window

Text Lightweight 32,768 ctx Tool use Streaming
View model →
◎
Qwen

Qwen-AgentWorld-35B-A3B

Qwen-AgentWorld

Agent-environment simulation, tool-interaction modeling, terminal and software-engineering trajectories, and research on language world models

Text Other 262,144 ctx Tool use Streaming
View model →
◎
Alibaba Cloud

qwen-audio-3.0-realtime-flash

Qwen-Audio-3.0-Realtime

Low-latency voice assistants, real-time customer service, duplex speech conversations, interactive voice agents, and applications requiring streaming audio responses.

Text Multimodal 40,960 ctx Audio input Tool use Web search
View model →
◎
Alibaba Cloud Model Studio

qwen-audio-3.0-realtime-plus

Qwen-Audio

Low-latency duplex voice assistants, real-time customer service, AI companions, and streamed speech-to-speech applications

Text Multimodal 40,960 ctx Audio input Tool use Web search
View model →
◎
Alibaba Cloud

Qwen-Audio-3.1-ASR-Flash-Filetrans

Qwen-Audio 3.1 ASR

Long-form offline transcription of meetings, interviews, calls, media files, and multilingual or dialect-rich recordings

Text Other 8,192 ctx Audio input
View model →
◎
Alibaba Cloud Model Studio

Qwen-Audio-3.1-ASR-Flash-Streaming

Qwen-Audio-3.1-ASR

Real-time multilingual speech transcription, live captions, meeting transcription, voice interfaces, streaming subtitles, and Chinese-dialect recognition.

Text Speech Recognition 8,192 ctx Audio input Streaming
View model →
◎
Alibaba Cloud Model Studio

qwen-audio-3.1-realtime-plus

Qwen-Audio

Low-latency voice assistants, customer service, AI companions, full-duplex spoken interaction, and voice applications using tools or cloned voices.

Text Multimodal 262,144 ctx Audio input Tool use Web search
View model →
◎
Alibaba Qwen

Qwen-Drive-1.0

Qwen-Drive

Autonomous-driving research, driving-scene VQA, 3D BEV perception, trajectory prediction, and embodied-AI experimentation

Text Multimodal Image input Streaming
View model →
◎
Alibaba Cloud Model Studio

Qwen-Flash

Qwen3

Fast, high-volume text generation; long-context analysis; summarization; extraction; structured outputs; and applications needing optional reasoning.

Text Lightweight 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Qwen

qwen-flash-character

Qwen Character

Low-latency character dialogue, virtual companions, game NPCs, role-playing applications, IP character replication, and conversational smart devices

Text Lightweight 32,768 ctx Web search Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen-Plus

Qwen-Plus

General-purpose text generation, long-context analysis, multilingual applications, structured business workflows, function-calling agents, and applications that need optional reasoning mode.

Text General Purpose 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

Qwen-VL-Max

Qwen-VL

Complex image and video understanding, document analysis, chart interpretation, visual question answering, and structured extraction

Text Multimodal 131,072 ctx Image input Video input Structured output
View model →
◎
Qwen

Qwen2.5-14B-Instruct

Qwen2.5

Self-hosted multilingual assistants, document processing, RAG, coding support, structured extraction, and cost-conscious production deployments.

Text General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen2.5-32B-Instruct

Qwen2.5

Self-hosted assistants, multilingual text generation, long-document processing, coding support, structured extraction, RAG, and agent applications

Text General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen2.5-72B-Instruct

Qwen2.5

Self-hosted multilingual assistants, coding, mathematics, document analysis, structured text generation, and long-context workloads

Text General Purpose 131,072 ctx Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen2.5-7B-Instruct

Qwen2.5

Self-hosted chat assistants, multilingual text generation, coding and mathematics assistance, long-context document work, structured text generation, and cost-sensitive private deployments

Text General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen2.5-Omni-7B

Qwen2.5-Omni

Multimodal assistants, audio and video understanding, visual question answering, voice interaction, speech instruction following, and local multimodal AI research

Text Multimodal 32,768 ctx Image input Audio input Video input
View model →
◎
Qwen

Qwen2.5-VL-72B-Instruct

Qwen2.5-VL

High-quality image, document, chart, screenshot, OCR, visual-grounding, and video analysis; multimodal agents and self-hosted experimentation

Text Multimodal 131,072 ctx Image input Video input Tool use
View model →
◎
Qwen

Qwen2.5-VL-7B-Instruct

Qwen2.5-VL

Local or self-hosted image and video understanding, OCR, document extraction, chart and diagram analysis, visual question answering, visual grounding, and multimodal research

Text Multimodal 32,768 ctx Image input Video input Structured output
View model →
◎
Qwen

Qwen3-0.6B

Qwen3

Lightweight local assistants, offline prototypes, embedded experimentation, education, simple text generation, and resource-constrained deployments

Text Lightweight 32,768 ctx Tool use Streaming
View model →
◎
Alibaba Cloud

Qwen3-1.7B

Qwen3

Efficient local inference, edge applications, multilingual chat, lightweight reasoning, coding assistance, and tool-enabled agents

Text Lightweight 32,768 ctx Tool use Streaming
View model →
◎
Qwen

Qwen3-14B

Qwen3

Local deployment, multilingual assistants, reasoning, mathematics, coding, structured text generation, research, and cost-sensitive agent workflows.

Text General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen3-235B-A22B

Qwen3

Complex reasoning, mathematics, software development, multilingual applications, function calling, agentic workflows, research and self-hosted open-weight deployment

Text Reasoning 131,072 ctx Tool use Structured output Streaming
View model →
◎
Qwen

Qwen3-30B-A3B

Qwen3

Self-hosted assistants, coding, mathematical and logical reasoning, multilingual applications, long-context document processing, and agentic tool-use systems

Text Reasoning 131,072 ctx Tool use Streaming
View model →
◎
Qwen

Qwen3-32B

Qwen3

Self-hosted reasoning assistants, coding agents, mathematics, multilingual applications, structured text generation, and tool-calling workflows

Text Reasoning 256,000 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen3-4B

Qwen3

Local and self-hosted chat, compact reasoning, coding assistance, multilingual applications, retrieval-augmented generation, and lightweight tool-using agents

Text General Purpose 32,768 ctx Tool use Streaming
View model →
◎
Qwen

Qwen3-8B

Qwen3

Cost-efficient reasoning, coding, multilingual assistants, local deployment, structured text generation, tool-enabled agents, and fine-tuned applications

Text General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Qwen

Qwen3-ASR-0.6B

Qwen3-ASR

Low-cost multilingual speech-to-text, language identification, offline transcription, real-time streaming ASR, and high-throughput deployments

Text Other 65,536 ctx Audio input Streaming
View model →
◎
Qwen

Qwen3-ASR-1.7B

Qwen3-ASR

Multilingual speech transcription, language identification, long-audio processing, and self-hosted or streaming ASR applications

Text Other Audio input Streaming
View model →
◎
Alibaba Cloud

Qwen3-Coder-Next

Qwen3-Coder

Repository-scale coding agents, code generation, code completion, debugging, refactoring, terminal workflows, and cost-sensitive self-hosted deployments.

Text Coding 262,144 ctx Tool use Streaming
View model →
◎
Alibaba Cloud Model Studio

Qwen3-Coder-Plus

Qwen3-Coder

Large-codebase analysis, code generation, refactoring, debugging, documentation and long-context coding-agent workflows

Text Coding 1,000,000 ctx Tool use Streaming
View model →
◎
Alibaba Cloud

Qwen3-LiveTranslate-Flash

Qwen3-LiveTranslate

Streaming translation of recorded or uploaded audio and video, multilingual subtitles, translated voice tracks, and applications requiring translated text or synthesized speech.

Text Other 53,248 ctx Audio input Video input Streaming
View model →
◎
Alibaba Cloud

Qwen3-LiveTranslate-Flash-Realtime

Qwen3-LiveTranslate

Real-time multilingual speech interpretation, live voice translation, conference translation, streaming media, and audiovisual translation with text or synthesized speech output

Text Other 53,248 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3-Max

Qwen3-Max

Complex reasoning, coding assistance, web-grounded agents, function calling, structured extraction, and long-context text analysis

Text Reasoning 262,144 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

Qwen3-VL-235B-A22B-Instruct

Qwen3-VL

High-quality image and video understanding, OCR, document intelligence, visual coding, spatial reasoning, long-context multimodal analysis and visual-agent applications

Text Multimodal 131,072 ctx Image input Video input Tool use
View model →
◎
Qwen

Qwen3-VL-2B-Instruct

Qwen3-VL

Local visual assistants, OCR, document and chart analysis, image question answering, lightweight video understanding, and multimodal prototyping

Text Multimodal 256,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3-VL-30B-A3B-Instruct

Qwen3-VL

Image and video understanding, OCR, document analysis, spatial reasoning, visual coding, long-context multimodal tasks, and visual-agent applications

Text Multimodal 256,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3-VL-32B-Instruct

Qwen3-VL

Document intelligence, OCR, image and video understanding, spatial reasoning, visual coding, and visual-agent applications

Text Multimodal 131,072 ctx Image input Video input Tool use
View model →
◎
Qwen

Qwen3-VL-4B-Instruct

Qwen3-VL

Local image and video understanding, OCR, document extraction, visual question answering, visual coding, and lightweight multimodal agents

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3-VL-8B-Instruct

Qwen3-VL

Local or hosted image and video understanding, OCR, document extraction, visual question answering, spatial reasoning, screenshot analysis, multimodal agents, and structured data extraction

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-122B-A10B

Qwen3.5

Advanced multimodal reasoning, image and video understanding, coding, long-context analysis, document and chart interpretation, function-calling agents, and web-grounded workflows.

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.5-27B

Qwen3.5

Long-context multimodal analysis, document understanding, video and image interpretation, general reasoning, coding, tool-enabled assistants, and self-hosted deployment

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.5-35B-A3B

Qwen3.5

Efficient multimodal assistants, coding, reasoning, long-context analysis, local deployment, and tool-using agents

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-397B-A17B

Qwen3.5

Advanced multimodal reasoning, coding, video and image understanding, long-context analysis, tool-using agents, and self-hosted open-weight deployments

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.5-Flash

Qwen3.5

Fast long-context text, image and video understanding; structured extraction; tool-enabled agents; web-grounded applications; and high-volume multimodal workloads.

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Flash

Qwen3.5-Omni

Fast multimodal analysis, long audio understanding, audiovisual question answering, voice assistants and spoken-response applications

Text Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Flash-Realtime

Qwen3.5-Omni

Low-latency voice assistants, speech-to-speech applications, realtime multimedia analysis, interactive agents, and multimodal conversations.

Text Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud

Qwen3.5-Omni-Plus

Qwen3.5-Omni

Multilingual voice assistants, speech-enabled multimodal applications, audio-visual analysis, spoken explanations, accessibility tools, and interactive media workflows.

Text Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-Omni-Plus-Realtime

Qwen3.5-Omni

Real-time voice assistants, speech-to-speech applications, multimodal customer service, visual conversational agents, live multimedia analysis, and interactive applications requiring controllable speech output.

Text Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud

Qwen3.5-Plus

Qwen3.5

Long-context reasoning, multimodal document and video analysis, coding, structured enterprise automation, function-calling agents, and web-grounded research.

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.6-27B

Qwen3.6

Coding agents, repository-level software engineering, visual document analysis, video understanding, STEM reasoning, long-context assistants, and self-hosted multimodal applications

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Qwen

Qwen3.6-35B-A3B

Qwen3.6

Agentic coding, long-context software engineering, multimodal analysis, tool-using agents, and self-hosted deployments

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.6-Flash

Qwen3.6

Fast multimodal assistants, coding agents, visual document analysis, video understanding, tool-using workflows, object localization, and large-context applications

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.6-Max-Preview

Qwen3.6

Advanced coding agents, front-end development, long-context analysis, structured API workflows, and text generation with web search

Text General Purpose 262,144 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

Qwen3.6-Plus

Qwen3.6

Long-context multimodal analysis, agentic coding, OCR, object localization, frontend development, visual reasoning, and tool-enabled enterprise assistants

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.7-Flash

Qwen3.7

Fast multimodal agents, visual coding, tool-use workflows, search agents, and long-context document or screen analysis

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.7-Plus

Qwen3.7

Long-context reasoning, multimodal document and video analysis, coding, tool-using agents, structured extraction, and enterprise productivity workflows

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.8-2.4T-A95B

Qwen3.8

Advanced reasoning, coding, scientific and professional research, long-context analysis, and long-horizon agent workflows

Text Reasoning 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

Qwen3.8-27B

Qwen3.8

Coding assistants, repository analysis, visual document workflows, long-context research, office automation, multimodal agents, and tool-using applications

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.8-Flash

Qwen3.8

Fast long-context reasoning, coding assistance, visual document and chart analysis, video understanding, function-calling agents, and high-concurrency applications

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba

Qwen3.8-Flash-Next

Qwen3.8

High-volume coding, agentic workflows, long-context document and codebase analysis, office automation, and multimodal text-image-video understanding

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.8-LiveTranslate-Flash-Realtime

Qwen3.8-LiveTranslate

Real-time speech translation, multilingual meetings, live interpretation, translated voice communication, and audiovisual translation with low latency.

Text Multimodal 53,248 ctx Image input Audio input Streaming
View model →
◎
Alibaba Cloud

Qwen3.8-Max

Qwen3.8

Complex coding, autonomous software engineering, long-horizon agent workflows, professional document analysis, visual reasoning, long videos, and demanding research tasks.

Text Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.8-Omni-Flash

Qwen3.8-Omni

Long-form audio and video understanding, multimedia analysis, audio-visual agents, content summarization, and tool-using workflows

Text Multimodal 1,000,000 ctx Image input Audio input Video input
View model →
◎
Alibaba Cloud Model Studio

Qwen3.8-Omni-Flash-Realtime

Qwen3.8-Omni

Real-time voice assistants, speech-to-speech applications, interactive video agents, live media analysis, multimodal customer service, meeting and collaboration interfaces, and applications requiring tool or MCP integration.

Text Multimodal 196,608 ctx Audio input Video input Tool use
View model →
◎
Reka

Reka Core

Reka Core

Complex multimodal analysis, long-context document understanding, image and video question answering, audio-aware workflows, coding, and enterprise applications requiring broad input support.

Text Multimodal 128,000 ctx Image input Audio input Video input
View model →
◎
Reka

Reka Edge

Reka Edge

Low-latency image and video understanding, object detection, physical AI, robotics, edge deployment, and on-device visual assistants

Text Multimodal 16,384 ctx Image input Video input Streaming
View model →
◎
Reka

Reka Flash

Reka Flash

Fast multimodal applications, document and image analysis, short-video understanding, structured extraction, multilingual chat, coding, and tool-using agents

Text Multimodal 128,000 ctx Image input Audio input Video input
View model →
◎
Reka AI

Reka Flash 3

Reka Flash

Low-latency reasoning, coding assistance, function calling, local deployment, on-device applications and cost-sensitive inference

Text Reasoning 65,536 ctx Tool use Streaming
View model →
◎
Reka AI

Reka Flash 3.1

Reka Flash

Coding, mathematical reasoning, local inference, research experimentation, and fine-tuning for agentic workflows

Text Reasoning 32,768 ctx
View model →
◎
Cohere

Rerank 4 Pro

Rerank 4.0

High-quality multilingual reranking for enterprise search, RAG pipelines, semantic retrieval, and semi-structured document ranking.

Text Other 32,768 ctx
View model →
◎
NVIDIA

Riva Translate 1.6b

Megatron NMT

Low-latency multilingual text translation, speech-translation pipelines, and self-hosted NVIDIA GPU deployments

Text Other
View model →
◎
NVIDIA

Riva-Translate-4B-Instruct-v1.1

Riva Translate

Multilingual sentence and document translation, localization, marketing content, and developer translation workflows

Text Other 8,192 ctx
View model →
◎
NVIDIA

Riva-Translate-4B-Instruct-v2

Riva Translate

Local or self-hosted multilingual sentence and document translation across English and 36 non-English languages

Text Other 8,192 ctx
View model →
◎
Meta

SeamlessM4T-Large v2

SeamlessM4T

Multilingual automatic speech recognition, speech-to-text translation, text translation, text-to-speech translation, and speech-to-speech translation

Text Multimodal Audio input
View model →
◎
ByteDance Seed

Seed Diffusion Preview

Seed Diffusion

High-throughput code generation research, diffusion-language-model evaluation, and code-editing experiments

Text Coding
View model →
◎
ByteDance Seed

Seed-OSS-36B-Base

Seed-OSS

Self-hosted general language modeling, long-context research, reasoning experiments, coding assistance, summarization, and foundation-model fine-tuning

Text General Purpose 524,288 ctx Streaming
View model →
◎
ByteDance Seed

Seed-OSS-36B-Base-woSyn

Seed-OSS-36B

Research, custom post-training, long-context processing, coding experiments, and self-hosted foundation-model deployment

Text General Purpose 512,000 ctx Tool use
View model →
◎
ByteDance Seed

Seed-OSS-36B-Instruct

Seed-OSS

Self-hosted long-context reasoning, coding, agentic tool use, research, question answering, summarization, and general text generation

Text Reasoning 524,288 ctx Tool use Streaming
View model →
◎
ByteDance Seed

Seed1.5 (Doubao-1.5-pro)

Seed1.5

General-purpose Chinese and multilingual assistance, coding, reasoning, image and document understanding, and voice-interaction applications.

Text Multimodal 32,768 ctx Image input Audio input Tool use
View model →
◎
ByteDance Seed

Seed1.5-VL

Seed1.5

Visual reasoning, image and video understanding, OCR, visual grounding, GUI-agent research, gameplay analysis, and multimodal benchmark evaluation

Text Multimodal 131,072 ctx Image input Video input Tool use
View model →
◎
ByteDance Seed

Seed1.6

Seed1.6

Multimodal document analysis, visual question answering, coding, mathematics, general reasoning, long-context analysis and adaptive-thinking applications

Text Multimodal 256,000 ctx Image input Tool use Structured output
View model →
◎
ByteDance

Seed1.6-Thinking

Seed1.6

Deep reasoning, coding, mathematics, logical analysis, document understanding, and visual reasoning over images or videos

Text Reasoning 256,000 ctx Image input Video input Tool use
View model →
◎
ByteDance

Seed1.8

Seed

Multimodal agent workflows, search and information retrieval, coding agents, GUI interaction, image and video understanding, complex instruction following, and long-context business tasks.

Text Multimodal 256,000 ctx Image input Video input Tool use
View model →
◎
ByteDance

Seed2.0 Lite

Seed2.0

Cost-conscious production applications requiring long-context multimodal understanding, document and video analysis, coding assistance, tool use, GUI automation, and structured extraction.

Text Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
ByteDance Seed

Seed2.0 Mini

Seed2.0

High-concurrency inference, batch generation, classification, extraction, summarization, and cost-sensitive multimodal workloads

Text Lightweight 256,000 ctx Image input Video input
View model →
◎
ByteDance Seed

Seed2.0 Pro

Seed2.0

Complex multimodal reasoning, long-chain agent workflows, visual and video analysis, document understanding, scientific research support, coding, and enterprise automation.

Text Multimodal 200,000 ctx Image input Audio input Video input
View model →
◎
ByteDance Seed

Seed2.1 Pro

Seed2.1

Complex agent workflows, high-value office and research tasks, long-horizon coding, document and visual analysis, video understanding, and tool-enabled productivity automation.

Text General Purpose Image input Video input Tool use
View model →
◎
ByteDance Seed

Seed2.1 Turbo

Seed2.1

Fast multimodal agents, coding assistants, document and video analysis, tool-calling workflows, structured data extraction, and cost-sensitive production applications

Text Multimodal 256,000 ctx Image input Video input Tool use
View model →
◎
ByteDance Seed

SeedRealtime

SeedRealtime

Real-time audio-visual assistants, scene-aware guidance, live explanation, interactive learning, accessibility, and proactive multimodal collaboration

Text Multimodal Image input Audio input Video input
View model →
◎
SenseTime

SenseNova 6.8 Flash Lite

SenseNova 6.8

Long-horizon multimodal agents, data analysis, deep research, complex information presentation, office automation, and tool-driven workflows

Text Multimodal Image input Video input Tool use
View model →
◎
SenseTime

SenseNova U1

SenseNova U1

Open-source visual understanding, image generation, image editing, infographic creation, visual reasoning, and continuous image-text workflows.

Text Multimodal Image input
View model →
◎
SenseTime

SenseNova-MARS-32B

SenseNova-MARS

Visual deep-search, fine-grained image understanding, multimodal agent research, and self-hosted tool-using vision-language applications

Text Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
SenseNova

SenseNova-MARS-8B

SenseNova-MARS

Visual question answering, high-resolution image understanding, multimodal search, image-grounded research, and tool-assisted agentic reasoning

Text Multimodal 262,144 ctx Image input Tool use Web search
View model →
◎
SenseNova

SenseNova-SI-1.5-InternVL3-8B

SenseNova-SI

Spatial intelligence research, image-text question answering, visual reasoning, 3D scene analysis, and solid-geometry problems

Text Multimodal 32,768 ctx Image input
View model →
◎
SenseTime

SenseNova-U1.5-8B-MoT

SenseNova-U1.5

Native image generation, high-resolution visual creation, image editing, infographic and layout generation, visual understanding, and multimodal research or creative workflows

Text Multimodal Image input
View model →
◎
SenseTime

SenseNova-Vision-7B-MoT

SenseNova-Vision

Unified computer-vision research, detection, OCR, segmentation, depth and normal estimation, visual grounding, and multi-view geometry

Text Multimodal Image input Structured output
View model →
◎
Mistral AI

Shieldstral 1.0

Shieldstral

Policy-based moderation of prompts, model responses, refusals, text, images, and text-image combinations

Text Other 32,768 ctx Image input Streaming
View model →
◎
StepFun

Step 3.5 Flash

Step 3

Coding agents, software engineering, long-context reasoning, tool-using agents, private local inference, and work-centric automation

Text Reasoning 262,144 ctx Tool use Streaming
View model →
◎
StepFun

Step 3.7 Flash

Step 3.7

High-throughput coding agents, visual document and UI understanding, search-heavy research workflows, long-context analysis, and tool-using autonomous agents

Text Multimodal 256,000 ctx Image input Tool use Web search
View model →
◎
StepFun

Step 5 Preview

Step 5

Long-running agentic workflows, software engineering, coding, professional knowledge work, financial analysis, research, and vision-assisted tasks

Text Reasoning 1,000,000 ctx Image input Tool use
View model →
◎
StepFun

Step Edge

Step Edge

On-device visual assistants, GUI grounding, screen understanding, spatial reasoning, automotive interaction, and low-latency local agent workflows

Text Multimodal 131,072 ctx Image input Video input Tool use
View model →
◎
StepFun

Step Edge Audio

Step Edge

On-device speech recognition, audio understanding, voice assistants, in-vehicle interaction, privacy-sensitive audio processing, and low-latency edge AI.

Text Multimodal Audio input
View model →
◎
StepFun

StepAudio 3 ASR Max

StepAudio 3

Multilingual transcription, meetings, subtitles, live content, customer-service recordings, specialized terminology, noisy audio, dialects, code-switching, and singing or music transcription

Text Speech Recognition Audio input Streaming
View model →
◎
OpenAI

text-moderation-stable

text-moderation

Legacy text-only safety classification and historical moderation integrations

Text Other Structured output
View model →
◎
Aleph Alpha

TFree-HAT-Pretrained-7B-Base

TFree-HAT

Research on tokenizer-free language modeling, English and German text generation, long-context experiments, and self-hosted non-commercial applications

Text General Purpose 32,900 ctx
View model →
◎
Cohere

Tiny Aya Earth

Tiny Aya

Multilingual translation, conversation, summarization, and text generation focused on African and West Asian languages; local and edge deployment

Text Lightweight 8,192 ctx
View model →
◎
Cohere

Tiny Aya Fire

Tiny Aya

South Asian multilingual conversation, translation, target-language generation, local inference, and edge or on-device applications

Text Lightweight 8,192 ctx
View model →
◎
Cohere

Tiny Aya Global

Tiny Aya

Multilingual translation, cross-lingual text generation, localized assistants, education, research, and efficient local or edge deployment

Text Lightweight 8,000 ctx
View model →
◎
Cohere

Tiny Aya Water

Tiny Aya

Efficient multilingual translation, text generation, localization, language learning, and local or edge deployment for European and Asia-Pacific languages

Text Lightweight 8,000 ctx
View model →
◎
ByteDance Seed

UI-TARS-1.5-7B

UI-TARS-1.5

Open-weight computer-use research, GUI grounding, browser automation prototypes, screenshot-based interface interaction, and visual action-model experimentation

Text Multimodal 128,000 ctx Image input
View model →
◎
Allen Institute for AI

Unified-IO

Unified-IO

Multimodal research, vision-language experiments, image generation, visual question answering, dense computer-vision tasks, and academic benchmarking

Text Multimodal Image input
View model →
◎
Allen Institute for AI

Unified-IO 2

Unified-IO

Multimodal research, image understanding and generation, audio and video understanding, spatial prediction, embodied AI and robotic-manipulation experiments, and self-hosted academic prototyping.

Text Multimodal Image input Audio input Video input
View model →
◎
Mistral AI

Voxtral Mini Transcribe 2

Voxtral

Batch transcription of meetings, interviews, calls, subtitles, compliance recordings, and searchable audio

Text Other Audio input
View model →
◎
Mistral AI

Voxtral Mini Transcribe Realtime

Voxtral

Live speech transcription, realtime captions, voice interfaces, realtime note-taking and speech-to-speech pipelines

Text Other Audio input Streaming
View model →
◎
Mistral AI

Voxtral Small

Voxtral

Production-scale audio understanding, multilingual transcription, audio Q&A, meeting and call summarization, speech translation, and voice-driven function calling

Text Multimodal 32,000 ctx Audio input Tool use Structured output
View model →
◎
Tencent

WAND-ASR-v1

WAND ASR

Audio and video transcription, subtitle generation, sentence-level timestamps, and short-form speech-to-text workflows

Text Other Audio input Video input
View model →
◎
OpenAI

Whisper

Whisper

Multilingual audio transcription, English speech translation, language identification, subtitles, captions, and word- or segment-level timestamps.

Text Other Audio input Streaming
View model →
◎
Yandex

YandexGPT 5 Lite

YandexGPT 5

Fast, cost-efficient Russian-language text generation and everyday conversational tasks

Text Lightweight
View model →
◎
Yandex

YandexGPT Pro 5.1

YandexGPT

Russian-language enterprise text generation, document analysis, RAG, structured extraction, rewriting, classification, reporting, and tool-augmented business assistants

Text General Purpose 32,768 ctx Tool use Web search Structured output
View model →
◎
01.AI

Yi Large FC

Yi Large

Function calling, tool selection, agent orchestration, and structured workflow automation

Text Coding 32,768 ctx Tool use Structured output Streaming
View model →
◎
01.AI

Yi Large Turbo

Yi Large

Low-cost short-context chat, text generation, summarization, classification, and general language applications

Text General Purpose 4,096 ctx Streaming
View model →
◎
01.AI

Yi-1.5-34B

Yi-1.5

Local bilingual text generation, research, domain adaptation, and fine-tuning

Text General Purpose 4,096 ctx
View model →
◎
01.AI

Yi-1.5-34B-Chat

Yi-1.5

Self-hosted bilingual assistants, English-Chinese text generation, coding support, mathematics, reasoning, fine-tuning, and privacy-sensitive deployments

Text General Purpose 4,096 ctx Streaming
View model →
◎
01.AI

Yi-1.5-6B

Yi-1.5

Local text generation, self-hosted applications, experimentation, domain adaptation, and fine-tuning

Text Lightweight 4,096 ctx Streaming
View model →
◎
01.AI

Yi-1.5-6B-Chat

Yi-1.5

Local conversational applications, instruction following, lightweight coding, bilingual English-Chinese text generation, experimentation, and self-hosted inference

Text General Purpose 4,096 ctx Streaming
View model →
◎
01.AI

Yi-1.5-9B

Yi-1.5

Local text generation, research, fine-tuning, Chinese-English applications, and cost-sensitive self-hosted deployments

Text General Purpose 4,096 ctx
View model →
◎
01.AI

Yi-1.5-9B-Chat

Yi-1.5

Self-hosted conversational assistants, local text generation, lightweight coding help, multilingual experimentation, and fine-tuning research.

Text General Purpose 4,096 ctx Streaming
View model →
◎
01.AI

Yi-34B

Yi

Self-hosted bilingual text generation, research, custom fine-tuning, coding experiments, and English-Chinese applications

Text General Purpose 4,096 ctx Streaming
View model →
◎
01.AI

Yi-34B-200K

Yi

Long-document analysis, English-Chinese generation, retrieval experiments, research, and self-hosted fine-tuning

Text General Purpose 200,000 ctx
View model →
◎
01.AI

Yi-34B-Chat

Yi

Self-hosted bilingual assistants, English-Chinese dialogue, open-weight LLM research, private inference, quantization, and fine-tuning

Text General Purpose 4,096 ctx Streaming
View model →
◎
01.AI

Yi-6B

Yi

Local bilingual text generation, research, fine-tuning, offline applications, and resource-conscious deployment

Text General Purpose 4,096 ctx
View model →
◎
01.AI

Yi-6B-200K

Yi

Long-document completion, local deployment, English-Chinese text generation, research, and downstream fine-tuning

Text General Purpose 200,000 ctx Streaming
View model →
◎
01.AI

Yi-6B-Chat

Yi

Local bilingual English-Chinese chat, personal projects, academic experimentation, and fine-tuning on modest open-weight infrastructure

Text General Purpose 4,096 ctx
View model →
◎
01.AI

Yi-9B

Yi

Local text generation, code completion, mathematics, bilingual English-Chinese applications, research, and downstream fine-tuning

Text General Purpose 4,096 ctx
View model →
◎
01.AI

Yi-9B-200K

Yi

Long-context document processing, code generation, mathematics, bilingual English-Chinese text generation, local inference, and domain fine-tuning

Text General Purpose 200,000 ctx Streaming
View model →
◎
01.AI

Yi-Coder-1.5B

Yi-Coder

Local code completion, code generation, multilingual programming tasks, long-context source-code analysis, and downstream fine-tuning

Text Coding 131,072 ctx Streaming
View model →
◎
01.AI

Yi-Coder-1.5B-Chat

Yi-Coder

Local code generation, completion, debugging, code explanation, lightweight IDE assistants, and model fine-tuning experiments

Text Coding 131,072 ctx Streaming
View model →
◎
01.AI

Yi-Coder-9B

Yi-Coder

Local code generation, completion, editing, repository-scale context, multilingual programming, and fine-tuning

Text Coding 131,072 ctx
View model →
◎
01.AI

Yi-Coder-9B-Chat

Yi-Coder

Self-hosted coding assistance, code generation, debugging, code explanation, code translation, and long-context repository analysis

Text Coding 131,072 ctx Streaming
View model →
◎
01.AI

Yi-Large

Yi

Long-context chat, complex text analysis, multilingual generation, prediction, and general-purpose enterprise language applications.

Text General Purpose 32,000 ctx Streaming
View model →
◎
01.AI

Yi-Lightning

Yi

Cost-sensitive hosted chat, Chinese-English generation, coding, mathematics, reasoning, summarization and high-volume API workloads

Text General Purpose 16,384 ctx Streaming
View model →
◎
01.AI

Yi-Vision

Yi-VL

Local bilingual image understanding, visual question answering, OCR-oriented image analysis, and image-to-text applications

Text Multimodal 4,096 ctx Image input
View model →
◎
01.AI

Yi-VL-34B

Yi-VL

Self-hosted bilingual image understanding, visual question answering, image text recognition, and research applications with substantial GPU capacity

Text Multimodal 4,096 ctx Image input
View model →
◎
01.AI

Yi-VL-6B

Yi-VL

Local bilingual image understanding, visual question answering, OCR-assisted extraction, image summarization, and lightweight multimodal experimentation

Text Multimodal 4,096 ctx Image input
View model →
◎
Tencent

YT-VITA

VITA

Video, image, audio, and text understanding; multimedia summarization; content tagging; structural video analysis; object localization; enterprise media workflows

Text Multimodal 128,000 ctx Image input Audio input Video input
View model →
◎
DeepSeek

DeepSeek-Prover-V1.5-RL

DeepSeek-Prover-V1.5

Lean 4 formal theorem proving, mathematical proof generation, proof-code completion, and research on verifier-guided language models

Text Reasoning 4,096 ctx
View model →
◎
DeepSeek

DeepSeek-Prover-V2-7B

DeepSeek-Prover-V2

Lean 4 theorem proving, formal mathematics, proof synthesis, lemma completion, and local proof-search research

Text Reasoning 32,768 ctx
View model →
◎
DeepSeek

DeepSeek-R1-0528-Qwen3-8B

DeepSeek-R1

Local mathematical reasoning, coding assistance, research, experimentation, and smaller-scale deployments requiring strong reasoning quality

Text Reasoning 131,072 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-R1-Distill-Llama-8B

DeepSeek-R1

Local mathematical reasoning, coding assistance, technical question answering, and private self-hosted inference

Text Reasoning 131,072 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-R1-Distill-Qwen-14B

DeepSeek-R1

Local mathematical reasoning, coding, technical analysis, research, and self-hosted text generation

Text Reasoning 131,072 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-R1-Distill-Qwen-32B

DeepSeek-R1

Mathematics, programming, complex reasoning, research, and self-hosted inference

Text Reasoning 32,768 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-R1-Distill-Qwen-7B

DeepSeek-R1

Local mathematical reasoning, coding assistance, research, and privacy-sensitive self-hosted applications

Text Reasoning 131,072 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-V3.1-Base

DeepSeek-V3.1

Self-hosted research, continued pretraining, fine-tuning, custom inference, and large-scale language or coding workloads

Text General Purpose 131,072 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-V3.2-Exp

DeepSeek-V3.2

Long-context text generation, reasoning, coding, research, document analysis, and self-hosted experimentation with sparse attention.

Text General Purpose 163,840 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V3.2-Exp-Base

DeepSeek V3.2

Long-context architecture research, self-hosted inference, continued pretraining, and custom model adaptation

Text General Purpose 163,840 ctx Streaming
View model →
◎
DeepSeek

DeepSeek-VL-1.3B-Chat

DeepSeek-VL

Local image-and-text chat, visual question answering, OCR, document and screenshot understanding, and multimodal research on modest hardware

Text Multimodal 4,096 ctx Image input Streaming
View model →
◎
DeepSeek

DeepSeek-VL-7B-Base

DeepSeek-VL

Local visual question answering, image and document understanding, multimodal research, and fine-tuning experiments

Text Multimodal 16,384 ctx Image input
View model →
◎
DeepSeek

DeepSeek-VL-7B-Chat

DeepSeek-VL

Self-hosted image understanding, visual question answering, document and webpage analysis, diagram interpretation, and research prototyping.

Text Multimodal 4,096 ctx Image input Streaming
View model →
◎
DeepSeek

DeepSeek-VL2

DeepSeek-VL2

Self-hosted image understanding, OCR, document and chart analysis, visual question answering, and visual grounding research

Text Multimodal 4,096 ctx Image input
View model →
◎
DeepSeek

DeepSeek-VL2-Small

DeepSeek-VL2

Local or self-hosted visual question answering, OCR, document and chart understanding, image-grounded conversation, and visual grounding

Text Multimodal 4,096 ctx Image input
View model →
◎
DeepSeek

DeepSeek-VL2-Tiny

DeepSeek-VL2

Local visual question answering, OCR, document and chart analysis, image understanding, and visual grounding

Text Multimodal 4,096 ctx Image input
View model →
◎
DeepSeek

DeepSeekMath-7B-Base

DeepSeekMath

Mathematical reasoning research, local inference, continued pretraining, and task-specific fine-tuning

Text Reasoning 4,096 ctx Streaming
View model →
◎
DeepSeek

DeepSeekMath-7B-Instruct

DeepSeekMath

Mathematical problem solving, educational assistants, local research, benchmark evaluation, and open-weight reasoning experiments

Text Reasoning 4,096 ctx Streaming
View model →
◎
DeepSeek

DeepSeekMath-7B-RL

DeepSeekMath

Open-weight mathematical reasoning, competition-math experiments, local deployment, and research on reinforcement learning for language models

Text Reasoning 4,096 ctx
View model →
◎
DeepSeek

DeepSeekMoE 16B Base

DeepSeekMoE

Local text completion, open-weight LLM research, domain adaptation, and fine-tuning

Text General Purpose 4,096 ctx
View model →
◎
Tencent

Hy-MT2-Plus

Hy-MT2

Professional multilingual translation, localization, terminology-sensitive workflows, and context-aware business translation

Text Translation 8,192 ctx Streaming
View model →
◎
DeepSeek

Janus-Pro-1B

Janus-Pro

Local multimodal research, image understanding, visual question answering, and compact text-to-image experimentation

Text Multimodal 4,096 ctx Image input
View model →
◎
DeepSeek

Janus-Pro-7B

Janus-Pro

Local image understanding, text-to-image generation, and unified multimodal research

Text Multimodal 4,096 ctx Image input
View model →
◎
Qwen

Qwen-Image-2.1-PE-I2I

Qwen-Image-2.1

Rewriting vague image-editing instructions, preserving source-image details, coordinating multi-image edits, and preparing prompts for Qwen-Image-2.1.

Text Multimodal 262,144 ctx Image input Structured output Streaming
View model →
◎
Qwen

Qwen-Image-2.1-PE-T2I

Qwen-Image-2.1

Expanding short or multilingual image requests into detailed prompts for Qwen-Image-2.1

Text Other 262,144 ctx Structured output
View model →
◎
Alibaba Cloud

Qwen-Max

Qwen-Max

Complex multilingual text generation, coding, logical reasoning, creative writing, structured extraction, and enterprise applications

Text General Purpose 32,768 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud Model Studio

qwen-plus-character

Qwen Character

Character role-play, virtual social applications, game NPCs, IP character replication, smart toys, in-car assistants, and empathetic conversational experiences

Text Other 32,768 ctx Web search Structured output
View model →
◎
Alibaba Cloud

Qwen-Turbo

Qwen-Turbo

High-volume customer support, simple to moderate question answering, summarization, rewriting, structured text extraction, and cost-sensitive applications

Text Lightweight 131,072 ctx Web search Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen-VL-Flash

Qwen-VL

Fast image and video understanding, visual question answering, document analysis, and multimodal extraction

Text Multimodal 32,768 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen-VL-Plus

Qwen-VL

High-resolution image and video understanding, OCR-style text recognition, document analysis, visual question answering, and multimodal assistants

Text Multimodal 131,072 ctx Image input Video input Structured output
View model →
◎
Alibaba Cloud

qwen3-vl-rerank

Qwen3-VL

Multimodal reranking, cross-modal search, image retrieval, video retrieval, image clustering, and multimodal RAG

Text Multimodal 120,000 ctx Image input Video input
View model →
◎
Alibaba Cloud

Qwen3.5-OCR

Qwen3.5-OCR

OCR, document parsing, text localization, table extraction, handwritten-text recognition, and key information extraction from images

Text Multimodal 65,536 ctx Image input
View model →
◎
Alibaba Cloud

Qwen3.7-Max

Qwen3.7

Complex reasoning, advanced coding, long-context analysis, tool-using agents, productivity automation, and long-horizon task execution

Text Reasoning 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

qwq-plus

QwQ

Mathematics, coding, research-style analysis, and complex multi-step reasoning

Text Reasoning 131,072 ctx Tool use Web search Streaming
View model →
◎
Meta

SeamlessStreaming

Seamless

Real-time multilingual speech recognition, simultaneous translation, speech-to-text translation, and speech-to-speech translation

Text Multimodal Audio input Streaming
View model →
◎
Alibaba Cloud Model Studio

text-embedding-v3

Qwen3-Embedding

Multilingual text embeddings, semantic search, RAG, recommendation, clustering, classification, and migration of existing v3 vector indexes

Text Embedding 8,192 ctx
View model →
Learn more

About text generation models

What text output means

Text output is the ability of an AI model to produce a sequence of textual symbols or tokens in response to an input, instruction, conversation, or internal generation process. In an API, the result may be returned as a plain string, a message containing text, streamed text fragments, or a text-encoded structure such as JSON.

The output does not have to be ordinary prose. Text-output models may return answers, explanations, summaries, translations, classifications, source code, SQL, HTML, Markdown, captions, extracted fields, or other content represented as text. The site's model taxonomy therefore treats text output as an output capability rather than as a guarantee that the model is text-only.

What a text-output model actually produces

A useful way to understand text generation is to think of the model as assembling a response piece by piece. Modern language models generally process the supplied context and predict successive tokens. A token may be a whole word, part of a word, punctuation, whitespace, or another small piece of text. The system continues until it reaches a stopping condition, an output limit, or a configured termination sequence.

Depending on the application, the resulting text may be:

  • Natural language: answers, explanations, dialogue, summaries, and drafts.
  • Code or markup: programming code, SQL, HTML, XML, Markdown, or configuration files.
  • Structured text: JSON or another schema-constrained representation.
  • Short labels: classifications, tags, routing decisions, or extracted values.
  • Streaming content: partial text delivered while the response is still being generated.
  • Status or refusal messages: explanations that a request cannot be completed or needs more information.

API responses may also contain metadata such as token counts, finish reasons, citations, safety annotations, or tool-call records. Those fields are part of the response envelope, but they are not necessarily text generated as the model's answer.

Input and output are separate capabilities

One of the most important distinctions in an AI model catalogue is the difference between what a model can accept and what it can produce. A model may accept an image and return a written description. It may accept an audio recording and return a transcript. It may process a PDF and return extracted fields or a summary. In each case, the native output can still be text.

Conversely, a model that accepts text does not automatically generate audio, images, or video. Text-to-speech, image generation, video generation, and music generation are separate output capabilities. A product may combine several models behind one interface, but the surrounding application should not be confused with the capability of the individual model.

When reviewing a model, ask two separate questions: What inputs can it understand? and What outputs can it create? This avoids treating a multimodal input model as though it necessarily has multimodal output.

How text output differs from related capabilities

Text understanding and text generation

Understanding and generation are related but not identical. A classifier may read text and return a label, while a generative model may write a paragraph, explanation, or structured record. Both can produce text, but they may differ greatly in controllability, latency, cost, and reliability.

Speech-to-text and text-to-speech

Speech-to-text converts audio into written language. Its defining input is audio and its output is text, so transcription is a text-output use case even though it is usually classified separately as an audio or speech capability. Text-to-speech works in the opposite direction: it accepts text and produces an audio waveform. A transcript returned alongside generated speech does not make the speech itself text output.

Image and video generation

Image and video models produce pixels or media files rather than ordinary text. They may also return textual metadata, prompts, captions, or status information, but those supporting fields do not turn the generated image or video into text output.

Embeddings

Embedding models produce numerical vectors that represent meaning for search, clustering, recommendation, or retrieval systems. The input may be text, an image, or another modality, but the native result is a vector rather than a textual answer.

Tool and function calling

A model can generate a structured request for another system, such as a weather service, database, search engine, or code interpreter. Function arguments are often encoded as JSON, but a tool call is not the same as an ordinary answer. It is an instruction or data package for external software. The tool's result comes from that external system, and the final response may be assembled by the application or generated by the model afterward.

Practical uses for text output

Text output is useful whenever information needs to be communicated, transformed, classified, or passed between software components. The right model depends on the required task, not simply on whether it can produce words.

  • Question answering: a user supplies a question and relevant context; the model returns an explanation or answer. If accuracy depends on current information, retrieval or citations may be needed.
  • Summarization: a document, meeting transcript, or webpage is supplied and the model returns a shorter version. Long inputs require an appropriate context window and careful checking for omitted details.
  • Extraction: an invoice, contract, or support message is provided and the model returns fields, labels, or JSON for a downstream system.
  • Translation and rewriting: the model receives source text and instructions about language, tone, length, or reading level, then returns a transformed version.
  • Coding assistance: a developer supplies a problem, specification, or existing code and receives code, an explanation, a test, or a proposed change. Human review and execution tests remain important.
  • Document and media understanding: an image, audio recording, video, or file is analyzed and represented as a caption, transcript, summary, or set of extracted facts.
  • Workflow automation: the model returns a classification, routing decision, or structured instruction that another application uses to continue a process.

For example, a customer-support workflow might provide a conversation and account context. The model could return a category, urgency label, suggested reply, and structured escalation data. A human agent or business system can then review or act on each part separately.

What matters when comparing text-output models

The best model is not necessarily the one with the longest context window or the most fluent prose. Compare models against the exact outputs your application needs.

  • Task quality: test the specific jobs involved, such as extraction, summarization, translation, coding, classification, or dialogue.
  • Factual accuracy and grounding: determine whether the model gives correct answers, uses supplied sources faithfully, and supports retrieval or citations where necessary.
  • Instruction following: check whether it follows constraints on tone, length, exclusions, language, and multi-step procedures.
  • Structured-output reliability: test required fields, data types, nested objects, enum values, schema adherence, and behavior when information is missing.
  • Context window: compare how much input the model can process and whether the usable capacity is sufficient for real documents, conversations, or retrieved material.
  • Output limits: check maximum output tokens, truncation behavior, stop sequences, and whether streaming is available.
  • Consistency: run the same prompts repeatedly to measure variation. Deterministic settings may improve repeatability, while sampling can produce more diverse drafts.
  • Latency and throughput: measure time to first token, total response time, tokens per second, concurrency, and delays introduced by retrieval or tool calls.
  • Language and domain coverage: evaluate the languages, terminology, reading levels, and document types that matter to the intended users.
  • Safety and refusal behavior: test legitimate sensitive use cases and determine whether refusals or filtering interfere with the application.
  • Cost: compare input and output token prices, caching, batch processing, reasoning charges, rate limits, and tool costs.

Benchmarks can help narrow the field, but a small evaluation set built from real examples is usually more useful. Include successful cases, ambiguous inputs, malformed documents, long contexts, repeated runs, and cases where the correct response is to say that information is unavailable.

Limitations and trade-offs

Fluent text is not proof of truth

Text-output models can produce confident but incorrect statements, invented references, faulty calculations, or plausible code that does not work. Retrieval, citations, deterministic checks, database validation, code execution, and human review can reduce these risks, but none is automatically present merely because a model generates text.

Formatting can fail

Unconstrained generation may add commentary, omit required fields, use invalid syntax, or return a value in the wrong format. Structured-output features and schema validation can improve reliability, but applications should still validate responses before using them in critical workflows.

Context and token limits affect results

Every model has limits on the amount of input and output it can process. Token limits do not correspond exactly to words or characters, and the relationship varies by language and tokenizer. Very long prompts can also increase latency and cost, even when they fit within the technical context window.

Quality varies by language and task

A model may perform strongly in one language, domain, or writing style and less reliably in another. Translation quality, terminology handling, instruction following, and safety behavior should be tested for the actual audience rather than inferred from general claims.

Privacy and integration matter

Sending prompts, documents, or media to a hosted model may involve provider retention, logging, access controls, or third-party processing. Review the relevant data-handling terms before using confidential information. Also account for authentication, rate limits, retries, response validation, monitoring, and version changes when integrating a model into production software.

Terminology is not completely standardized

Providers use overlapping terms such as text generation, text response, language generation, completion, and structured output. These labels can describe different scopes. “Completion” may refer specifically to continuing a prompt, while “structured output” often refers to constraints applied to a text response rather than to a separate output modality.

For a model catalogue, the most useful interpretation is that text output means the model can return text-encoded content. Task-specific capabilities, structured response support, input modalities, tool use, and other output types should be recorded separately when they are relevant.

Who needs a text-output model?

Almost any application that needs an AI-generated explanation, transformation, decision label, extracted record, or machine-readable response may need text output. It is particularly important when the result must be read by people, stored in a database, passed to another service, or used to guide a workflow.

Choose a text-output model when the desired result is fundamentally written or text-encoded. Choose a different or additional model when the application must create speech, images, video, embeddings, or physical actions. In multimodal systems, the practical solution may be a combination: one model interprets an image or audio file, a text-output model produces the explanation or structured record, and application code validates or routes the result.

Bottom line

Text output describes the form of an AI model's result: text-encoded content such as prose, code, labels, markup, or JSON. It does not guarantee factual accuracy, text-only input, speech, images, video, or autonomous action. To choose effectively, separate native model output from tool results and application features, then evaluate the model on the exact tasks, formats, languages, latency, reliability, privacy requirements, and costs that matter to the intended use.