{}
Models by output type

Structured Output Models: JSON Schema, Uses, Limits, and Reliability

Structured output models return information in a predictable format instead of unrestricted prose. A developer can define fields, data types, allowed values, required properties, and nested objects so that software can consume the response more reliably. This capability is useful for extracting data, classifying content, creating database records, generating application state, and preparing tool arguments. It is a formatting and integration capability—not a guarantee of factual accuracy, external action, or image, audio, or video generation.
What this means

Structured output lets software request responses that conform to a defined structure or schema rather than receiving unrestricted prose.

Structured Output models

221 models currently match this capability.

View all models →
◎
NVIDIA

Active Speaker Detection

Active Speaker Detection

Real-time or batch speaker identification and active-speaker tagging in broadcast, video communication, dubbing, localization, conferencing, and media analytics workflows.

Structured Output Multimodal Audio input Video input Structured output
View model →
◎
AI21

AI21-Jamba-Large-1.6

Jamba 1.6

Long-context retrieval-augmented generation, enterprise document analysis, grounded question answering, structured extraction, classification, and private deployment

Structured Output General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
AI21 Labs

AI21-Jamba-Mini-1.5

Jamba 1.5

Long-context document analysis, RAG, structured generation, function calling, multilingual enterprise assistants, and private deployment

Structured Output General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
AI21

AI21-Jamba-Mini-1.6

Jamba 1.6

Long-context RAG, grounded question answering, enterprise document processing, classification, structured text generation, function calling and privacy-sensitive private deployments.

Structured Output General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
AI21

AI21-Jamba-Mini-1.7

Jamba 1.7

Long-document analysis, enterprise RAG, grounded question answering, structured text generation, private deployment, and cost-sensitive text workflows

Structured Output General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
Amazon

Amazon Nova 2 Sonic

Amazon Nova 2

Real-time voice assistants, customer-service automation, telephony, interactive learning, multilingual conversations, and tool-enabled speech agents.

Structured Output Multimodal 1,000,000 ctx Audio input Tool use Structured output
View model →
◎
Amazon

Amazon Nova Premier

Amazon Nova

Long-context multimodal analysis, enterprise document workflows, complex tool calling, agentic orchestration, codebase analysis, and teacher-model distillation before retirement.

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
OpenAI

Chat Latest

Chat Latest

ChatGPT-style instant responses, general-purpose writing and analysis, image-aware conversations, and tool-assisted workflows

Structured Output General Purpose 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

ChatGPT-4o

GPT-4o

Fast general-purpose conversations, vision, voice interactions, coding, and everyday productivity

Structured Output Multimodal 128,000 ctx Image input Audio input Video input
View model →
◎
Anthropic

Claude Fable 5

Claude Fable 5

Complex reasoning, long-running autonomous agents, advanced coding, multi-stage research, document-heavy analysis, and high-value enterprise workflows

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Fable 5.1

Claude Fable 5

Long-running agentic coding, demanding reasoning, multistep research, complex document analysis, spreadsheets, presentations, and high-stakes knowledge work

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Haiku 4.5

Claude 4.5

Low-latency assistants, customer support, high-volume text processing, coding assistance, image understanding, and parallel subagent workloads

Structured Output Lightweight 200,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Mythos 5

Claude Mythos

Advanced cybersecurity, vulnerability research, biology, healthcare, life-sciences research, and long-running technical workflows requiring extensive context and reasoning

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Mythos 5.1

Claude Mythos

Vetted cybersecurity defense, advanced biology and life-sciences research, complex coding, long-running agents, technical investigation, and high-value research workflows

Structured Output Reasoning 1,000,000 ctx Image input Tool use Structured output
View model →
◎
Anthropic

Claude Opus 4.5

Claude Opus 4

Complex software engineering, coding agents, multi-step research, computer-use workflows, enterprise analysis, visual document understanding, and high-value tool-using applications.

Structured Output Reasoning 200,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 4.6

Claude Opus

Complex reasoning, agentic coding, repository-scale software work, long-context analysis, research, document and spreadsheet workflows, and enterprise automation

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 4.7

Claude Opus 4

Complex reasoning, agentic coding, repository-scale software engineering, research, document analysis, computer-use workflows, and high-stakes knowledge work

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 4.8

Claude Opus 4

Complex reasoning, agentic coding, long-horizon workflows, enterprise knowledge work, document analysis, vision tasks, and computer-use agents

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 5

Claude Opus

Complex agentic coding, code review, enterprise analysis, long-context research, document workflows, financial and legal reasoning, and multi-step tool-using applications.

Structured Output General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Opus 5.5

Claude Opus

Long-running agentic coding, repository-scale software engineering, code review, knowledge work, multimodal document analysis, and tool-using applications

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Sonnet 4.5

Claude Sonnet

Complex coding, software agents, computer-use workflows, visual document analysis, research, and long-running multi-step tasks

Structured Output Coding 200,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Sonnet 4.6

Claude Sonnet 4.6

Agentic coding, computer use, long-context document analysis, enterprise knowledge work, structured extraction, and high-volume applications needing strong reasoning at Sonnet-tier pricing

Structured Output General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Sonnet 5

Claude Sonnet

Agentic coding, software engineering, browser and computer-use workflows, long-context analysis, tool-driven automation, and high-volume assistants

Structured Output General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
Anthropic

Claude Sonnet 5.5

Claude Sonnet 5.5

Fast coding agents, long-context analysis, image-aware workflows, document creation, and tool-enabled business applications

Structured Output General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
Mistral AI

Codestral 25.08

Codestral

Low-latency IDE autocomplete, fill-in-the-middle completion, code generation, code editing, test generation and developer assistants

Structured Output Coding 128,000 ctx Tool use Structured output Streaming
View model →
◎
OpenAI

codex-mini-latest

Codex

Codex CLI coding workflows, code question answering, code editing, repository tasks, and low-latency software-engineering assistance

Structured Output Coding 200,000 ctx Image input Tool use Structured output
View model →
◎
Cohere

Command A

Command A

Enterprise RAG, long-context document analysis, multilingual applications, tool use, agentic workflows, financial text processing and structured text generation

Structured Output General Purpose 256,000 ctx Tool use Structured output Streaming
View model →
◎
Cohere

Command A Reasoning

Command A

Complex enterprise agents, tool use, retrieval-augmented generation, multilingual reasoning, long-context analysis, and workflow automation.

Structured Output Reasoning 256,000 ctx Tool use Web search Structured output
View model →
◎
Cohere

Command A Translate

Command A

High-quality multilingual text translation, enterprise document translation, and privacy-sensitive translation workflows.

Structured Output Specialized 8,000 ctx Tool use Structured output Streaming
View model →
◎
Cohere

Command A Vision

Command A

Enterprise document intelligence, OCR, chart and table analysis, visual question answering, and multilingual image understanding

Structured Output Multimodal 128,000 ctx Image input Structured output Streaming
View model →
◎
Cohere

Command A+

Command A

Enterprise agents, multimodal document and image analysis, multilingual workflows, reasoning-intensive automation, retrieval-augmented generation, and tool-using applications.

Structured Output Multimodal 128,000 ctx Image input Tool use Structured output
View model →
◎
Cohere

Command R 08-2024

Command R

Low-cost enterprise RAG, multilingual document workflows, long-context chat, structured extraction, and tool-using agents

Structured Output General Purpose 128,000 ctx Tool use Structured output Streaming
View model →
◎
Cohere

Command R+ 08-2024

Command R+

Complex enterprise RAG, long-context document analysis, multilingual assistants, citations, structured data tasks and multi-step tool-use agents.

Structured Output General Purpose 128,000 ctx Tool use Structured output Streaming
View model →
◎
Cohere

Command R7B

Command R

Cost-sensitive RAG, enterprise chat, tool use, coding assistance, and fast multi-step agents

Structured Output Lightweight 128,000 ctx Tool use Web search Structured output
View model →
◎
DeepSeek

DeepSeek-V3.1

DeepSeek-V3

Open-weight reasoning, coding, tool-calling, long-context analysis, and self-hosted agent systems

Structured Output General Purpose 128,000 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V3.1-Terminus

DeepSeek-V3.1

Open-weight deployment, coding assistance, long-context text processing, reasoning workflows, search agents, and terminal-oriented automation

Structured Output General Purpose 128,000 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V3.2

DeepSeek V3

Open-weight reasoning, coding, long-context analysis, tool-using agents, research workflows, and cost-sensitive deployments

Structured Output Reasoning 131,072 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V4-Pro

DeepSeek V4

Complex reasoning, coding agents, long-context analysis, tool-using workflows, and large document or codebase processing

Structured Output General Purpose 1,000,000 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V4.1-Flash

DeepSeek V4.1

Low-cost, high-throughput reasoning and coding, long-context analysis, agentic workflows, tool calling, and text-plus-image understanding.

Structured Output Lightweight 1,048,576 ctx Image input Tool use Structured output
View model →
◎
Baidu

ERNIE 5.0

ERNIE

Multimodal understanding, long-context Chinese and English applications, complex reasoning, coding, tool-enabled agents, and enterprise workloads

Structured Output Multimodal 248,832 ctx Image input Audio input Video input
View model →
◎
Baidu

ERNIE X1 Turbo

ERNIE X1

Chinese-language reasoning, long-form analysis, complex calculations, literary and document writing, agent workflows, and function calling

Structured Output Reasoning 32,768 ctx Tool use Web search Structured output
View model →
◎
Technology Innovation Institute

Falcon Perception

Falcon Perception

Natural-language object grounding, open-vocabulary detection, promptable instance segmentation, crowded-scene perception, robotics and visual inspection pipelines

Structured Output Multimodal 8,192 ctx Image input Structured output
View model →
◎
Google DeepMind

Gemini 2.5 Computer Use

Gemini 2.5

Browser automation, visual UI interaction, repetitive web workflows, form filling, and user-interface testing

Structured Output Multimodal 128,000 ctx Image input Tool use Structured output
View model →
◎
Google DeepMind

Gemini 2.5 Flash

Gemini 2.5

Large-scale multimodal processing, low-latency reasoning, coding, data extraction, tool-using agents, and applications requiring a very large context window.

Structured Output General Purpose 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 2.5 Flash-Lite

Gemini 2.5

High-volume classification, simple extraction, lightweight multimodal analysis, routing, tagging, summarization, and extremely latency-sensitive applications.

Structured Output Lightweight 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 2.5 Pro

Gemini 2.5

Advanced coding, complex reasoning, mathematics, STEM analysis, long documents, large codebases, multimodal analysis, and tool-using agents.

Structured Output Reasoning 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3 Flash

Gemini 3

Agentic workflows, everyday coding, reasoning and planning, multimodal analysis, long-context document work, and cost-sensitive tool-using applications

Structured Output Multimodal 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.1 Flash-Lite

Gemini 3.1

High-volume translation, classification, extraction, summarization, document processing, and lightweight tool-using agent workflows

Structured Output Lightweight 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.1 Pro

Gemini 3

Complex reasoning, advanced software engineering, long-context multimodal analysis, and agentic workflows requiring reliable tool use

Structured Output Reasoning 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.5 Flash

Gemini 3.5

Agentic workflows, coding agents, long-context multimodal analysis, tool use, and scaled production applications

Structured Output General Purpose 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google

Gemini 3.5 Flash-Lite

Gemini 3.5

High-volume, latency-sensitive agentic workflows, document parsing, translation, classification, data extraction, search-backed applications, and multimodal sub-agents.

Structured Output Lightweight 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.6 Flash

Gemini 3

Fast multimodal applications, coding assistants, long-context document and video analysis, tool-using agents, enterprise workflows, and rapid agentic execution

Structured Output General Purpose 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.7 Flash

Gemini 3

Coding, software engineering, tool-using agents, long-context multimodal analysis, structured extraction and high-throughput enterprise workflows

Structured Output General Purpose 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini 3.8 Flash

Gemini 3

Long-horizon software engineering, autonomous agents, multimodal document workflows, enterprise knowledge work, and tool-using applications.

Structured Output Multimodal 1,048,576 ctx Image input Audio input Video input
View model →
◎
Google DeepMind

Gemini Robotics ER 2

Gemini Robotics ER

High-level robot planning, spatial reasoning, video progress tracking, tool orchestration, and multi-robot collaboration

Structured Output Reasoning 131,072 ctx Image input Audio input Video input
View model →
◎
Z.AI

GLM-5.2

GLM-5

Long-horizon software engineering, repository-scale coding, code agents, complex refactoring, large-document analysis, research implementation, and tool-using workflows.

Structured Output Reasoning 1,000,000 ctx Tool use Structured output Streaming
View model →
◎
Z.ai

GLM-5.3

GLM-5

Complex software engineering, large codebases, long-horizon coding agents, terminal workflows, technical research, and authorized cybersecurity analysis

Structured Output Reasoning 1,000,000 ctx Tool use Structured output Streaming
View model →
◎
Z.ai

GLM-5.3-Flash

GLM-5

Multimodal coding agents, screenshot and interface understanding, long-context document work, visual verification, tool-assisted research, and cost-sensitive agent workflows

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Z.ai

GLM-OCR

GLM-OCR

High-volume OCR, PDF parsing, table and formula recognition, handwriting, receipts, forms, information extraction, and local document-processing pipelines

Structured Output Other Image input Structured output
View model →
◎
OpenAI

GPT-4.1

GPT-4.1

Software engineering, long-context document analysis, precise instruction following, structured extraction, tool-enabled agents, and image understanding

Structured Output General Purpose 1,047,576 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-4.1 Mini

GPT-4.1

Fast, cost-efficient instruction following, coding assistance, image understanding, structured extraction, tool calling, and long-context API applications

Structured Output Lightweight 1,047,576 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-4.1 nano

GPT-4.1

High-volume, latency-sensitive classification, extraction, routing, summarization, lightweight assistants, image-assisted analysis, and simple tool-calling workflows

Structured Output Lightweight 1,047,576 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-4.5 Preview

GPT-4.5

Historical general-purpose writing, creative work, nuanced communication, image understanding, programming assistance, and applications needing function calling or structured outputs.

Structured Output General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-4o

GPT-4o

General-purpose assistants, image understanding, coding help, structured extraction, multilingual generation, and latency-sensitive API workflows

Structured Output Multimodal 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-4o Mini

GPT-4o

Low-cost, high-volume text and image understanding, classification, extraction, translation, tagging, customer support, routing, and structured data generation

Structured Output Lightweight 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-4o Mini Search Preview

GPT-4o Mini Search Preview

Legacy Chat Completions applications requiring low-cost, search-grounded text responses

Structured Output Lightweight 128,000 ctx Tool use Web search Structured output
View model →
◎
OpenAI

GPT-4o Search Preview

GPT-4o

Historical web-search applications built around OpenAI Chat Completions

Structured Output Other 128,000 ctx Tool use Structured output Streaming
View model →
◎
OpenAI

GPT-5

GPT-5

Complex coding, reasoning, research, long-context analysis, visual understanding, tool-using agents, and structured professional workflows

Structured Output Reasoning 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5 Chat

GPT-5

ChatGPT-aligned conversational applications, text generation, image-aware question answering, structured outputs, and tool-enabled workflows requiring GPT-5 compatibility

Structured Output General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5 Mini

GPT-5

Cost-sensitive reasoning, coding assistance, structured extraction, document processing, high-volume automation, and tool-enabled workflows

Structured Output Lightweight 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5 nano

GPT-5

High-volume classification, summarization, extraction, ranking, routing, image-assisted analysis, and lightweight coding subagents

Structured Output Lightweight 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5 Pro

GPT-5

Difficult research, mathematics, science, complex coding, high-stakes analysis, and tool-using workflows where maximum answer quality matters more than latency or cost.

Structured Output Reasoning 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5-Codex

GPT-5

Agentic software engineering, repository-level coding, code review, debugging, refactoring, test generation, and frontend work using screenshots

Structured Output Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.1

GPT-5

Coding, long-context analysis, tool-using agents, structured outputs, and multi-step workflows

Structured Output General Purpose 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.1 Chat

GPT-5.1

Conversational assistants, instruction following, image-grounded chat, structured extraction, streaming responses, and tool-using API workflows.

Structured Output General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.1-Codex

GPT-5.1

Agentic software engineering, code generation, debugging, refactoring, testing, code review, and long-running Codex workflows

Structured Output Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.1-Codex Mini

GPT-5.1-Codex

Cost-sensitive agentic coding, code editing, repository maintenance, and Codex-style workflows

Structured Output Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.1-Codex-Max

GPT-5.1-Codex

Long-running agentic coding, repository-scale refactoring, multi-file implementation, debugging, code review, pull-request creation, and extended Codex workflows.

Structured Output Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.2

GPT-5.2

Complex professional work, long-context analysis, coding, document and spreadsheet workflows, visual understanding, and multi-step agents

Structured Output Reasoning 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.2 Chat

GPT-5.2

ChatGPT-aligned conversational applications, general writing, summarization, translation, vision-enabled assistants, and tool-calling workflows

Structured Output General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.2-Codex

GPT-5.2

Long-horizon agentic coding, large refactors, code migrations, repository-scale changes, terminal workflows, Windows development and defensive cybersecurity

Structured Output Coding 400,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.3 Chat

GPT-5.3

Fast general-purpose conversation, writing, summarization, text-and-image understanding, streaming responses, and function-calling applications

Structured Output General Purpose 128,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

GPT-5.3-Codex

GPT-5.3

Long-running agentic software engineering, codebase maintenance, debugging, testing, web development, tool-driven development, and technical computer workflows

Structured Output Coding 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.4

GPT-5.4

Complex professional work, advanced reasoning, software engineering, long-horizon agents, visual document analysis, computer use, research, and tool-heavy workflows

Structured Output Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.4 Mini

GPT-5.4

High-volume coding assistants, computer-use agents, subagents, tool calling, image reasoning, document workflows, and latency-sensitive applications

Structured Output Lightweight 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.4 nano

GPT-5.4

High-volume classification, data extraction, ranking, image understanding, routing, and lightweight coding subagents

Structured Output Lightweight 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.5

GPT-5.5

Complex coding, long-context research, professional analysis, tool-heavy agents, computer use, and multi-step workflow execution

Structured Output General Purpose 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.5 Pro

GPT-5.5

High-accuracy reasoning, complex coding, long-context research, data analysis, and multi-step professional workflows

Structured Output Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.6 Cyber

GPT-5.6

Authorized vulnerability research, exploit validation, exploit-chain development, advanced security testing, vulnerability triage, and defensive cybersecurity agents

Structured Output Coding 400,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.6 Luna

GPT-5.6

High-volume classification, summarization, routing, extraction, document understanding, agent automation, routine coding assistance, and cost-sensitive tool-using applications.

Structured Output Lightweight 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.6 Sol

GPT-5.6

Complex reasoning, coding, research, cybersecurity, science, long-context analysis, document-heavy workflows, and tool-using agents

Structured Output Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-5.6 Terra

GPT-5.6

Cost-conscious reasoning, coding agents, long-context analysis, structured business automation, research workflows, and tool-enabled production applications

Structured Output General Purpose 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-6 Astra

GPT-6

Complex reasoning, agentic coding, computer use, web research, scientific and professional workflows, and long-context document tasks

Structured Output Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-6 Luna

GPT-6

High-volume reasoning, document analysis, coding assistance, retrieval-augmented generation, and repeatable agent workflows

Structured Output Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

GPT-6 Sol

GPT-6

Complex coding, long-context reasoning, software engineering, research, computer use, and agentic workflows with tools.

Structured Output Reasoning 1,050,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

gpt-oss-120b

GPT-OSS

Self-hosted reasoning, coding, agentic workflows, private deployments, research, and fine-tuning

Structured Output Reasoning 131,072 ctx Tool use Web search Structured output
View model →
◎
OpenAI

gpt-oss-20b

gpt-oss

Local and private reasoning applications, coding assistants, agentic workflows, on-device or edge inference, fine-tuning, and cost-sensitive deployments with suitable hardware.

Structured Output Reasoning 131,072 ctx Tool use Web search Structured output
View model →
◎
OpenAI

gpt-oss-safeguard-120b

gpt-oss-safeguard

Custom-policy safety classification, LLM input and output filtering, trust and safety labeling, nuanced moderation review, and offline safety analysis

Structured Output Safety 131,072 ctx Structured output
View model →
◎
OpenAI

gpt-oss-safeguard-20b

GPT-OSS-Safeguard

Policy-based safety classification, LLM input/output filtering, content labeling, trust and safety review, and self-hosted moderation workflows

Structured Output Other 131,072 ctx Structured output
View model →
◎
IBM

Granite 4.0 3B Vision

Granite Vision

Chart extraction, table parsing, semantic key-value extraction, visual document processing, and local enterprise RAG pipelines

Structured Output Multimodal Image input Structured output
View model →
◎
IBM

Granite 4.1 30B

Granite 4.1

Self-hosted enterprise assistants, long-context RAG, multilingual applications, coding, structured extraction, and tool-calling agents

Structured Output General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite 4.1 3B

Granite 4.1

Efficient local or private deployment, multilingual enterprise text processing, RAG, summarization, extraction, coding assistance, function calling, and lightweight AI assistants.

Structured Output Lightweight 131,072 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite 4.1 8B

Granite 4.1

Self-hosted enterprise assistants, multilingual text generation, RAG, coding assistance, structured extraction, and tool-calling agents

Structured Output General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite 4.2 8B

Granite 4.2

Local or self-hosted reasoning, coding assistants, tool calling, multilingual dialogue, retrieval-augmented generation, and agentic workflows

Structured Output Reasoning 131,072 ctx Tool use Structured output Streaming
View model →
◎
IBM

Granite Vision 4.1 4B

Granite Vision 4.1

Structured extraction from charts, tables, invoices, forms, and enterprise document images

Structured Output Multimodal Image input Structured output Streaming
View model →
◎
IBM

Granite-4.0-H-Small

Granite 4.0

Enterprise RAG, multi-tool agents, function calling, customer-support automation, multilingual instruction following, and long-context workloads

Structured Output General Purpose 131,072 ctx Tool use Structured output
View model →
◎
IBM

Granite-4.0-H-Tiny

Granite 4.0

Efficient enterprise assistants, multilingual text generation, RAG, classification, extraction, summarization, coding assistance, fill-in-the-middle completion, structured JSON, and tool-calling workflows

Structured Output Lightweight 128,000 ctx Tool use Structured output Streaming
View model →
◎
xAI

Grok 4.20 Multi-Agent Beta

Grok 4.20

Deep research, parallel investigation, long-context analysis, and tool-assisted synthesis

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.20-0309-non-reasoning

Grok 4.20

Fast general-purpose text generation, image-aware analysis, coding assistance, structured extraction, tool-calling agents and large-context workflows

Structured Output General Purpose 1,000,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.20-0309-reasoning

Grok 4.20

Complex reasoning, coding, technical research, long-context document analysis, image understanding, structured responses, and tool-enabled agentic workflows

Structured Output Reasoning 1,000,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.3

Grok 4

Long-context analysis, enterprise agents, research, coding assistance, structured extraction and tool-enabled workflows

Structured Output Multimodal 1,000,000 ctx Image input Tool use Web search
View model →
◎
SpaceXAI

Grok 4.5

Grok 4

Software engineering, codebase analysis, technical reasoning, long-context work, tool-using agents, document analysis, and structured workflow automation

Structured Output Coding 500,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.6

Grok 4.6

Agentic coding, long-context software engineering, research, knowledge work, visual analysis, structured extraction, and tool-using workflows.

Structured Output Reasoning 500,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.7

Grok 4

Advanced software engineering, long-context reasoning, agentic tool use, research with web or X search, and professional knowledge work.

Structured Output Reasoning 500,000 ctx Image input Tool use Web search
View model →
◎
xAI

Grok 4.7 Fast

Grok 4.7

Low-latency coding, interactive development, and agentic workflows in Cursor or Grok Build

Structured Output General Purpose 500,000 ctx Image input Tool use Structured output
View model →
◎
SpaceXAI

Grok Build 0.1

Grok Build

Agentic coding, web development, debugging, software engineering workflows, MCP integrations, and fast tool-calling applications.

Structured Output Coding 256,000 ctx Image input Tool use Structured output
View model →
◎
NAVER

HCX-007

HyperCLOVA X

Complex reasoning, mathematics, science, language reasoning, writing, long-context text generation, and Korean-language enterprise applications

Structured Output Reasoning 128,000 ctx Tool use Structured output Streaming
View model →
◎
Tencent

HunyuanOCR-1.5

HunyuanOCR

Multilingual OCR, document parsing, text spotting, table and formula extraction, structured information extraction, and local visual-document processing

Structured Output Multimodal 131,072 ctx Image input Structured output Streaming
View model →
◎
Tencent Hunyuan

HY-Vision-1.5-Thinking

Hunyuan Vision 1.5

Image-grounded reasoning, OCR, chart and document analysis, visual localization, educational problem solving, and multilingual visual question answering

Structured Output Multimodal 40,000 ctx Image input Video input Tool use
View model →
◎
Tencent

Hy3

Hy

Coding agents, long-context analysis, complex reasoning, productivity automation, structured workflows, and multi-step tool use

Structured Output Reasoning 256,000 ctx Tool use Web search Structured output
View model →
◎
Tencent

Hy4 preview

Hy4

Long-context coding agents, complex tool-use workflows, productivity automation, document analysis, game development, and scientific reasoning

Structured Output General Purpose 1,000,000 ctx Tool use Structured output Streaming
View model →
◎
AI21

Jamba 1.5 Large

Jamba 1.5

Long-context document analysis, retrieval-augmented generation, enterprise assistants, structured text generation, multilingual workflows, and self-hosted deployments

Structured Output General Purpose 262,144 ctx Tool use Structured output Streaming
View model →
◎
Moonshot AI

Kimi K2.6

Kimi K2

Long-horizon software engineering, agentic coding, visual document understanding, tool-using workflows, and multi-agent orchestration

Structured Output Multimodal 262,144 ctx Image input Tool use Web search
View model →
◎
Moonshot AI

Kimi K3

Kimi K3

Long-context coding, software engineering, multimodal document and video understanding, agentic workflows, technical research, and complex reasoning

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Mistral AI

Leanstral 1.5

Leanstral

Lean 4 theorem proving, formal verification, autoformalization, proof debugging, and agentic proof engineering

Structured Output Other 256,000 ctx Tool use Structured output
View model →
◎
Meta

Llama 4 Scout

Llama 4

Long-context document and code analysis, visual question answering, multimodal assistants, multilingual applications, self-hosted inference, and customized deployments

Structured Output Multimodal 10,000,000 ctx Image input Tool use Structured output
View model →
◎
Microsoft

MedImageParse

BiomedParse

Text-guided biomedical image segmentation, annotation assistance, organ and tumor delineation, pathology-cell analysis, and research-oriented medical imaging pipelines

Structured Output Other Image input Structured output
View model →
◎
Xiaomi

MiMo-V2.5-Pro

MiMo V2.5

Long-horizon agent workflows, repository-scale coding, complex software engineering, tool-driven automation, and very large documents

Structured Output Reasoning 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Xiaomi

MiMo-V2.6-Flash

MiMo-V2.6

High-volume multimodal API workloads, coding assistants, agent automation, long-context document and repository analysis, tool-using workflows, and cost-sensitive professional applications.

Structured Output Multimodal 1,000,000 ctx Image input Audio input Video input
View model →
◎
Xiaomi MiMo

MiMo-V2.6-Pro

MiMo-V2.6

Long-horizon agents, coding, cybersecurity, research, computer use, multimodal analysis, and complex multi-step workflows

Structured Output Multimodal 1,000,000 ctx Image input Audio input Video input
View model →
◎
Xiaomi

MiMo-V2.6-Pro UltraSpeed

MiMo-V2.6

Latency-sensitive multimodal reasoning, coding, tool-use, long-context research, and interactive agent workflows

Structured Output Reasoning 1,000,000 ctx Image input Audio input Video input
View model →
◎
MiniMax

MiniMax M2.5-highspeed

MiniMax M2.5

Low-latency coding assistants, software-engineering agents, tool-using workflows, search tasks, and long-context productivity automation

Structured Output Coding 204,800 ctx Tool use Web search Structured output
View model →
◎
Mistral AI

Ministral 3 14B

Ministral 3

Private assistants, local vision-language applications, multilingual workloads, document and image analysis, and cost-efficient agentic systems

Structured Output Multimodal 262,144 ctx Image input Tool use Structured output
View model →
◎
Mistral AI

Ministral 3 3B

Ministral 3

Low-cost edge and local inference, image-aware assistants, document analysis, structured extraction, lightweight agents, task routing, and privacy-sensitive deployments.

Structured Output Lightweight 256,000 ctx Image input Tool use Structured output
View model →
◎
Mistral AI

Ministral 3 8B

Ministral 3

Efficient edge and local inference, image understanding, document workflows, structured extraction, lightweight agents, and high-volume text generation.

Structured Output Lightweight 256,000 ctx Image input Tool use Structured output
View model →
◎
Mistral AI

Mistral Large 3

Mistral Large

Long-context enterprise assistants, multilingual applications, image-aware document analysis, agentic workflows, coding, RAG and self-hosted sovereign deployments

Structured Output Multimodal 256,000 ctx Image input Tool use Web search
View model →
◎
Mistral AI

Mistral Medium 3.5

Mistral Medium

Agentic coding, software engineering, long-context analysis, multimodal document workflows, structured outputs and multi-step tool use

Structured Output Multimodal 256,000 ctx Image input Tool use Web search
View model →
◎
Mistral AI

Mistral Moderation 2

Mistral Moderation

Text moderation, conversational safety classification, content filtering, policy enforcement, guardrails, and jailbreaking detection

Structured Output Moderation 131,072 ctx Structured output
View model →
◎
Mistral AI

Mistral Small 4

Mistral Small

Cost-efficient general chat, multimodal document analysis, coding, agentic workflows, and configurable reasoning

Structured Output Multimodal 256,000 ctx Image input Tool use Web search
View model →
◎
Meta

Muse Spark 1.1

Muse Spark

Agentic workflows, coding agents, computer-use automation, multimodal document and media analysis, long-context reasoning, tool orchestration, and web-grounded applications.

Structured Output Reasoning 1,048,576 ctx Image input Audio input Video input
View model →
◎
Meta

Muse Spark 1.2

Muse Spark

Long-horizon coding agents, repository-scale software engineering, multimodal code generation, debugging, refactoring, and tool-driven workflows

Structured Output Coding 1,048,576 ctx Image input Audio input Video input
View model →
◎
Meta

Muse Spark 1.3

Muse Spark

Long-horizon coding agents, software engineering, browser and computer-use workflows, tool orchestration, large repositories, document analysis, and multimodal reasoning

Structured Output Multimodal 1,048,576 ctx Image input Audio input Video input
View model →
◎
NVIDIA

Nemotron OCR v1

Nemotron OCR

English OCR, document ingestion, layout-aware text extraction, multimodal retrieval, RAG preprocessing, and enterprise document intelligence

Structured Output Other Image input Structured output
View model →
◎
NVIDIA

Nemotron Table Structure v1

Nemotron

Detecting table cells, rows, columns, and merged-cell structure in document images for OCR alignment, table reconstruction, document ingestion, and retrieval systems

Structured Output Other Image input Structured output
View model →
◎
NVIDIA

nemotron-graphic-elements-v1

Nemotron Graphic Elements

Detecting and localizing chart titles, axis labels, legends, mark labels, and value labels in document images

Structured Output Other Image input Structured output
View model →
◎
NVIDIA

nemotron-page-elements-v3

Nemotron Page Elements

Document page-layout detection before OCR, table extraction, indexing, enterprise document processing, and multimodal RAG pipelines

Structured Output Other Image input Structured output
View model →
◎
Cohere

North Mini Code

North

Agentic software engineering, repository-level code changes, terminal-based coding agents, code review, local inference, and private deployment.

Structured Output Coding 256,000 ctx Image input Tool use Structured output
View model →
◎
Cohere

North Small Translate

North

Enterprise machine translation, multilingual documentation, localization, internal communications, safety procedures, and private or self-hosted translation workflows

Structured Output Other 16,000 ctx Tool use Structured output Streaming
View model →
◎
NVIDIA

NVIDIA Cosmos3-Nano-Reasoner

Cosmos 3

Physical-world video and image understanding, robotic perception, embodied-agent planning, spatial-temporal reasoning, and Physical AI research

Structured Output Reasoning 256,000 ctx Image input Video input Structured output
View model →
◎
NVIDIA

NVIDIA Nemotron OCR v2

Nemotron OCR

Multilingual OCR, scanned documents, forms, reports, charts, tables, image-based search, document ingestion, and retrieval-augmented generation preprocessing

Structured Output Other Image input Structured output
View model →
◎
OpenAI

o1

o1

Complex reasoning, mathematics, science, coding analysis, visual reasoning, and high-accuracy multi-step tasks

Structured Output Reasoning 200,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

o1 Preview

o1

Historically, difficult mathematics, science, coding, and other multi-step reasoning tasks requiring extended deliberation

Structured Output Reasoning 128,000 ctx Tool use Structured output Streaming
View model →
◎
OpenAI

o1-pro

o1

Complex reasoning, difficult technical analysis, advanced programming, research workflows, and tasks where answer consistency matters more than latency or cost.

Structured Output Reasoning 200,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

o3

o-series

Complex reasoning, advanced coding, mathematics, science, technical research, visual analysis and multi-step tool workflows

Structured Output Reasoning 200,000 ctx Image input Tool use Web search
View model →
◎
OpenAI

o3-mini

o3

Coding, mathematics, science, technical analysis, structured extraction, text-to-SQL, and multi-step reasoning

Structured Output Reasoning 200,000 ctx Tool use Structured output Streaming
View model →
◎
OpenAI

o3-pro

o3

High-reliability reasoning, advanced mathematics, scientific analysis, complex coding, research, and multi-step professional work.

Structured Output Reasoning 200,000 ctx Image input Tool use Structured output
View model →
◎
OpenAI

o4-mini

o4

Fast, cost-sensitive reasoning; coding; mathematics; visual analysis; structured extraction; high-volume tool-using agents

Structured Output Reasoning 200,000 ctx Image input Tool use Web search
View model →
◎
Mistral AI

OCR 3

OCR

High-volume document extraction, scanned forms, handwriting, invoices, complex tables, archival digitization, and document-to-knowledge pipelines.

Structured Output Other Image input Structured output
View model →
◎
Mistral AI

OCR 4.0

Mistral OCR

High-volume OCR, structured document extraction, enterprise search, RAG ingestion, invoice processing, compliance workflows, and document automation

Structured Output Other Image input Structured output
View model →
◎
Mistral AI

OCR 4.1

OCR

OCR, document parsing, structured extraction, enterprise search, RAG ingestion, invoice processing, and document AI workflows

Structured Output Other Image input Structured output
View model →
◎
Baidu

PP-StructureV3

PP-Structure

Document parsing, OCR, layout analysis, table extraction, and structured document understanding

Structured Output Multimodal Image input Structured output
View model →
◎
Alibaba Cloud Model Studio

Qwen-Flash

Qwen3

Fast, high-volume text generation; long-context analysis; summarization; extraction; structured outputs; and applications needing optional reasoning.

Structured Output Lightweight 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Qwen

qwen-flash-character

Qwen Character

Low-latency character dialogue, virtual companions, game NPCs, role-playing applications, IP character replication, and conversational smart devices

Structured Output Lightweight 32,768 ctx Web search Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen-Plus

Qwen-Plus

General-purpose text generation, long-context analysis, multilingual applications, structured business workflows, function-calling agents, and applications that need optional reasoning mode.

Structured Output General Purpose 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

Qwen-VL-Max

Qwen-VL

Complex image and video understanding, document analysis, chart interpretation, visual question answering, and structured extraction

Structured Output Multimodal 131,072 ctx Image input Video input Structured output
View model →
◎
Qwen

Qwen2.5-14B-Instruct

Qwen2.5

Self-hosted multilingual assistants, document processing, RAG, coding support, structured extraction, and cost-conscious production deployments.

Structured Output General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen2.5-32B-Instruct

Qwen2.5

Self-hosted assistants, multilingual text generation, long-document processing, coding support, structured extraction, RAG, and agent applications

Structured Output General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen2.5-72B-Instruct

Qwen2.5

Self-hosted multilingual assistants, coding, mathematics, document analysis, structured text generation, and long-context workloads

Structured Output General Purpose 131,072 ctx Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen2.5-7B-Instruct

Qwen2.5

Self-hosted chat assistants, multilingual text generation, coding and mathematics assistance, long-context document work, structured text generation, and cost-sensitive private deployments

Structured Output General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Qwen

Qwen2.5-VL-72B-Instruct

Qwen2.5-VL

High-quality image, document, chart, screenshot, OCR, visual-grounding, and video analysis; multimodal agents and self-hosted experimentation

Structured Output Multimodal 131,072 ctx Image input Video input Tool use
View model →
◎
Qwen

Qwen2.5-VL-7B-Instruct

Qwen2.5-VL

Local or self-hosted image and video understanding, OCR, document extraction, chart and diagram analysis, visual question answering, visual grounding, and multimodal research

Structured Output Multimodal 32,768 ctx Image input Video input Structured output
View model →
◎
Qwen

Qwen3-14B

Qwen3

Local deployment, multilingual assistants, reasoning, mathematics, coding, structured text generation, research, and cost-sensitive agent workflows.

Structured Output General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen3-235B-A22B

Qwen3

Complex reasoning, mathematics, software development, multilingual applications, function calling, agentic workflows, research and self-hosted open-weight deployment

Structured Output Reasoning 131,072 ctx Tool use Structured output Streaming
View model →
◎
Qwen

Qwen3-32B

Qwen3

Self-hosted reasoning assistants, coding agents, mathematics, multilingual applications, structured text generation, and tool-calling workflows

Structured Output Reasoning 256,000 ctx Tool use Structured output Streaming
View model →
◎
Qwen

Qwen3-8B

Qwen3

Cost-efficient reasoning, coding, multilingual assistants, local deployment, structured text generation, tool-enabled agents, and fine-tuned applications

Structured Output General Purpose 131,072 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud Model Studio

Qwen3-Max

Qwen3-Max

Complex reasoning, coding assistance, web-grounded agents, function calling, structured extraction, and long-context text analysis

Structured Output Reasoning 262,144 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

Qwen3-VL-235B-A22B-Instruct

Qwen3-VL

High-quality image and video understanding, OCR, document intelligence, visual coding, spatial reasoning, long-context multimodal analysis and visual-agent applications

Structured Output Multimodal 131,072 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3-VL-30B-A3B-Instruct

Qwen3-VL

Image and video understanding, OCR, document analysis, spatial reasoning, visual coding, long-context multimodal tasks, and visual-agent applications

Structured Output Multimodal 256,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3-VL-32B-Instruct

Qwen3-VL

Document intelligence, OCR, image and video understanding, spatial reasoning, visual coding, and visual-agent applications

Structured Output Multimodal 131,072 ctx Image input Video input Tool use
View model →
◎
Qwen

Qwen3-VL-4B-Instruct

Qwen3-VL

Local image and video understanding, OCR, document extraction, visual question answering, visual coding, and lightweight multimodal agents

Structured Output Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3-VL-8B-Instruct

Qwen3-VL

Local or hosted image and video understanding, OCR, document extraction, visual question answering, spatial reasoning, screenshot analysis, multimodal agents, and structured data extraction

Structured Output Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-122B-A10B

Qwen3.5

Advanced multimodal reasoning, image and video understanding, coding, long-context analysis, document and chart interpretation, function-calling agents, and web-grounded workflows.

Structured Output Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.5-27B

Qwen3.5

Long-context multimodal analysis, document understanding, video and image interpretation, general reasoning, coding, tool-enabled assistants, and self-hosted deployment

Structured Output Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.5-35B-A3B

Qwen3.5

Efficient multimodal assistants, coding, reasoning, long-context analysis, local deployment, and tool-using agents

Structured Output Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.5-397B-A17B

Qwen3.5

Advanced multimodal reasoning, coding, video and image understanding, long-context analysis, tool-using agents, and self-hosted open-weight deployments

Structured Output Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.5-Flash

Qwen3.5

Fast long-context text, image and video understanding; structured extraction; tool-enabled agents; web-grounded applications; and high-volume multimodal workloads.

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.5-Plus

Qwen3.5

Long-context reasoning, multimodal document and video analysis, coding, structured enterprise automation, function-calling agents, and web-grounded research.

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.6-27B

Qwen3.6

Coding agents, repository-level software engineering, visual document analysis, video understanding, STEM reasoning, long-context assistants, and self-hosted multimodal applications

Structured Output Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Qwen

Qwen3.6-35B-A3B

Qwen3.6

Agentic coding, long-context software engineering, multimodal analysis, tool-using agents, and self-hosted deployments

Structured Output Multimodal 262,144 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.6-Flash

Qwen3.6

Fast multimodal assistants, coding agents, visual document analysis, video understanding, tool-using workflows, object localization, and large-context applications

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.6-Max-Preview

Qwen3.6

Advanced coding agents, front-end development, long-context analysis, structured API workflows, and text generation with web search

Structured Output General Purpose 262,144 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

Qwen3.6-Plus

Qwen3.6

Long-context multimodal analysis, agentic coding, OCR, object localization, frontend development, visual reasoning, and tool-enabled enterprise assistants

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.7-Flash

Qwen3.7

Fast multimodal agents, visual coding, tool-use workflows, search agents, and long-context document or screen analysis

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.7-Plus

Qwen3.7

Long-context reasoning, multimodal document and video analysis, coding, tool-using agents, structured extraction, and enterprise productivity workflows

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.8-2.4T-A95B

Qwen3.8

Advanced reasoning, coding, scientific and professional research, long-context analysis, and long-horizon agent workflows

Structured Output Reasoning 1,000,000 ctx Tool use Web search Structured output
View model →
◎
Alibaba Cloud

Qwen3.8-27B

Qwen3.8

Coding assistants, repository analysis, visual document workflows, long-context research, office automation, multimodal agents, and tool-using applications

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.8-Flash

Qwen3.8

Fast long-context reasoning, coding assistance, visual document and chart analysis, video understanding, function-calling agents, and high-concurrency applications

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud

Qwen3.8-Max

Qwen3.8

Complex coding, autonomous software engineering, long-horizon agent workflows, professional document analysis, visual reasoning, long videos, and demanding research tasks.

Structured Output Multimodal 1,000,000 ctx Image input Video input Tool use
View model →
◎
Alibaba Cloud Model Studio

Qwen3.8-Omni-Flash

Qwen3.8-Omni

Long-form audio and video understanding, multimedia analysis, audio-visual agents, content summarization, and tool-using workflows

Structured Output Multimodal 1,000,000 ctx Image input Audio input Video input
View model →
◎
Reka

Reka Flash

Reka Flash

Fast multimodal applications, document and image analysis, short-video understanding, structured extraction, multilingual chat, coding, and tool-using agents

Structured Output Multimodal 128,000 ctx Image input Audio input Video input
View model →
◎
Cohere

rerank-english-v3.0

Rerank 3

English semantic reranking for enterprise search, hybrid retrieval, RAG pipelines, FAQs, knowledge bases, documents, code retrieval, and semi-structured records

Structured Output Other 4,096 ctx Structured output
View model →
◎
Meta

SAM 3.1

Segment Anything

Open-vocabulary object detection, pixel-level image segmentation, and multi-object video tracking

Structured Output Other Image input Video input Structured output
View model →
◎
ByteDance Seed

Seed1.6

Seed1.6

Multimodal document analysis, visual question answering, coding, mathematics, general reasoning, long-context analysis and adaptive-thinking applications

Structured Output Multimodal 256,000 ctx Image input Tool use Structured output
View model →
◎
ByteDance

Seed1.8

Seed

Multimodal agent workflows, search and information retrieval, coding agents, GUI interaction, image and video understanding, complex instruction following, and long-context business tasks.

Structured Output Multimodal 256,000 ctx Image input Video input Tool use
View model →
◎
ByteDance

Seed2.0 Lite

Seed2.0

Cost-conscious production applications requiring long-context multimodal understanding, document and video analysis, coding assistance, tool use, GUI automation, and structured extraction.

Structured Output Multimodal 262,144 ctx Image input Audio input Video input
View model →
◎
ByteDance Seed

Seed2.0 Pro

Seed2.0

Complex multimodal reasoning, long-chain agent workflows, visual and video analysis, document understanding, scientific research support, coding, and enterprise automation.

Structured Output Multimodal 200,000 ctx Image input Audio input Video input
View model →
◎
ByteDance Seed

Seed2.1 Turbo

Seed2.1

Fast multimodal agents, coding assistants, document and video analysis, tool-calling workflows, structured data extraction, and cost-sensitive production applications

Structured Output Multimodal 256,000 ctx Image input Video input Tool use
View model →
◎
SenseTime

SenseNova-Vision-7B-MoT

SenseNova-Vision

Unified computer-vision research, detection, OCR, segmentation, depth and normal estimation, visual grounding, and multi-view geometry

Structured Output Multimodal Image input Structured output
View model →
◎
NVIDIA

StreamPETR

StreamPETR

Camera-only multi-view 3D perception, autonomous-driving scene analysis, bird's-eye-view visualization, and object tracking

Structured Output Other Image input Video input Structured output
View model →
◎
OpenAI

text-moderation-stable

text-moderation

Legacy text-only safety classification and historical moderation integrations

Structured Output Other Structured output
View model →
◎
Mistral AI

Voxtral Small

Voxtral

Production-scale audio understanding, multilingual transcription, audio Q&A, meeting and call summarization, speech translation, and voice-driven function calling

Structured Output Multimodal 32,000 ctx Audio input Tool use Structured output
View model →
◎
Yandex

YandexGPT Pro 5.1

YandexGPT

Russian-language enterprise text generation, document analysis, RAG, structured extraction, rewriting, classification, reporting, and tool-augmented business assistants

Structured Output General Purpose 32,768 ctx Tool use Web search Structured output
View model →
◎
01.AI

Yi Large FC

Yi Large

Function calling, tool selection, agent orchestration, and structured workflow automation

Structured Output Coding 32,768 ctx Tool use Structured output Streaming
View model →
◎
DeepSeek

DeepSeek-V3.2-Exp

DeepSeek-V3.2

Long-context text generation, reasoning, coding, research, document analysis, and self-hosted experimentation with sparse attention.

Structured Output General Purpose 163,840 ctx Tool use Structured output Streaming
View model →
◎
Qwen

Qwen-Image-2.1-PE-I2I

Qwen-Image-2.1

Rewriting vague image-editing instructions, preserving source-image details, coordinating multi-image edits, and preparing prompts for Qwen-Image-2.1.

Structured Output Multimodal 262,144 ctx Image input Structured output Streaming
View model →
◎
Qwen

Qwen-Image-2.1-PE-T2I

Qwen-Image-2.1

Expanding short or multilingual image requests into detailed prompts for Qwen-Image-2.1

Structured Output Other 262,144 ctx Structured output
View model →
◎
Alibaba Cloud

Qwen-Max

Qwen-Max

Complex multilingual text generation, coding, logical reasoning, creative writing, structured extraction, and enterprise applications

Structured Output General Purpose 32,768 ctx Tool use Structured output Streaming
View model →
◎
Alibaba Cloud Model Studio

qwen-plus-character

Qwen Character

Character role-play, virtual social applications, game NPCs, IP character replication, smart toys, in-car assistants, and empathetic conversational experiences

Structured Output Other 32,768 ctx Web search Structured output
View model →
◎
Alibaba Cloud

Qwen-Turbo

Qwen-Turbo

High-volume customer support, simple to moderate question answering, summarization, rewriting, structured text extraction, and cost-sensitive applications

Structured Output Lightweight 131,072 ctx Web search Structured output Streaming
View model →
◎
Alibaba Cloud

Qwen-VL-Plus

Qwen-VL

High-resolution image and video understanding, OCR-style text recognition, document analysis, visual question answering, and multimodal assistants

Structured Output Multimodal 131,072 ctx Image input Video input Structured output
View model →
◎
Alibaba Cloud

Qwen3.7-Max

Qwen3.7

Complex reasoning, advanced coding, long-context analysis, tool-using agents, productivity automation, and long-horizon task execution

Structured Output Reasoning 1,000,000 ctx Tool use Web search Structured output
View model →
Learn more

About structured output models

What structured output means

Structured output is a model capability that constrains a response to follow a developer-defined data structure. The most common structure is a JSON object or array described with JSON Schema. A schema can specify fields such as names, dates, numbers, booleans, categories, arrays, nested objects, and required properties.

For example, an application might ask a model to review a support message and return:

{
  "category": "billing",
  "priority": "high",
  "customer_name": "Sam Lee",
  "needs_human_review": true
}

Without structured output, the model might express the same information in several different ways. With a supported schema, the response is intended to have a consistent shape that software can validate and process.

The terminology is not completely standardized. Providers may describe related features as structured outputs, schema-constrained generation, JSON Schema responses, typed responses, or strict tool schemas. The exact guarantees and supported schema features depend on the provider, model, endpoint, and SDK.

What the model actually produces

In most cases, structured output is still generated through a text-generation interface. The response usually contains machine-readable text representing a JSON object or array. An SDK may then parse that response into a typed object in Python, JavaScript, or another programming language.

This does not make structured output a separate physical output modality. A JSON object containing a caption, classification, or list of detected items is still structured text. The model has not generated an image, audio file, or video simply because the response contains fields describing one.

Input and output are different questions

A model can accept images, audio, documents, or video as input and return structured text as output. For example, a vision-capable model may inspect an invoice image and return a JSON object containing the invoice number, supplier, line items, tax, and total. The image is the input; the JSON is the output.

Likewise, a text-only model may produce structured JSON from a written request. To determine whether a model belongs in this category, evaluate what it directly generates rather than only what it can understand.

How structured output works

A request normally supplies both instructions and a schema. The provider may use constrained decoding, schema-guided generation, validation, or a combination of these techniques to reduce the chance that the response violates the requested structure.

At a conceptual level, the schema acts like a form that the model must fill in. It can restrict the available fields and values, but it cannot independently verify whether the values are true. A schema might require a valid-looking date or restrict a priority field to low, medium, or high; it cannot determine whether the chosen date or priority is appropriate without reliable source information and application logic.

SDKs add another layer. Some allow developers to define schemas using typed classes, Pydantic models, Zod objects, or similar declarations. The SDK may convert those definitions into a provider-compatible schema and parse the result. That convenience should not be confused with a provider guarantee: client-side parsing, provider-side enforcement, and post-processing are separate layers.

Structured output versus related capabilities

JSON mode

JSON mode generally asks the model to return syntactically valid JSON. It may not require a particular set of fields, data types, nesting structure, or allowed values. Structured output usually goes further by targeting adherence to a specified schema. The distinction matters when downstream software expects a particular contract.

Function calling and tool use

Function calling or tool use lets a model emit structured arguments for an external function. For example, the model might produce a customer ID and date range for a reporting function. The application then decides whether to execute that function.

A structured response, by contrast, is generally the final answer returned in the requested format. The two capabilities overlap because both can use schemas, and some providers offer strict schemas for tool arguments. However, a tool call does not mean that the model itself completed the action. The external application, database, search service, or API must execute it and return a result.

Application-level parsing

An application can ask for ordinary prose and then use a parser or another model to convert that prose into JSON. The application may end up with structured data, but this does not prove that the underlying model natively supports structured output. Native schema enforcement, SDK parsing, and application post-processing should be evaluated separately.

Multimodal generation

Structured output concerns the shape of a response. Multimodal generation concerns the direct creation of media such as images, audio, or video. A model can analyze an image and return structured JSON without generating an image, and an image-generation model can create an image without supporting schema-constrained text responses.

Practical uses

Structured output is most valuable when another system needs predictable data rather than an answer intended only for reading.

Extracting information from documents

A user can provide an invoice, contract, résumé, receipt, or form. The model can return fields such as dates, names, totals, clauses, or line items in a predefined structure. The application can then validate the result, store it in a database, or route it for human review.

This is particularly useful when documents vary in layout but the application needs a consistent record. The model handles interpretation; the schema provides a stable destination for the extracted values.

Classification and routing

A customer-support system can ask the model to classify a message into an approved category, assign a priority, identify sentiment, and indicate whether escalation is needed. The structured result can route the message to a queue or trigger a workflow without requiring fragile text matching.

Meeting and project data

Given meeting notes or a transcript, a model can return action items with fields for the task, owner, deadline, status, and priority. A project-management system can then display those items or request confirmation before creating them.

Application interfaces

Structured responses can populate dynamic forms, user-interface components, search filters, content-management fields, or application state. A model might return a list of suggested filters, a set of form fields, or a structured report that the interface renders consistently.

Tool and workflow integration

Agents often need structured arguments for search, database queries, calendars, ticketing systems, or business APIs. Schema-constrained arguments reduce integration errors, but the application should still check permissions, validate values, and decide whether execution is safe.

What matters when comparing structured-output models

The important question is not simply whether a model can produce JSON. Compare the strength and usefulness of its structured-output contract.

  • Guarantee level: Determine whether the provider promises valid JSON, adherence to a JSON Schema, strict tool arguments, or only prompt-based formatting.
  • Schema support: Check support for nested objects, arrays, enums, nullable fields, unions, recursion, additional properties, numeric constraints, and string formats. Providers commonly support only a subset of JSON Schema.
  • Failure behavior: Find out how the API represents refusals, unsupported schemas, invalid requests, truncation, timeouts, and incomplete responses.
  • Semantic accuracy: Measure whether extracted fields and classifications are correct. A response can match the schema while containing invented or misinterpreted values.
  • Consistency: Test repeated requests, ambiguous inputs, missing information, and model or prompt changes. Structural consistency does not necessarily mean identical field values.
  • Latency and streaming: Consider normal response time, first-use schema processing, schema caching, and whether partially streamed responses are practical for the application.
  • Limits: Check maximum schema size, nesting depth, response length, array size, context window, and behavior with large documents.
  • Tool integration: If tools are involved, verify whether strict schemas apply to function calls, whether parallel calls are supported, and how tool results are represented.
  • SDK support: Typed parsing can simplify development, but confirm that the SDK version, language, and endpoint support the features you need.
  • Operational cost: Account for schema overhead, token usage, retries, validation, human review, and additional tool or service calls.

Limitations and trade-offs

Correct format does not mean correct information

The most important limitation is that structure is not truth. A model can return a perfectly valid object with the wrong customer, date, amount, diagnosis, category, or conclusion. Applications should validate important values against source documents, business rules, databases, or human review.

Schema restrictions can reduce flexibility

Provider implementations may reject or ignore unsupported schema keywords. Some require every field to be present, while others encourage nullable fields or explicit alternatives. Deeply nested or highly complex schemas can increase latency and make errors harder to diagnose.

Refusals and incomplete responses still occur

A model may refuse an unsafe request instead of returning the requested object. Token limits, interruptions, and timeouts can also leave a response incomplete. Production code should handle refusal representations, missing fields, null values, parsing failures, and tool errors explicitly rather than assuming every request succeeds.

Strict structure is not determinism

Schema enforcement can make the shape of a response predictable while the contents still vary. Ambiguous source material may lead to different classifications or omissions across attempts. Testing should measure both structural validity and substantive correctness.

Privacy and security remain application responsibilities

Structured output does not make sensitive data safe by itself. Applications should consider what documents are sent to a provider, how long data is retained, who can access the resulting records, and whether generated tool arguments could cause unauthorized actions. Strict schemas can reduce accidental integration errors, but they are not a replacement for authentication, authorization, or security testing.

How to evaluate one in practice

  1. Define a schema that reflects the real application rather than a simplified demonstration.
  2. Confirm which schema features the selected provider and endpoint officially support.
  3. Test ordinary examples along with missing, conflicting, ambiguous, malformed, and adversarial inputs.
  4. Measure schema validity, required-field compliance, enum accuracy, null handling, refusal handling, and truncation.
  5. Score semantic correctness separately from formatting correctness.
  6. Check whether the model invents values when the source does not provide enough information.
  7. Measure latency, token usage, throughput, retries, and behavior near response-size limits.
  8. Validate every response in application code and define safe behavior for uncertainty, failure, and external actions.
  9. Repeat the evaluation after changing the model, SDK, schema, prompt, or provider configuration.

Who needs structured output?

Structured output is a strong fit when AI-generated information must enter software reliably. It is useful for document extraction, classification, workflow routing, database population, user-interface generation, and agent integrations.

It may be unnecessary when a person only wants a natural-language explanation, brainstorming, or creative writing. In those cases, a rigid schema can add complexity without much benefit. The capability becomes valuable when predictable fields, validation, and downstream automation matter more than conversational flexibility.

The practical rule is simple: use structured output when the application needs a dependable shape for an AI response, then add separate checks for whether the contents are accurate, safe, complete, and authorized for use.