What structured output means
Structured output is a model capability that constrains a response to follow a developer-defined data structure. The most common structure is a JSON object or array described with JSON Schema. A schema can specify fields such as names, dates, numbers, booleans, categories, arrays, nested objects, and required properties.
For example, an application might ask a model to review a support message and return:
{
"category": "billing",
"priority": "high",
"customer_name": "Sam Lee",
"needs_human_review": true
}Without structured output, the model might express the same information in several different ways. With a supported schema, the response is intended to have a consistent shape that software can validate and process.
The terminology is not completely standardized. Providers may describe related features as structured outputs, schema-constrained generation, JSON Schema responses, typed responses, or strict tool schemas. The exact guarantees and supported schema features depend on the provider, model, endpoint, and SDK.
What the model actually produces
In most cases, structured output is still generated through a text-generation interface. The response usually contains machine-readable text representing a JSON object or array. An SDK may then parse that response into a typed object in Python, JavaScript, or another programming language.
This does not make structured output a separate physical output modality. A JSON object containing a caption, classification, or list of detected items is still structured text. The model has not generated an image, audio file, or video simply because the response contains fields describing one.
Input and output are different questions
A model can accept images, audio, documents, or video as input and return structured text as output. For example, a vision-capable model may inspect an invoice image and return a JSON object containing the invoice number, supplier, line items, tax, and total. The image is the input; the JSON is the output.
Likewise, a text-only model may produce structured JSON from a written request. To determine whether a model belongs in this category, evaluate what it directly generates rather than only what it can understand.
How structured output works
A request normally supplies both instructions and a schema. The provider may use constrained decoding, schema-guided generation, validation, or a combination of these techniques to reduce the chance that the response violates the requested structure.
At a conceptual level, the schema acts like a form that the model must fill in. It can restrict the available fields and values, but it cannot independently verify whether the values are true. A schema might require a valid-looking date or restrict a priority field to low, medium, or high; it cannot determine whether the chosen date or priority is appropriate without reliable source information and application logic.
SDKs add another layer. Some allow developers to define schemas using typed classes, Pydantic models, Zod objects, or similar declarations. The SDK may convert those definitions into a provider-compatible schema and parse the result. That convenience should not be confused with a provider guarantee: client-side parsing, provider-side enforcement, and post-processing are separate layers.
Structured output versus related capabilities
JSON mode
JSON mode generally asks the model to return syntactically valid JSON. It may not require a particular set of fields, data types, nesting structure, or allowed values. Structured output usually goes further by targeting adherence to a specified schema. The distinction matters when downstream software expects a particular contract.
Function calling and tool use
Function calling or tool use lets a model emit structured arguments for an external function. For example, the model might produce a customer ID and date range for a reporting function. The application then decides whether to execute that function.
A structured response, by contrast, is generally the final answer returned in the requested format. The two capabilities overlap because both can use schemas, and some providers offer strict schemas for tool arguments. However, a tool call does not mean that the model itself completed the action. The external application, database, search service, or API must execute it and return a result.
Application-level parsing
An application can ask for ordinary prose and then use a parser or another model to convert that prose into JSON. The application may end up with structured data, but this does not prove that the underlying model natively supports structured output. Native schema enforcement, SDK parsing, and application post-processing should be evaluated separately.
Multimodal generation
Structured output concerns the shape of a response. Multimodal generation concerns the direct creation of media such as images, audio, or video. A model can analyze an image and return structured JSON without generating an image, and an image-generation model can create an image without supporting schema-constrained text responses.
Practical uses
Structured output is most valuable when another system needs predictable data rather than an answer intended only for reading.
Extracting information from documents
A user can provide an invoice, contract, résumé, receipt, or form. The model can return fields such as dates, names, totals, clauses, or line items in a predefined structure. The application can then validate the result, store it in a database, or route it for human review.
This is particularly useful when documents vary in layout but the application needs a consistent record. The model handles interpretation; the schema provides a stable destination for the extracted values.
Classification and routing
A customer-support system can ask the model to classify a message into an approved category, assign a priority, identify sentiment, and indicate whether escalation is needed. The structured result can route the message to a queue or trigger a workflow without requiring fragile text matching.
Meeting and project data
Given meeting notes or a transcript, a model can return action items with fields for the task, owner, deadline, status, and priority. A project-management system can then display those items or request confirmation before creating them.
Application interfaces
Structured responses can populate dynamic forms, user-interface components, search filters, content-management fields, or application state. A model might return a list of suggested filters, a set of form fields, or a structured report that the interface renders consistently.
Tool and workflow integration
Agents often need structured arguments for search, database queries, calendars, ticketing systems, or business APIs. Schema-constrained arguments reduce integration errors, but the application should still check permissions, validate values, and decide whether execution is safe.
What matters when comparing structured-output models
The important question is not simply whether a model can produce JSON. Compare the strength and usefulness of its structured-output contract.
- Guarantee level: Determine whether the provider promises valid JSON, adherence to a JSON Schema, strict tool arguments, or only prompt-based formatting.
- Schema support: Check support for nested objects, arrays, enums, nullable fields, unions, recursion, additional properties, numeric constraints, and string formats. Providers commonly support only a subset of JSON Schema.
- Failure behavior: Find out how the API represents refusals, unsupported schemas, invalid requests, truncation, timeouts, and incomplete responses.
- Semantic accuracy: Measure whether extracted fields and classifications are correct. A response can match the schema while containing invented or misinterpreted values.
- Consistency: Test repeated requests, ambiguous inputs, missing information, and model or prompt changes. Structural consistency does not necessarily mean identical field values.
- Latency and streaming: Consider normal response time, first-use schema processing, schema caching, and whether partially streamed responses are practical for the application.
- Limits: Check maximum schema size, nesting depth, response length, array size, context window, and behavior with large documents.
- Tool integration: If tools are involved, verify whether strict schemas apply to function calls, whether parallel calls are supported, and how tool results are represented.
- SDK support: Typed parsing can simplify development, but confirm that the SDK version, language, and endpoint support the features you need.
- Operational cost: Account for schema overhead, token usage, retries, validation, human review, and additional tool or service calls.
Limitations and trade-offs
Correct format does not mean correct information
The most important limitation is that structure is not truth. A model can return a perfectly valid object with the wrong customer, date, amount, diagnosis, category, or conclusion. Applications should validate important values against source documents, business rules, databases, or human review.
Schema restrictions can reduce flexibility
Provider implementations may reject or ignore unsupported schema keywords. Some require every field to be present, while others encourage nullable fields or explicit alternatives. Deeply nested or highly complex schemas can increase latency and make errors harder to diagnose.
Refusals and incomplete responses still occur
A model may refuse an unsafe request instead of returning the requested object. Token limits, interruptions, and timeouts can also leave a response incomplete. Production code should handle refusal representations, missing fields, null values, parsing failures, and tool errors explicitly rather than assuming every request succeeds.
Strict structure is not determinism
Schema enforcement can make the shape of a response predictable while the contents still vary. Ambiguous source material may lead to different classifications or omissions across attempts. Testing should measure both structural validity and substantive correctness.
Privacy and security remain application responsibilities
Structured output does not make sensitive data safe by itself. Applications should consider what documents are sent to a provider, how long data is retained, who can access the resulting records, and whether generated tool arguments could cause unauthorized actions. Strict schemas can reduce accidental integration errors, but they are not a replacement for authentication, authorization, or security testing.
How to evaluate one in practice
- Define a schema that reflects the real application rather than a simplified demonstration.
- Confirm which schema features the selected provider and endpoint officially support.
- Test ordinary examples along with missing, conflicting, ambiguous, malformed, and adversarial inputs.
- Measure schema validity, required-field compliance, enum accuracy, null handling, refusal handling, and truncation.
- Score semantic correctness separately from formatting correctness.
- Check whether the model invents values when the source does not provide enough information.
- Measure latency, token usage, throughput, retries, and behavior near response-size limits.
- Validate every response in application code and define safe behavior for uncertainty, failure, and external actions.
- Repeat the evaluation after changing the model, SDK, schema, prompt, or provider configuration.
Who needs structured output?
Structured output is a strong fit when AI-generated information must enter software reliably. It is useful for document extraction, classification, workflow routing, database population, user-interface generation, and agent integrations.
It may be unnecessary when a person only wants a natural-language explanation, brainstorming, or creative writing. In those cases, a rigid schema can add complexity without much benefit. The capability becomes valuable when predictable fields, validation, and downstream automation matter more than conversational flexibility.
The practical rule is simple: use structured output when the application needs a dependable shape for an AI response, then add separate checks for whether the contents are accurate, safe, complete, and authorized for use.
