GPT-5.6 Terra is OpenAI's balanced model in the GPT-5.6 family. Its purpose is to provide a practical middle ground between capability, response speed, and operating cost. It is aimed at production software, coding agents, research workflows, document analysis, and business automation where a model must handle substantial reasoning and tool use but does not need the most expensive model in the family.
Terra is not a general-purpose consumer product by itself. It is an API model that developers can call using its canonical model ID, gpt-5.6-terra. The specifications and prices below describe the supplied OpenAI model documentation and catalog information. Comparative comments about value, speed, or suitability are editorial evaluations based on those specifications rather than provider-published ratings.
What is GPT-5.6 Terra?
GPT-5.6 Terra is a general-purpose language model from OpenAI with support for reasoning, coding, structured responses, and tool-enabled workflows. OpenAI positions it between the flagship GPT-5.6 Sol and the faster, lower-cost GPT-5.6 Luna. The supplied research describes Terra as roughly comparable to the mini tier used in earlier GPT-5 families.
That positioning matters because Terra is designed for workloads where the cheapest available model may produce insufficiently reliable reasoning, while the flagship model would make every request unnecessarily expensive. Typical examples include an agent that reads long technical documents, a coding assistant that must use external tools, or an internal automation system that produces consistent structured records.
Technical specifications and limits
Terra has a documented context window of 1,050,000 tokens. A context window is the amount of input material the model can consider in one request, including instructions, conversation history, retrieved documents, tool results, and other supplied content. A context size of 1.05 million tokens is suitable for very large document collections and extended multi-step workflows, although sending very large requests can increase cost.
The maximum output is 128,000 tokens. This is an upper limit rather than a recommended size for every response. Most ordinary answers, code changes, and structured records will use much less, but the limit gives developers room for long analyses or substantial generated content.
| Specification | GPT-5.6 Terra |
|---|---|
| Provider | OpenAI |
| API model ID | gpt-5.6-terra |
| Availability | Generally available |
| Release date | July 9, 2026 |
| Context window | 1,050,000 tokens |
| Maximum output | 128,000 tokens |
| Knowledge cutoff | February 16, 2026 |
| Input types | Text and images |
| Output type | Text |
The knowledge cutoff is separate from tool access. Terra's underlying training knowledge is documented through February 16, 2026. If an application uses web search or another retrieval system, the model can process newer information supplied through that tool, but this does not change the model's underlying cutoff.
Pricing and cost considerations
The supplied OpenAI model catalog lists GPT-5.6 Terra at $2.00 per million input tokens and $12.00 per million output tokens. Cached input tokens are listed at $0.20 per million tokens. Cache writes are billed at 1.25 times the uncached input rate. Requests containing more than 272,000 input tokens receive higher pricing for the full request.
| Token category | Listed price |
|---|---|
| Input tokens | $2.00 per 1 million |
| Cached input tokens | $0.20 per 1 million |
| Output tokens | $12.00 per 1 million |
| Cache writes | 1.25 times the uncached input rate |
These are usage-based API prices, not a monthly subscription fee. The difference between input and output pricing is important for applications that generate long responses: output tokens cost substantially more than ordinary input tokens. Reusing stable instructions or reference material through prompt caching can reduce the cost of repeated requests, while very large prompts need special attention because the higher-pricing threshold applies above 272,000 input tokens.
Compared with GPT-5.6 Sol, Terra is the cost-conscious choice within the same family. Compared with a smaller or speed-oriented model such as GPT-5.6 Luna, Terra is better suited to tasks where deeper reasoning, longer context, or more capable tool use justifies additional spending. The supplied material does not provide benchmark figures, so these are positioning and cost trade-offs rather than measured performance claims.
Reasoning and coding capabilities
Terra supports configurable reasoning effort levels: none, low, medium, high, xhigh, and max. Medium is the documented default. Reasoning effort controls how much deliberate processing the model applies to a request. Lower settings can be appropriate for straightforward classification, extraction, or short answers; higher settings are more relevant to complex analysis, planning, debugging, and multi-step problem solving.
This flexibility helps developers balance latency and cost against answer quality. A production workflow might use a lower setting for routine records and increase the setting only when a task involves ambiguous requirements, difficult code, or several dependent decisions. The supplied research does not give benchmark scores for reasoning or coding, so Terra should not be described as having a verified numerical advantage over other models.
For coding applications, Terra can generate and explain code, analyze technical material, work through debugging tasks, and participate in tool-enabled coding agents. Its support for function calling, structured outputs, code interpretation, hosted shell access, and apply-patch workflows makes it suitable for applications that need more than plain text completion. Developers should still test generated code and apply normal security controls before allowing an agent to change files, execute commands, or access production systems.
Input, output, and tool support
GPT-5.6 Terra accepts text and image input and produces text output. It is therefore multimodal on the input side: an application can provide an image alongside instructions for analysis. It is not documented as a native image, audio, video, speech, music, or embedding-output model. Image generation can be accessed as a supported tool, but that is different from Terra directly producing native image output.
The documented capability set includes streaming, function calling, structured outputs, batch processing, and prompt caching. Structured outputs are useful when an application needs predictable fields, such as an invoice record, a classification result, or a machine-readable workflow decision. Function calling allows the model to request application-defined operations, while streaming lets an application display or process generated text as it arrives.
The supplied Responses API tool list includes web search, file search, code interpreter, hosted shell, computer use, MCP, tool search, image-generation tool access, apply patch, and skills. Tool availability and safe use depend on the surrounding implementation. A tool-enabled model does not independently grant access to a user's files, computer, databases, or external services; the application must provide and control those connections.
Main strengths and limitations
Where Terra is strongest
- Balanced cost and capability: It is positioned below the flagship GPT-5.6 Sol in cost while retaining substantial reasoning and tool support.
- Very long context: The 1.05-million-token window supports large document, repository, and research workflows.
- Adjustable reasoning: Six reasoning effort levels allow developers to tune the trade-off between deliberation, latency, and spending.
- Production-oriented interfaces: Streaming, function calling, structured outputs, batch processing, and caching support application development beyond simple chat.
- Image understanding: Text-and-image input allows the model to analyze visual material while keeping text as its output format.
Where Terra is not the right fit
- It is not a native media generator: Terra does not directly produce image, audio, video, speech, music, or embedding outputs. Image generation is available through a tool rather than as the model's native output mode.
- No documented fine-tuning: The supplied model information lists fine-tuning as unsupported, so applications that require provider-hosted customization should consider another approach.
- High-end tasks may favor the flagship: If maximum capability is more important than cost, GPT-5.6 Sol may be the more appropriate sibling option.
- Simple, high-volume tasks may not need Terra: For short classification, basic extraction, or latency-sensitive requests, a smaller or faster model may offer better economics.
- Tool use adds operational risk: Computer use, shell access, file operations, and external tools require permissions, monitoring, validation, and safeguards.
Best use cases for GPT-5.6 Terra
Terra is a good fit when the task combines meaningful reasoning with enough volume or duration that flagship pricing would be difficult to justify. Practical examples include:
- Long-context analysis of contracts, technical manuals, research collections, or internal policies.
- Coding agents that inspect repositories, plan changes, run tools, and return structured results.
- Research assistants that combine document retrieval, web search, synthesis, and citations or evidence records.
- Business automation that extracts information from text and images into a fixed schema.
- Multi-step workflows that require function calls, code execution, hosted shell operations, or computer-use actions.
- Batch processing where reasoning quality matters but every item does not require the highest-priced model.
When to choose this model
Choose GPT-5.6 Terra when you need a strong general-purpose API model with long context, adjustable reasoning, coding ability, image understanding, and broad tool support, but need to control per-request cost. It is especially suitable when the same system must handle both routine production requests and more demanding analysis without switching constantly between unrelated models.
Choose GPT-5.6 Sol instead when the application consistently prioritizes the highest capability available within the GPT-5.6 family and can accept higher pricing. Choose GPT-5.6 Luna or another speed- and cost-focused option when requests are short, predictable, and relatively simple, or when latency and request volume matter more than maximum reasoning depth. Choose a specialized image, audio, video, speech, or embedding model when native non-text output or vector generation is the central requirement.
Before deployment, test Terra with the application's real prompts, document sizes, tool calls, and required output schemas. Pay particular attention to output-token usage, requests above the 272,000-token pricing threshold, the effect of reasoning settings on latency and cost, and the safeguards needed for tools that can affect external systems.

