What is GPT-5.6 Luna?
GPT-5.6 Luna is an OpenAI model for applications that need to process many requests at relatively low cost. It belongs to the GPT-5.6 family and is positioned as the family's fastest and least expensive model. OpenAI describes it as suitable for high-volume workloads, while the supplied model documentation identifies it as a current API model released on July 9, 2026.
The model accepts text and images as input and produces text as output. That combination makes it useful for tasks such as extracting information from documents, classifying images or text, summarizing long material, routing requests to other systems, and generating structured responses for software to consume.
Luna is not an image, audio, or video generation model. Some Responses API tools can connect a request to capabilities such as image generation or code execution, but those tools should not be confused with native non-text output from Luna itself.
Where Luna fits in OpenAI's lineup
Within the GPT-5.6 family, Luna is the cost and speed-oriented option. The research describes it as roughly corresponding to the nano tier used in earlier GPT-5 families. Its role is therefore different from a model selected primarily for maximum reasoning depth, frontier coding performance, or demanding research tasks.
This positioning creates a straightforward trade-off. Luna is a strong candidate when the application needs many moderately complex operations and the cost of every token matters. A higher-end model may be more appropriate when a smaller improvement in answer quality, difficult multi-step reasoning, or advanced coding performance is worth higher latency and pricing. The supplied research does not provide a complete specification or price comparison for every sibling model, so those choices should be evaluated using the current documentation for the alternatives.
Key specifications
| Specification | GPT-5.6 Luna |
|---|---|
| Provider | OpenAI |
| Model family | GPT-5.6 |
| Status | Current |
| Context window | 1,050,000 tokens |
| Maximum output | 128,000 tokens |
| Input types | Text and images |
| Output type | Text |
| Knowledge cutoff | February 16, 2026 |
| Reasoning controls | None, low, medium, high, xhigh, and max |
The context window is the amount of input and generated material the model can handle in one request. At 1.05 million tokens, Luna can work with unusually large collections of text, although practical limits can also depend on application design, tool calls, and how much output is requested. The maximum output allowance is 128,000 tokens, but applications generally should request only the amount of output they actually need.
The February 16, 2026 knowledge cutoff describes the information contained in the underlying model. Web search and other external tools can supply newer information during use, but they do not change the model's cutoff.
Pricing and cost behavior
OpenAI's listed price is $0.20 per 1 million input tokens and $1.20 per 1 million output tokens. Cached input is priced at $0.02 per 1 million tokens. Cache writes are billed at 1.25 times the uncached input rate.
There is an important long-context pricing condition. Requests containing more than 272,000 input tokens are charged at twice the input rate for the full request and at 1.5 times the output rate for the full request. Consequently, the large context window should not be treated as an unlimited low-cost workspace. If a request crosses that threshold, its effective price changes substantially.
For workloads that repeat the same instructions, reference material, or system context, caching can reduce the cost of repeated input. Batch processing is also supported, making Luna suitable for jobs that do not require an immediate response. The supplied research does not specify a separate batch discount, so no additional reduction should be assumed.
Reasoning and coding capabilities
Luna supports selectable reasoning effort levels: none, low, medium, high, xhigh, and max, with medium identified as the default. Reasoning effort controls how much computation the model applies to a request. Lower settings can favor speed and cost, while higher settings may be useful for more involved problems. The setting is not a guarantee of correctness, and the supplied research does not provide benchmark results that would quantify the difference between levels.
Its most appropriate coding uses are routine coding assistance, code transformation, extraction of structured information from source files, and tool-using automation. The editorial coding score supplied for Luna is 8 out of 10, but this is a comparative editorial estimate rather than an OpenAI-published rating. It should not be interpreted as an official benchmark result.
For difficult software architecture, complex debugging, or tasks where a subtle reasoning error is expensive, a higher-end model may be a better choice. Luna's advantage is usually the ability to handle many coding-related requests economically, not a claim of being the strongest coding model in every situation.
Tools and API support
Luna supports function calling, which lets an application describe external functions and allow the model to request them with structured arguments. This is useful for operations such as looking up an order, querying a database, creating a ticket, or invoking an internal workflow. The application remains responsible for executing the function and validating its arguments.
The model also supports structured outputs, allowing developers to request responses that follow a defined schema. Structured output is useful for classification, extraction, routing, and data pipelines because software can consume the result more reliably than free-form prose. It should not automatically be treated as a separate JSON mode; the supplied research lists JSON mode as unspecified while separately confirming structured-output support.
According to the supplied notes, the Responses API documents support for web search, file search, image-generation tools, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Streaming is supported as well, so applications can receive generated text progressively rather than waiting for the complete response.
These integrations expand what an application can do with Luna, but they do not change its native modality. Luna itself accepts text and images and returns text. For example, an image-generation tool may be available in a workflow, while Luna remains the model that interprets the request or coordinates the tool call.
Main strengths and limitations
Strengths
- Low token pricing: The listed input and output rates make Luna suitable for high-volume processing.
- Fast positioning: OpenAI positions it as the fastest model in the GPT-5.6 family.
- Very long context: The 1.05-million-token context window supports large document collections and extended workflows.
- Image understanding: Image input enables visual document analysis and other text-and-image tasks.
- Flexible automation: Function calling, structured outputs, streaming, caching, batch processing, and Responses API tools support production pipelines.
- Adjustable reasoning: Multiple reasoning effort settings let developers trade computation against speed and cost.
Limitations
- Text-only output: Luna does not natively generate images, audio, or video.
- No fine-tuning: Fine-tuning is listed as unsupported, so customization must use prompting, tools, retrieval, or application-level logic.
- Long-context surcharge: Requests above 272,000 input tokens receive higher full-request input and output rates.
- Not the highest-end option: The model is optimized for speed and cost, so users seeking maximum research, reasoning, or frontier coding performance may prefer another model.
- Knowledge cutoff: Its built-in knowledge ends on February 16, 2026; current facts require web search or another external data source.
- Potential model error: Tool support and structured responses improve application reliability but do not guarantee that the model's interpretation or conclusions are correct.
Best use cases for GPT-5.6 Luna
Luna is particularly well matched to workloads where each individual request is manageable but the total request volume is large. Examples include:
- Classifying support tickets, documents, images, or user requests.
- Summarizing reports, transcripts, research collections, and long business documents.
- Extracting fields from invoices, forms, contracts, or image-based documents.
- Routing requests to specialist agents, tools, queues, or business systems.
- Running routine coding assistance, code explanation, and transformation workflows.
- Building tool-using agents that need structured function arguments and text responses.
- Processing large batches of records when immediate results are unnecessary.
- Analyzing long context while keeping per-token costs under control.
For example, a document-processing pipeline could send a scanned form and its accompanying text to Luna, request a defined schema, and then use function calling to store the extracted fields. A support system could classify thousands of incoming messages, summarize each conversation, and route only the difficult cases to a more capable model or a human reviewer.
When to choose GPT-5.6 Luna
Choose GPT-5.6 Luna when speed, high throughput, long context, and predictable token economics matter more than obtaining the deepest possible reasoning on every request. It is a practical default for classification, summarization, extraction, routing, document understanding, and routine agent automation.
Consider a different type of model when the task is unusually ambiguous, requires advanced mathematical or scientific reasoning, depends on frontier-level coding, or has a high cost of failure. A more capable model may justify its higher price when each answer needs extensive verification or when a single difficult problem matters more than processing volume. Conversely, a smaller or more specialized model may be preferable for simple deterministic transformations where Luna's broader capabilities are unnecessary.
The most useful deployment pattern may be a tiered workflow: use Luna for the large majority of ordinary requests, then escalate difficult, uncertain, or high-impact cases. This approach takes advantage of Luna's speed and cost profile without assuming that the least expensive model is appropriate for every task.
Bottom line
GPT-5.6 Luna is a text-output model built for efficient scale. Its combination of a 1.05-million-token context window, image input, adjustable reasoning, tool calling, structured outputs, caching, batch support, and low listed token rates makes it well suited to production workloads that process many documents or requests. Its trade-offs are equally clear: it does not natively produce non-text media, it cannot be fine-tuned according to the supplied specifications, long requests above 272,000 input tokens cost more, and higher-end alternatives may be better for the hardest reasoning and coding tasks.

