GPT-5.6

GPT-5.6 Luna

by OpenAI · current

OpenAI's GPT-5.6 Luna is a current cost-optimized model for high-volume applications. It combines text generation, image understanding, long-context processing, reasoning controls, tool use, structured outputs, streaming, caching and batch API support at $0.20 per million input tokens and $1.20 per million output tokens.

Text Reasoning Coding
GPT-5.6 Luna is a cost-optimized OpenAI model for high-volume inference, agent workflows, classification, summarization, routing, document understanding, coding assistance, and other applications where speed and price efficiency matter. It supports a 1,050,000-token context window, up to 128,000 output tokens, image input, web search, function calling, structured outputs, streaming, caching, batch processing, and several Responses API tools.
Outputs

What GPT-5.6 Luna can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

7/10 Reasoning
8/10 Coding
9/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family GPT-5.6
Model type Lightweight
Context window 1.05M tokens
Maximum output 128K tokens
Knowledge cutoff 2026-02-16
Release date 2026-07-09
Status current
Knowledge cutoff notes

The official model documentation lists February 16, 2026 as the knowledge cutoff. Web search and other external tools can provide newer information during use but do not change the underlying cutoff.

Model notes

GPT-5.6 Luna is the fastest and lowest-cost model in the GPT-5.6 family and roughly corresponds to the nano tier in earlier GPT-5 families. The model supports reasoning effort settings of none, low, medium, high, xhigh and max, with medium as the default. It accepts text and image input and produces text output. OpenAI documents support for function calling, structured outputs, web search, file search, image generation tools, code interpreter, hosted shell, apply patch, skills, computer use, MCP and tool search through the Responses API; these tools do not mean that Luna natively generates images, audio or video. The editorial capability scores are comparative estimates rather than official provider ratings.

Cost

Model pricing

Input $0.20 per 1 million input tokens; cached input $0.02 per 1 million tokens; cache writes billed at 1.25x the uncached input rate. Requests with more than 272,000 input tokens are priced at 2x input for the full request.
Output $1.20 per 1 million output tokens. Requests with more than 272,000 input tokens are priced at 1.5x output for the full request.
Model guide

GPT-5.6 Luna: Pricing, Context Window, Capabilities and API Support

GPT-5.6 Luna is OpenAI's fastest and most cost-efficient GPT-5.6 model, designed for high-volume workloads requiring text and image understanding, long context, reasoning, tool use, structured outputs, and low per-token cost.

What is GPT-5.6 Luna?

GPT-5.6 Luna is an OpenAI model for applications that need to process many requests at relatively low cost. It belongs to the GPT-5.6 family and is positioned as the family's fastest and least expensive model. OpenAI describes it as suitable for high-volume workloads, while the supplied model documentation identifies it as a current API model released on July 9, 2026.

The model accepts text and images as input and produces text as output. That combination makes it useful for tasks such as extracting information from documents, classifying images or text, summarizing long material, routing requests to other systems, and generating structured responses for software to consume.

Luna is not an image, audio, or video generation model. Some Responses API tools can connect a request to capabilities such as image generation or code execution, but those tools should not be confused with native non-text output from Luna itself.

Where Luna fits in OpenAI's lineup

Within the GPT-5.6 family, Luna is the cost and speed-oriented option. The research describes it as roughly corresponding to the nano tier used in earlier GPT-5 families. Its role is therefore different from a model selected primarily for maximum reasoning depth, frontier coding performance, or demanding research tasks.

This positioning creates a straightforward trade-off. Luna is a strong candidate when the application needs many moderately complex operations and the cost of every token matters. A higher-end model may be more appropriate when a smaller improvement in answer quality, difficult multi-step reasoning, or advanced coding performance is worth higher latency and pricing. The supplied research does not provide a complete specification or price comparison for every sibling model, so those choices should be evaluated using the current documentation for the alternatives.

Key specifications

SpecificationGPT-5.6 Luna
ProviderOpenAI
Model familyGPT-5.6
StatusCurrent
Context window1,050,000 tokens
Maximum output128,000 tokens
Input typesText and images
Output typeText
Knowledge cutoffFebruary 16, 2026
Reasoning controlsNone, low, medium, high, xhigh, and max

The context window is the amount of input and generated material the model can handle in one request. At 1.05 million tokens, Luna can work with unusually large collections of text, although practical limits can also depend on application design, tool calls, and how much output is requested. The maximum output allowance is 128,000 tokens, but applications generally should request only the amount of output they actually need.

The February 16, 2026 knowledge cutoff describes the information contained in the underlying model. Web search and other external tools can supply newer information during use, but they do not change the model's cutoff.

Pricing and cost behavior

OpenAI's listed price is $0.20 per 1 million input tokens and $1.20 per 1 million output tokens. Cached input is priced at $0.02 per 1 million tokens. Cache writes are billed at 1.25 times the uncached input rate.

There is an important long-context pricing condition. Requests containing more than 272,000 input tokens are charged at twice the input rate for the full request and at 1.5 times the output rate for the full request. Consequently, the large context window should not be treated as an unlimited low-cost workspace. If a request crosses that threshold, its effective price changes substantially.

For workloads that repeat the same instructions, reference material, or system context, caching can reduce the cost of repeated input. Batch processing is also supported, making Luna suitable for jobs that do not require an immediate response. The supplied research does not specify a separate batch discount, so no additional reduction should be assumed.

Reasoning and coding capabilities

Luna supports selectable reasoning effort levels: none, low, medium, high, xhigh, and max, with medium identified as the default. Reasoning effort controls how much computation the model applies to a request. Lower settings can favor speed and cost, while higher settings may be useful for more involved problems. The setting is not a guarantee of correctness, and the supplied research does not provide benchmark results that would quantify the difference between levels.

Its most appropriate coding uses are routine coding assistance, code transformation, extraction of structured information from source files, and tool-using automation. The editorial coding score supplied for Luna is 8 out of 10, but this is a comparative editorial estimate rather than an OpenAI-published rating. It should not be interpreted as an official benchmark result.

For difficult software architecture, complex debugging, or tasks where a subtle reasoning error is expensive, a higher-end model may be a better choice. Luna's advantage is usually the ability to handle many coding-related requests economically, not a claim of being the strongest coding model in every situation.

Tools and API support

Luna supports function calling, which lets an application describe external functions and allow the model to request them with structured arguments. This is useful for operations such as looking up an order, querying a database, creating a ticket, or invoking an internal workflow. The application remains responsible for executing the function and validating its arguments.

The model also supports structured outputs, allowing developers to request responses that follow a defined schema. Structured output is useful for classification, extraction, routing, and data pipelines because software can consume the result more reliably than free-form prose. It should not automatically be treated as a separate JSON mode; the supplied research lists JSON mode as unspecified while separately confirming structured-output support.

According to the supplied notes, the Responses API documents support for web search, file search, image-generation tools, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. Streaming is supported as well, so applications can receive generated text progressively rather than waiting for the complete response.

These integrations expand what an application can do with Luna, but they do not change its native modality. Luna itself accepts text and images and returns text. For example, an image-generation tool may be available in a workflow, while Luna remains the model that interprets the request or coordinates the tool call.

Main strengths and limitations

Strengths

  • Low token pricing: The listed input and output rates make Luna suitable for high-volume processing.
  • Fast positioning: OpenAI positions it as the fastest model in the GPT-5.6 family.
  • Very long context: The 1.05-million-token context window supports large document collections and extended workflows.
  • Image understanding: Image input enables visual document analysis and other text-and-image tasks.
  • Flexible automation: Function calling, structured outputs, streaming, caching, batch processing, and Responses API tools support production pipelines.
  • Adjustable reasoning: Multiple reasoning effort settings let developers trade computation against speed and cost.

Limitations

  • Text-only output: Luna does not natively generate images, audio, or video.
  • No fine-tuning: Fine-tuning is listed as unsupported, so customization must use prompting, tools, retrieval, or application-level logic.
  • Long-context surcharge: Requests above 272,000 input tokens receive higher full-request input and output rates.
  • Not the highest-end option: The model is optimized for speed and cost, so users seeking maximum research, reasoning, or frontier coding performance may prefer another model.
  • Knowledge cutoff: Its built-in knowledge ends on February 16, 2026; current facts require web search or another external data source.
  • Potential model error: Tool support and structured responses improve application reliability but do not guarantee that the model's interpretation or conclusions are correct.

Best use cases for GPT-5.6 Luna

Luna is particularly well matched to workloads where each individual request is manageable but the total request volume is large. Examples include:

  • Classifying support tickets, documents, images, or user requests.
  • Summarizing reports, transcripts, research collections, and long business documents.
  • Extracting fields from invoices, forms, contracts, or image-based documents.
  • Routing requests to specialist agents, tools, queues, or business systems.
  • Running routine coding assistance, code explanation, and transformation workflows.
  • Building tool-using agents that need structured function arguments and text responses.
  • Processing large batches of records when immediate results are unnecessary.
  • Analyzing long context while keeping per-token costs under control.

For example, a document-processing pipeline could send a scanned form and its accompanying text to Luna, request a defined schema, and then use function calling to store the extracted fields. A support system could classify thousands of incoming messages, summarize each conversation, and route only the difficult cases to a more capable model or a human reviewer.

When to choose GPT-5.6 Luna

Choose GPT-5.6 Luna when speed, high throughput, long context, and predictable token economics matter more than obtaining the deepest possible reasoning on every request. It is a practical default for classification, summarization, extraction, routing, document understanding, and routine agent automation.

Consider a different type of model when the task is unusually ambiguous, requires advanced mathematical or scientific reasoning, depends on frontier-level coding, or has a high cost of failure. A more capable model may justify its higher price when each answer needs extensive verification or when a single difficult problem matters more than processing volume. Conversely, a smaller or more specialized model may be preferable for simple deterministic transformations where Luna's broader capabilities are unnecessary.

The most useful deployment pattern may be a tiered workflow: use Luna for the large majority of ordinary requests, then escalate difficult, uncertain, or high-impact cases. This approach takes advantage of Luna's speed and cost profile without assuming that the least expensive model is appropriate for every task.

Bottom line

GPT-5.6 Luna is a text-output model built for efficient scale. Its combination of a 1.05-million-token context window, image input, adjustable reasoning, tool calling, structured outputs, caching, batch support, and low listed token rates makes it well suited to production workloads that process many documents or requests. Its trade-offs are equally clear: it does not natively produce non-text media, it cannot be fine-tuned according to the supplied specifications, long requests above 272,000 input tokens cost more, and higher-end alternatives may be better for the hardest reasoning and coding tasks.


Answers to Frequently Asked Questions

Does GPT-5.6 Luna generate images, audio, or video?
No. GPT-5.6 Luna is a text-output model and does not natively generate images, audio, or video. An application may connect it to tools such as image generation, but those tools are separate from Luna's native capabilities.
Does GPT-5.6 Luna support images, function calling, and structured outputs?
Yes. Luna accepts text and images as input, supports function calling and structured outputs, and can be used with tools through the Responses API, including web search, file search, code interpreter, and other integrations. Its native output is text.
How much does GPT-5.6 Luna cost?
GPT-5.6 Luna costs $0.20 per 1 million input tokens and $1.20 per 1 million output tokens. Cached input costs $0.02 per 1 million tokens, while cache writes cost 1.25 times the standard uncached input rate.
What is GPT-5.6 Luna's context window and maximum output?
GPT-5.6 Luna has a context window of 1,050,000 tokens and supports a maximum output of 128,000 tokens. Requests containing more than 272,000 input tokens are subject to higher full-request input and output pricing.
What is GPT-5.6 Luna best used for?
GPT-5.6 Luna is designed for high-volume, cost-sensitive workloads such as document classification, summarization, information extraction, image understanding, request routing, routine coding assistance, and tool-using automation.


Sources 5
Provider

About OpenAI