What is GPT-6 Luna?
GPT-6 Luna is OpenAI’s efficiency-focused model in the GPT-6 family. Its primary purpose is to handle focused, repeatable tasks at high volume while keeping token costs low. The model is available through the OpenAI API under the canonical model ID gpt-6-luna.
OpenAI announced GPT-6 Luna on September 22, 2026, alongside GPT-6 Sol. The company positions Luna below its higher-capability GPT-6 models on the cost-capability curve: it is intended for workloads where consistent reasoning, coding, tool use, and long-context processing matter, but maximum frontier capability is not required for every request.
This positioning makes GPT-6 Luna different from a general-purpose consumer chatbot description. It is primarily an API model for building applications and automated workflows, although OpenAI also announced access through ChatGPT Work and Codex for selected paid, business, education, and desktop-app users.
GPT-6 Luna specifications at a glance
| Specification | GPT-6 Luna |
|---|---|
| Provider | OpenAI |
| Model ID | gpt-6-luna |
| Model type | Reasoning model |
| Native input | Text and images |
| Native output | Text |
| Context window | 1,050,000 tokens |
| Maximum output | 128,000 tokens |
| Reasoning effort | None, low, medium, high, xhigh, or max |
| Default reasoning effort | Medium |
| Fine-tuning | Not supported |
| Knowledge cutoff | May 18, 2026 |
The specifications above describe the model itself. Platform tools can extend what an application does, but they should not be confused with native model output. For example, a workflow can use image generation or computer-use tools through the Responses API, while GPT-6 Luna itself still returns text rather than images or action-control signals.
Modalities and context window
GPT-6 Luna accepts text and image input and produces text output. This allows an application to combine ordinary prompts with screenshots, photographs, diagrams, scanned pages, or other supported image content. Audio and video are not listed as native input modalities, and the model does not directly generate images, audio, or video.
The model has a 1,050,000-token context window. A context window is the amount of input and generated material the model can consider in one request or conversation state. In practical terms, this gives developers room to send very long documents, large collections of related files, substantial code repositories, or accumulated agent context without splitting every task into small independent calls.
The maximum output is 128,000 tokens. That limit is substantially larger than the response length needed for most summaries, classifications, or coding tasks, but it can be useful when the model must produce a long structured result, extensive analysis, or a large code change. The maximum does not mean every request will use or need that many tokens; applications should still set sensible output limits to control cost and latency.
OpenAI documents May 18, 2026 as GPT-6 Luna’s knowledge cutoff. Web search and other tools may provide newer information during a request, but tool access does not change the model’s underlying training cutoff.
Reasoning and coding capabilities
GPT-6 Luna supports configurable reasoning effort. Developers can choose none, low, medium, high, xhigh, or max, with medium documented as the default. Reasoning effort is a practical control over how much processing the model applies before producing an answer. Lower settings can reduce latency and cost for straightforward work; higher settings are more appropriate for difficult analysis, complex code, or multi-step decisions.
The choice is a trade-off rather than a universal quality setting. A high reasoning level may be unnecessary for routine extraction or classification, while a low level may be unsuitable for complicated debugging or an agent task that must reconcile many constraints. Production applications can route simple requests to lower effort and reserve higher settings for cases that fail validation or require deeper analysis.
GPT-6 Luna is suitable for coding assistance, code explanation, repository analysis, test generation, structured code changes, and software-agent sub-tasks. OpenAI’s launch material reports that Luna improved over GPT-5.6 Luna on AutomationBench and achieved a 66.6% score on DeepSWE v1.1 at maximum reasoning effort. These are provider-reported evaluation results, not guarantees of performance on a particular codebase. Developers should test representative languages, frameworks, repository sizes, and failure modes before relying on the model for unattended changes.
Tools, function calling and API features
GPT-6 Luna supports streaming, function calling, structured outputs, prompt caching, and batch processing. Structured outputs allow an application to request a defined data structure instead of relying only on free-form prose. This is useful for extraction pipelines, classification systems, database preparation, and workflows that pass model results to another program.
OpenAI recommends the Responses API when using built-in tools or function calling with non-none reasoning effort. Function calling is also available through Chat Completions when reasoning effort is set to none. Function calling does not mean that the model independently performs arbitrary actions. Instead, the model proposes a tool call, and the surrounding application decides whether to execute it and what result to return.
The Responses API lists support for web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. These capabilities can make a Luna-based application more useful, but they are platform-level tools. They should not be recorded as direct image, audio, video, or executable output from the base model.
Fine-tuning is not supported for the exact GPT-6 Luna model. Structured outputs are supported. JSON mode is a separate API capability from structured outputs, and the supplied model documentation does not independently confirm JSON mode for GPT-6 Luna. Developers who specifically require JSON mode should verify the current endpoint documentation rather than assuming that structured-output support means the two features are identical.
GPT-6 Luna pricing
GPT-6 Luna’s standard short-context pricing is $0.10 per 1 million input tokens and $0.50 per 1 million output tokens. Cached input is priced at $0.01 per 1 million tokens, while cache writes cost $0.125 per 1 million tokens.
OpenAI lists separate long-context rates: $0.20 per 1 million input tokens, $0.02 per 1 million cached input tokens, $0.25 per 1 million cache writes, and $0.75 per 1 million output tokens. Requests with more than 272,000 input tokens receive this additional pricing treatment. A long-context request can therefore remain inexpensive relative to larger models, but developers should estimate costs using the applicable long-context rates when sending very large prompts.
| Usage type | Short-context rate | Long-context rate |
|---|---|---|
| Input | $0.10 per 1M tokens | $0.20 per 1M tokens |
| Cached input | $0.01 per 1M tokens | $0.02 per 1M tokens |
| Cache writes | $0.125 per 1M tokens | $0.25 per 1M tokens |
| Output | $0.50 per 1M tokens | $0.75 per 1M tokens |
OpenAI states that Batch and Flex processing are priced at 50% of standard rates, while Fast mode costs twice the applicable rates. Regional processing may add a 10% premium where available. European Union data residency for GPT-6 Luna is limited to Standard processing. These options make the model’s economics dependent not only on token volume but also on latency requirements, processing mode, caching behavior, and request size.
Strengths and trade-offs
GPT-6 Luna’s main strength is the combination of low token pricing, a very large context window, configurable reasoning, and broad tool support. That combination is valuable when an application processes many similar requests or must repeatedly inspect large amounts of information. Prompt caching can further reduce the cost of repeated instructions or shared context.
- Cost: Standard input and output prices are substantially lower than the listed rates for GPT-6 Sol and GPT-6 Astra.
- Scale: Batch processing, caching, and low token rates support high request volumes.
- Long context: The 1.05-million-token window can reduce the need to divide large documents or repositories into many separate calls.
- Control: Reasoning effort settings let developers balance response quality, speed, and cost by task.
- Integration: Function calling, structured outputs, streaming, web search, file search, code execution, and other Responses API tools support application workflows.
The trade-off is that Luna is not OpenAI’s highest-capability GPT-6 model. OpenAI recommends GPT-6 Astra when maximum capability is the priority and GPT-6 Sol for stronger reasoning on especially demanding tasks. Choosing Luna makes more sense when the workload is large, repeatable, or cost-sensitive and does not require the best available result on every difficult prompt.
Luna also has clear modality limits. It can inspect images but does not natively return images, audio, or video. It is not fine-tunable according to the supplied model documentation. Its knowledge cutoff means that current facts may require web search or another retrieval system, and all model-generated results should be checked when errors have meaningful consequences.
Best use cases for GPT-6 Luna
GPT-6 Luna is a strong fit for workloads where each individual request is important but the overall system must process many requests efficiently. Suitable examples include:
- Classifying and routing large volumes of support tickets, documents, or business records.
- Extracting fields, entities, decisions, or compliance information into structured outputs.
- Summarizing long reports, contracts, technical documentation, or collections of related files.
- Analyzing code repositories and generating explanations, tests, or proposed patches.
- Supporting retrieval-augmented generation systems that combine retrieved source material with model reasoning.
- Handling sub-tasks inside agents, such as planning a bounded step, calling a tool, or checking an intermediate result.
- Processing repeated image-and-text requests, such as reviewing diagrams or document scans alongside instructions.
For example, a document-processing service could use image input for scanned pages, structured outputs for extracted fields, prompt caching for a repeated schema and policy definition, and Batch processing for non-urgent work. A coding system could use low or medium reasoning for routine code explanations and escalate difficult repository changes to a higher reasoning setting.
When to choose GPT-6 Luna
Choose GPT-6 Luna when you need a capable reasoning model for high-volume work and the cost of every request matters. It is particularly attractive when your application needs a long context, text-and-image input, structured results, tool calling, or batch processing, but does not require native media generation.
Consider a different option when maximum reasoning quality is more important than price or throughput. Within OpenAI’s GPT-6 lineup, GPT-6 Astra is the better comparison for frontier capability, while GPT-6 Sol is positioned for stronger reasoning on demanding tasks. A practical architecture may use Luna for routine requests and reserve a more capable sibling model for escalations, difficult edge cases, or quality-sensitive outputs.
GPT-6 Luna is also a poor fit when the required output is natively audio, video, or an image. Although the Responses API can expose tools such as image generation and computer use, those tools do not change Luna’s native text-output modality. Similarly, teams that need to customize model weights through fine-tuning should not select Luna based on an assumption that fine-tuning is available.
Availability and final assessment
GPT-6 Luna is available in the OpenAI API as gpt-6-luna. OpenAI announced access in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, with access for Free and Go users through the desktop app. The launch announcement stated that the models were not initially available broadly in ChatGPT and would roll out gradually, so availability may depend on account type, product surface, and rollout status.
Overall, GPT-6 Luna is best understood as an efficiency-oriented reasoning model rather than a universal replacement for every GPT-6 option. Its unusually large context window, low standard token prices, adjustable reasoning effort, and extensive tool support make it well suited to production pipelines and repeatable agent workflows. Its limitations—text-only output, no documented fine-tuning, a fixed knowledge cutoff, and lower positioning than GPT-6 Sol and GPT-6 Astra—define when its lower cost is worth choosing over a higher-capability alternative.

