GPT-6

GPT-6 Luna

by OpenAI · Current and available

GPT-6 Luna is OpenAI’s cost-efficient GPT-6 reasoning model for high-volume workloads. It supports text and image input, a 1.05-million-token context window, 128,000-token outputs, configurable reasoning, coding, tool calling, structured outputs, web search, caching, streaming, and batch processing.

Text Reasoning Coding
GPT-6 Luna is an efficient reasoning model from OpenAI designed for applications that need capable analysis, coding, long-context processing, and tool use without the cost of OpenAI’s higher-capability GPT-6 models. It is especially suited to repeatable workloads such as document processing, classification, extraction, coding assistance, retrieval-augmented generation, and agent sub-tasks.
Outputs

What GPT-6 Luna can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Structured output Prompt caching Batch API
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family GPT-6
Model type Reasoning
Context window 1.05M tokens
Maximum output 128K tokens
Knowledge cutoff 2026-05-18
Release date 2026-09-22
Status Current and available
Knowledge cutoff notes

The official GPT-6 Luna model documentation specifies May 18, 2026 as the model’s knowledge cutoff. Web search and other tools can provide newer information during use but do not change the underlying cutoff.

Model notes

GPT-6 Luna is OpenAI’s efficiency-focused GPT-6 model for focused, high-volume tasks. The canonical API model ID is gpt-6-luna. It supports reasoning effort values none, low, medium, high, xhigh, and max, with medium as the documented default. The model accepts text and image input and returns text. Audio and video are not supported as native modalities. Platform tools including image generation and computer use are available through the Responses API, but those tools do not make the base model a native image, video, audio, or action-output model. OpenAI lists structured outputs as supported and fine-tuning as unsupported. JSON mode is not independently confirmed for this exact model page and is therefore left unknown. Batch and Flex processing are priced at 50% of standard rates; Fast mode is priced at twice the applicable rates. Editorial scores are comparative estimates, not official OpenAI ratings.

Cost

Model pricing

Input $0.10 per 1 million input tokens; $0.01 cached input; $0.125 cache writes. Long-context rates are $0.20 input, $0.02 cached input, and $0.25 cache writes per 1 million tokens.
Output $0.50 per 1 million output tokens for short context; $0.75 per 1 million output tokens for long context.
Model guide

GPT-6 Luna: Pricing, Context, Capabilities and API Availability

GPT-6 Luna is OpenAI’s efficiency-focused GPT-6 reasoning model for high-volume, repeatable workloads. It accepts text and image input, produces text, supports a 1.05-million-token context window and 128,000-token outputs, and provides reasoning controls, tool calling, web search, structured outputs, caching, streaming, and batch processing at substantially lower token prices than GPT-6 Sol and GPT-6 Astra.

What is GPT-6 Luna?

GPT-6 Luna is OpenAI’s efficiency-focused model in the GPT-6 family. Its primary purpose is to handle focused, repeatable tasks at high volume while keeping token costs low. The model is available through the OpenAI API under the canonical model ID gpt-6-luna.

OpenAI announced GPT-6 Luna on September 22, 2026, alongside GPT-6 Sol. The company positions Luna below its higher-capability GPT-6 models on the cost-capability curve: it is intended for workloads where consistent reasoning, coding, tool use, and long-context processing matter, but maximum frontier capability is not required for every request.

This positioning makes GPT-6 Luna different from a general-purpose consumer chatbot description. It is primarily an API model for building applications and automated workflows, although OpenAI also announced access through ChatGPT Work and Codex for selected paid, business, education, and desktop-app users.

GPT-6 Luna specifications at a glance

SpecificationGPT-6 Luna
ProviderOpenAI
Model IDgpt-6-luna
Model typeReasoning model
Native inputText and images
Native outputText
Context window1,050,000 tokens
Maximum output128,000 tokens
Reasoning effortNone, low, medium, high, xhigh, or max
Default reasoning effortMedium
Fine-tuningNot supported
Knowledge cutoffMay 18, 2026

The specifications above describe the model itself. Platform tools can extend what an application does, but they should not be confused with native model output. For example, a workflow can use image generation or computer-use tools through the Responses API, while GPT-6 Luna itself still returns text rather than images or action-control signals.

Modalities and context window

GPT-6 Luna accepts text and image input and produces text output. This allows an application to combine ordinary prompts with screenshots, photographs, diagrams, scanned pages, or other supported image content. Audio and video are not listed as native input modalities, and the model does not directly generate images, audio, or video.

The model has a 1,050,000-token context window. A context window is the amount of input and generated material the model can consider in one request or conversation state. In practical terms, this gives developers room to send very long documents, large collections of related files, substantial code repositories, or accumulated agent context without splitting every task into small independent calls.

The maximum output is 128,000 tokens. That limit is substantially larger than the response length needed for most summaries, classifications, or coding tasks, but it can be useful when the model must produce a long structured result, extensive analysis, or a large code change. The maximum does not mean every request will use or need that many tokens; applications should still set sensible output limits to control cost and latency.

OpenAI documents May 18, 2026 as GPT-6 Luna’s knowledge cutoff. Web search and other tools may provide newer information during a request, but tool access does not change the model’s underlying training cutoff.

Reasoning and coding capabilities

GPT-6 Luna supports configurable reasoning effort. Developers can choose none, low, medium, high, xhigh, or max, with medium documented as the default. Reasoning effort is a practical control over how much processing the model applies before producing an answer. Lower settings can reduce latency and cost for straightforward work; higher settings are more appropriate for difficult analysis, complex code, or multi-step decisions.

The choice is a trade-off rather than a universal quality setting. A high reasoning level may be unnecessary for routine extraction or classification, while a low level may be unsuitable for complicated debugging or an agent task that must reconcile many constraints. Production applications can route simple requests to lower effort and reserve higher settings for cases that fail validation or require deeper analysis.

GPT-6 Luna is suitable for coding assistance, code explanation, repository analysis, test generation, structured code changes, and software-agent sub-tasks. OpenAI’s launch material reports that Luna improved over GPT-5.6 Luna on AutomationBench and achieved a 66.6% score on DeepSWE v1.1 at maximum reasoning effort. These are provider-reported evaluation results, not guarantees of performance on a particular codebase. Developers should test representative languages, frameworks, repository sizes, and failure modes before relying on the model for unattended changes.

Tools, function calling and API features

GPT-6 Luna supports streaming, function calling, structured outputs, prompt caching, and batch processing. Structured outputs allow an application to request a defined data structure instead of relying only on free-form prose. This is useful for extraction pipelines, classification systems, database preparation, and workflows that pass model results to another program.

OpenAI recommends the Responses API when using built-in tools or function calling with non-none reasoning effort. Function calling is also available through Chat Completions when reasoning effort is set to none. Function calling does not mean that the model independently performs arbitrary actions. Instead, the model proposes a tool call, and the surrounding application decides whether to execute it and what result to return.

The Responses API lists support for web search, file search, image generation, code interpreter, hosted shell, apply patch, skills, computer use, MCP, and tool search. These capabilities can make a Luna-based application more useful, but they are platform-level tools. They should not be recorded as direct image, audio, video, or executable output from the base model.

Fine-tuning is not supported for the exact GPT-6 Luna model. Structured outputs are supported. JSON mode is a separate API capability from structured outputs, and the supplied model documentation does not independently confirm JSON mode for GPT-6 Luna. Developers who specifically require JSON mode should verify the current endpoint documentation rather than assuming that structured-output support means the two features are identical.

GPT-6 Luna pricing

GPT-6 Luna’s standard short-context pricing is $0.10 per 1 million input tokens and $0.50 per 1 million output tokens. Cached input is priced at $0.01 per 1 million tokens, while cache writes cost $0.125 per 1 million tokens.

OpenAI lists separate long-context rates: $0.20 per 1 million input tokens, $0.02 per 1 million cached input tokens, $0.25 per 1 million cache writes, and $0.75 per 1 million output tokens. Requests with more than 272,000 input tokens receive this additional pricing treatment. A long-context request can therefore remain inexpensive relative to larger models, but developers should estimate costs using the applicable long-context rates when sending very large prompts.

Usage typeShort-context rateLong-context rate
Input$0.10 per 1M tokens$0.20 per 1M tokens
Cached input$0.01 per 1M tokens$0.02 per 1M tokens
Cache writes$0.125 per 1M tokens$0.25 per 1M tokens
Output$0.50 per 1M tokens$0.75 per 1M tokens

OpenAI states that Batch and Flex processing are priced at 50% of standard rates, while Fast mode costs twice the applicable rates. Regional processing may add a 10% premium where available. European Union data residency for GPT-6 Luna is limited to Standard processing. These options make the model’s economics dependent not only on token volume but also on latency requirements, processing mode, caching behavior, and request size.

Strengths and trade-offs

GPT-6 Luna’s main strength is the combination of low token pricing, a very large context window, configurable reasoning, and broad tool support. That combination is valuable when an application processes many similar requests or must repeatedly inspect large amounts of information. Prompt caching can further reduce the cost of repeated instructions or shared context.

  • Cost: Standard input and output prices are substantially lower than the listed rates for GPT-6 Sol and GPT-6 Astra.
  • Scale: Batch processing, caching, and low token rates support high request volumes.
  • Long context: The 1.05-million-token window can reduce the need to divide large documents or repositories into many separate calls.
  • Control: Reasoning effort settings let developers balance response quality, speed, and cost by task.
  • Integration: Function calling, structured outputs, streaming, web search, file search, code execution, and other Responses API tools support application workflows.

The trade-off is that Luna is not OpenAI’s highest-capability GPT-6 model. OpenAI recommends GPT-6 Astra when maximum capability is the priority and GPT-6 Sol for stronger reasoning on especially demanding tasks. Choosing Luna makes more sense when the workload is large, repeatable, or cost-sensitive and does not require the best available result on every difficult prompt.

Luna also has clear modality limits. It can inspect images but does not natively return images, audio, or video. It is not fine-tunable according to the supplied model documentation. Its knowledge cutoff means that current facts may require web search or another retrieval system, and all model-generated results should be checked when errors have meaningful consequences.

Best use cases for GPT-6 Luna

GPT-6 Luna is a strong fit for workloads where each individual request is important but the overall system must process many requests efficiently. Suitable examples include:

  • Classifying and routing large volumes of support tickets, documents, or business records.
  • Extracting fields, entities, decisions, or compliance information into structured outputs.
  • Summarizing long reports, contracts, technical documentation, or collections of related files.
  • Analyzing code repositories and generating explanations, tests, or proposed patches.
  • Supporting retrieval-augmented generation systems that combine retrieved source material with model reasoning.
  • Handling sub-tasks inside agents, such as planning a bounded step, calling a tool, or checking an intermediate result.
  • Processing repeated image-and-text requests, such as reviewing diagrams or document scans alongside instructions.

For example, a document-processing service could use image input for scanned pages, structured outputs for extracted fields, prompt caching for a repeated schema and policy definition, and Batch processing for non-urgent work. A coding system could use low or medium reasoning for routine code explanations and escalate difficult repository changes to a higher reasoning setting.

When to choose GPT-6 Luna

Choose GPT-6 Luna when you need a capable reasoning model for high-volume work and the cost of every request matters. It is particularly attractive when your application needs a long context, text-and-image input, structured results, tool calling, or batch processing, but does not require native media generation.

Consider a different option when maximum reasoning quality is more important than price or throughput. Within OpenAI’s GPT-6 lineup, GPT-6 Astra is the better comparison for frontier capability, while GPT-6 Sol is positioned for stronger reasoning on demanding tasks. A practical architecture may use Luna for routine requests and reserve a more capable sibling model for escalations, difficult edge cases, or quality-sensitive outputs.

GPT-6 Luna is also a poor fit when the required output is natively audio, video, or an image. Although the Responses API can expose tools such as image generation and computer use, those tools do not change Luna’s native text-output modality. Similarly, teams that need to customize model weights through fine-tuning should not select Luna based on an assumption that fine-tuning is available.

Availability and final assessment

GPT-6 Luna is available in the OpenAI API as gpt-6-luna. OpenAI announced access in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users, with access for Free and Go users through the desktop app. The launch announcement stated that the models were not initially available broadly in ChatGPT and would roll out gradually, so availability may depend on account type, product surface, and rollout status.

Overall, GPT-6 Luna is best understood as an efficiency-oriented reasoning model rather than a universal replacement for every GPT-6 option. Its unusually large context window, low standard token prices, adjustable reasoning effort, and extensive tool support make it well suited to production pipelines and repeatable agent workflows. Its limitations—text-only output, no documented fine-tuning, a fixed knowledge cutoff, and lower positioning than GPT-6 Sol and GPT-6 Astra—define when its lower cost is worth choosing over a higher-capability alternative.


Answers to Frequently Asked Questions

Where is GPT-6 Luna available?
GPT-6 Luna is available in the OpenAI API under the canonical model ID "gpt-6-luna". OpenAI also announced access through ChatGPT Work and Codex for eligible Plus, Pro, Business, Enterprise, and Edu users, with access for Free and Go users through the desktop app. Availability may vary by account type, product surface, and rollout status.
What can GPT-6 Luna do, and does it support image generation or fine-tuning?
GPT-6 Luna supports reasoning, coding, image understanding, streaming, function calling, structured outputs, prompt caching, batch processing, and Responses API tools such as web search and code interpreter. It does not natively generate images, audio, or video, and fine-tuning is not supported for the model.
What is GPT-6 Luna’s context window and maximum output?
GPT-6 Luna has a 1,050,000-token context window and supports a maximum output of 128,000 tokens. This makes it suitable for processing long documents, large code repositories, and extensive agent context.
What is GPT-6 Luna?
GPT-6 Luna is OpenAI’s efficiency-focused reasoning model for high-volume, repeatable tasks. It is available through the API under the model ID "gpt-6-luna" and supports text and image input with text output.
How much does GPT-6 Luna cost?
Standard pricing is $0.10 per 1 million input tokens and $0.50 per 1 million output tokens. Cached input costs $0.01 per 1 million tokens, while cache writes cost $0.125 per 1 million tokens. Requests with more than 272,000 input tokens use higher long-context rates: $0.20 per 1 million input tokens and $0.75 per 1 million output tokens.


Sources 5
Provider

About OpenAI