What is DALL·E 3?
DALL·E 3 was a specialized text-to-image model from OpenAI. It accepted a written prompt and generated a new image rather than returning a text completion. A prompt could describe a subject, setting, visual style, composition, objects, colors, or text that should appear in the image.
OpenAI introduced DALL·E 3 in September 2023 as a successor to DALL·E 2. Its main practical improvement was more reliable interpretation of detailed, conversational instructions. Instead of depending as heavily on short, highly specialized prompts, users could describe an intended scene in ordinary language and expect the model to preserve more of the requested relationships and details.
DALL·E 3 was also integrated with ChatGPT, where a conversational system could help expand or refine a user’s description before image generation. That made it useful to people who wanted to create images without learning a dedicated prompting syntax.
Where DALL·E 3 fit in OpenAI’s lineup
DALL·E 3 belonged to OpenAI’s image-generation product line rather than its general-purpose language-model family. It was not a reasoning model, coding model, conversational completion model, or multimodal assistant that analyzed arbitrary files. Its core task was prompt-conditioned image synthesis: turning text input into an image output.
During its supported period, users could access it through ChatGPT and through the OpenAI image-generation service with the model identifier dall-e-3. OpenAI deprecated the model on November 14, 2025 and removed it from the API on May 12, 2026. OpenAI recommends newer GPT Image models for current image-generation and editing workloads, including the newer image-model options covered elsewhere in this catalog.
Key capabilities
- Natural-language image generation: DALL·E 3 created images from written descriptions.
- Detailed prompt following: It was designed to better interpret long or nuanced prompts and preserve multiple requested elements.
- Several image formats: Supported API sizes included 1024×1024, 1792×1024, and 1024×1792, covering square, landscape, and portrait-oriented output.
- Improved visual detail: Compared with DALL·E 2, OpenAI emphasized improvements in detail, faces, hands, and other fine visual features.
- Text in images: DALL·E 3 improved the rendering of readable text within generated images, although generated lettering should still be checked rather than assumed to be exact.
- Safety controls: OpenAI used moderation and additional mitigations for harmful imagery, public-figure requests, and attempts to imitate the styles of living artists.
Input, output and technical profile
DALL·E 3 accepted text input and returned generated images. The supplied specifications identify text input and image output, but do not provide a conventional context-window size or maximum output-token limit. Those limits are not applicable in the same way they are for text-generation models.
| Capability | DALL·E 3 support |
|---|---|
| Text input | Yes |
| Image output | Yes |
| Text output | No |
| Audio or video input/output | No |
| Web search and tool calling | No |
| Structured JSON output | No |
| Streaming | No |
| Fine-tuning | No |
| Context length | Not published or not applicable as a text-model limit |
These restrictions matter when selecting a model. DALL·E 3 could create an illustration from a description, but it could not independently browse for facts, execute code, return a structured data object, generate audio or video, or serve as a general-purpose language assistant. A surrounding product such as ChatGPT could provide a conversational interface, but those additional capabilities should not be attributed to DALL·E 3 itself.
Pricing and API history
DALL·E 3 API pricing was based on generated images rather than input and output tokens. Historical pricing started at $0.04 per 1024×1024 standard-quality image. HD quality and the larger landscape or portrait formats cost more. The supplied research does not establish a current price because the model is no longer available through the API.
That pricing structure made DALL·E 3 different from text models whose bills are calculated from token counts. Each image request represented a discrete generation cost, so image dimensions and quality settings were important when estimating usage. Because the API has been shut down, the historical prices should not be treated as an available purchasing option.
Strengths and limitations
What DALL·E 3 did well
DALL·E 3’s strongest feature was the connection between ordinary language and image composition. It was a good fit for prompts that required several visible details to coexist, such as a particular subject, environment, arrangement, mood, and piece of lettering. Its ChatGPT integration also lowered the barrier for users who preferred to describe an idea conversationally and revise it through follow-up instructions.
Compared with earlier DALL·E systems, it offered improved prompt adherence, visual detail, and text rendering. The availability of square, landscape, and portrait formats made it suitable for different creative and marketing layouts without restricting every request to a square canvas.
What DALL·E 3 could not do
DALL·E 3 was not a general-purpose AI model. It did not produce text answers, embeddings, audio, video, code, web-search results, or structured JSON as native outputs. It also was not designed for fine-tuning or token-based context-window workflows. The model’s function was image creation, so applications requiring image editing, document analysis, reasoning, or reliable structured data should use a different system.
Safety controls could also affect results. OpenAI applied safeguards to harmful-image requests, certain public-figure prompts, and requests that attempted to imitate the styles of living artists. A prompt could therefore be declined or altered even when the user considered the intended image legitimate.
Generated text and visual details were improvements rather than guarantees. Users creating posters, labels, diagrams, or other text-heavy images needed to inspect the result and be prepared for inaccuracies. The supplied research does not provide benchmark scores or a formal accuracy guarantee.
Reasoning, coding and speed trade-offs
DALL·E 3 had no native reasoning or coding capability in the sense used for language models. It could interpret a prompt describing a visual problem, but it did not reason through a task with a text response or write and execute code. Editorial capability scores classify its reasoning and coding performance as minimal because it was a specialized image model, not because OpenAI presented those scores as official benchmarks.
Its speed and cost profile should likewise be understood in relation to its purpose. Image generation involved a separate per-image charge and a generation wait, while a text model may be more appropriate for rapid drafting, classification, extraction, or code generation. For image work, the choice between standard and HD output, and between smaller and larger formats, affected historical cost. The available research does not establish a precise response-time guarantee.
When to choose this model
DALL·E 3 made sense historically when the main requirement was producing a new image from a detailed natural-language description. Suitable uses included:
- Concept art and early visual ideation.
- Illustrations for creative projects.
- Marketing or campaign-image exploration.
- Scene visualization and mood-board development.
- Prompt-following research.
- Images that required a particular aspect ratio or some readable in-image text.
For a new deployment today, however, DALL·E 3 is not the practical choice because its API has been removed. A current GPT Image model is more appropriate when the project needs supported image generation or editing. A general-purpose language model is a better fit for text, coding, reasoning, web research, or structured output. A dedicated audio or video model is needed for those media types.
Why DALL·E 3 remains important
DALL·E 3 helped make conversational image prompting more accessible. Its integration with ChatGPT showed how a user could describe an idea in ordinary language, refine it through dialogue, and then generate an image without manually tuning a specialized prompt format.
The model also marked a notable step in commercially available text-to-image systems because it combined improved instruction following with multiple output formats and stronger safety measures. Although it is no longer available through the OpenAI API, DALL·E 3 remains a useful reference point for understanding the development of OpenAI’s image-generation systems and the transition toward newer GPT Image models.

