Sora 2

Sora 2

by OpenAI · Deprecated; currently accessible through the API as of September 23, 2026, with shutdown scheduled for September 24, 2026

OpenAI's Sora 2 generates short videos from text prompts or image references with synchronized dialogue, sound effects, music, and ambient audio. It focuses on physical realism and rapid iteration, but is deprecated and scheduled for API shutdown on September 24, 2026.

Video generation Speech Reasoning Coding
Sora 2 is OpenAI's video-and-audio generation model for creating short clips from natural-language descriptions and image references. It is designed for concept development, social video, storyboards, prototypes, and other short-form creative work where synchronized audio and quick iteration are useful. OpenAI describes Sora 2 as the faster, more flexible member of the Sora 2 generation, while Sora 2 Pro targets higher-fidelity production work. However, Sora 2 is deprecated, with API access scheduled to end on September 24, 2026.
Outputs

What Sora 2 can produce

Video generation Speech
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Batch API Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family Sora 2
Model type Multimodal
Release date 2025-09-30
Status Deprecated; currently accessible through the API as of September 23, 2026, with shutdown scheduled for September 24, 2026
Deprecation date 2026-03-24
Shutdown date 2026-09-24
Knowledge cutoff notes

OpenAI does not publish a directly verifiable knowledge-cutoff date for Sora 2. As a generative video model, its documented behavior is based on trained visual and audiovisual data rather than a conventional conversational knowledge cutoff.

Model notes

Sora 2 is OpenAI's faster and more flexible second-generation video model. It accepts natural-language prompts and image references and outputs video with synchronized audio, including dialogue and sound effects. The model page lists the canonical alias sora-2 and dated snapshots sora-2-2025-10-06 and sora-2-2025-12-08. OpenAI documentation states that the Sora 2 models and Videos API are deprecated and scheduled to shut down on September 24, 2026. The consumer Sora product was discontinued on April 26, 2026, separately from the API shutdown. The model is not a general-purpose text model, so token context and maximum output-token fields are not applicable.

Cost

Model pricing

Input Not token-priced; image and text inputs are included in video-generation requests
Output $0.10 per generated video second for 720x1280 portrait or 1280x720 landscape output
Model guide

Sora 2: OpenAI Video Generation, Pricing, Capabilities and API Status

Sora 2 is OpenAI's second-generation video generation model for creating short clips from text prompts or image references. It produces video with synchronized dialogue, sound effects, music, and environmental audio, with an emphasis on physical realism, scene continuity, and rapid creative iteration. The model is deprecated, and OpenAI has scheduled the Sora 2 models and Videos API for shutdown on September 24, 2026.

What is Sora 2?

Sora 2 is OpenAI's second-generation video generation model. It creates short video clips from natural-language prompts or image references and generates synchronized audio alongside the visuals. Depending on the scene, that audio can include dialogue, sound effects, music, and ambient soundscapes.

Unlike OpenAI's text-generation models, Sora 2 is primarily a media-generation system. Text and images guide the request, while the main outputs are video and audio. This makes it suitable for producing visual concepts and short audiovisual sequences rather than answering questions, writing code, or generating structured text.

OpenAI positions Sora 2 as the faster and more flexible model in the Sora 2 generation. Sora 2 Pro is intended for higher-fidelity production work, so the standard Sora 2 model is more appropriate when iteration speed and lower generation cost matter more than maximum fidelity.

What can Sora 2 generate?

Sora 2 is designed to turn descriptions such as a scene, action, camera direction, or visual style into a short video. It can also use an image as a visual reference, allowing an existing still image to guide the appearance or composition of the resulting clip.

  • Text-to-video generation from natural-language prompts
  • Image-guided video generation
  • Video with synchronized dialogue, sound effects, music, and environmental audio
  • Multi-shot sequences based on more detailed instructions
  • Video extension and targeted editing through the Videos API
  • Reusable character assets for more consistent subjects
  • Portrait and landscape output at 720p-oriented dimensions

The model is intended for short-form generation and rapid iteration. A creator can explore several versions of a scene, adjust the prompt or reference image, and use the results as concept material, a rough cut, storyboard content, or a prototype rather than treating every generated clip as a finished production asset.

Physical realism and continuity

OpenAI highlights improvements in motion realism, physical behavior, and scene continuity compared with earlier video-generation systems. The documented target areas include difficult actions such as gymnastics, impacts, interactions between objects, and movement involving water or rigid surfaces.

Sora 2 also supports multi-shot instruction following. In practical terms, a prompt can describe more than one shot or a sequence of related visual events, rather than only a single isolated frame or action. Reusable character assets are intended to help maintain more consistent subjects across generations.

These capabilities should be understood as model goals and provider-described strengths, not a guarantee that every generated clip will be physically correct. Video generation can still produce visual, temporal, physical, or audio inconsistencies. Important footage therefore requires human review, editing, and potentially multiple regeneration attempts.

Inputs and outputs

Sora 2 accepts text and image inputs. The supplied model data does not identify audio or video input as supported input modalities, so it should not be treated as a general-purpose video-to-video or audio-conditioned model.

Its direct outputs are video and audio. The audio may be synchronized with the generated action and can include spoken dialogue, sound effects, music, and environmental sounds. Text output, image output, embeddings, and structured data are not the purpose of this model.

CategorySora 2 support
Text inputYes
Image inputYes
Audio inputNo, according to the supplied model data
Video inputNo, according to the supplied model data
Video outputYes
Audio outputYes, synchronized with generated video
Text outputNo

The documented output dimensions are 720x1280 for portrait video and 1280x720 for landscape video. The supplied research does not specify a maximum clip duration, context window, maximum output-token limit, frame rate, or file-size limit. Context length and maximum output tokens are not applicable in the same way they are for conversational language models.

API access and pricing

Sora 2 is available through OpenAI's Videos API for creating, extending, editing, downloading, and managing generated videos. Its pricing is based on generated video duration rather than language-model input and output tokens.

OpenAI lists Sora 2 at $0.10 per generated video second for portrait output at 720x1280 and landscape output at 1280x720. For example, a 10-second generation would cost $1.00 at that listed rate before any applicable account-specific considerations. The research does not provide a separate input price because text and image inputs are included in video-generation requests.

The model's cost and speed profile makes it more suitable for exploratory workflows than a model selected solely for maximum production fidelity. However, price should be considered alongside the model's scheduled retirement: a new integration built around Sora 2 may have only a limited operating window.

Reasoning, coding, and tool support

Sora 2 is not a general-purpose reasoning or coding model. Its reasoning and coding scores in the supplied model data are editorial evaluation fields, not OpenAI-published benchmark results. They should not be interpreted as evidence that Sora 2 can replace a text model for analysis, programming, or planning.

The model does not provide general tool or function-calling support. The Videos API does provide operations specific to video workflows, including generation, extension, editing, downloading, and management. Those API operations should not be confused with broad agent tools such as web search, code execution, or external application actions.

Sora 2 also does not support structured output or JSON generation as a primary model capability. If an application needs reliable structured data, text completion, software development, web research, or embeddings, a dedicated language, reasoning, coding, search, or embedding model is more appropriate.

Deprecation and important limitations

OpenAI's supplied documentation marks Sora 2 as legacy and deprecated. The Sora 2 models and Videos API are scheduled to shut down on September 24, 2026. This date concerns API access to the model and should be distinguished from the separate discontinuation of the consumer Sora product on April 26, 2026.

The planned shutdown is the most important practical limitation for developers. Existing workflows may be able to continue until the stated date, but Sora 2 is a poor foundation for a new long-lived production integration unless the project can be completed or migrated before access ends.

Other limitations include the absence of a documented conventional context window, maximum output-token value, and general-purpose tool support. The model can also produce inconsistent motion, object interactions, timing, character appearance, or audio synchronization. Human review is especially important for scenes involving precise physical actions, continuity across shots, branded assets, or dialogue that must match the intended script.

Best use cases for Sora 2

  • Rapid video concepting: Explore visual directions before committing to a full production.
  • Social media clips: Create short portrait or landscape audiovisual content for experimentation and publishing workflows.
  • Storyboards and pitch materials: Turn written concepts into moving visual references.
  • Image-to-video experiments: Animate or develop an existing image into a short sequence.
  • Prototype scenes: Test cinematic ideas, camera movement, environments, or character concepts.
  • Rough cuts and creative iteration: Generate multiple alternatives where speed is more valuable than maximum fidelity.

These use cases benefit from short generation cycles and native synchronized audio. They are less dependent on exact factual correctness than applications such as education, legal work, software development, or automated business reporting.

When to choose Sora 2

Choose Sora 2 when you need short text- or image-guided video with accompanying audio, and when fast creative exploration is more important than the highest available fidelity. It is particularly reasonable for temporary prototypes, preproduction work, social concepts, and experiments that can be completed before the scheduled API shutdown.

Choose a higher-fidelity video option, including OpenAI's Sora 2 Pro where available and appropriate, when production quality is more important than speed or cost. Choose a conventional language or reasoning model for text, planning, analysis, coding, structured responses, or tool-driven workflows. A dedicated image model is a better fit when the required output is a still image rather than a moving audiovisual clip.

The shutdown schedule should influence the decision as much as output quality. If the project requires a stable video-generation API beyond September 24, 2026, another currently supported provider or model is a safer choice. Sora 2 is best treated as a capable but time-limited option for short-form audiovisual generation.

Bottom line

Sora 2 combines text- and image-guided video generation with synchronized dialogue, sound effects, music, and ambient audio. Its strongest documented advantages are short-form audiovisual creation, improved physical behavior, multi-shot instruction following, and rapid iteration at a listed price of $0.10 per generated second. It is not a conversational, coding, reasoning, image, or general tool-use model. Because OpenAI has deprecated it and scheduled the Sora 2 models and Videos API for shutdown on September 24, 2026, it is mainly suitable for existing workflows, short-lived integrations, and creative projects that can be completed before that date.


Answers to Frequently Asked Questions

What are the best use cases and limitations of Sora 2?
Sora 2 is best suited to rapid video concepting, social media clips, storyboards, pitch materials, image-to-video experiments, prototype scenes, and rough cuts. It can still produce inconsistent motion, object interactions, continuity, timing, or audio synchronization, so important footage requires human review and editing.
When will the Sora 2 API be discontinued?
OpenAI has marked Sora 2 and its Videos API as legacy and deprecated, with shutdown scheduled for September 24, 2026. New long-term integrations should consider a currently supported alternative if they need access beyond that date.
Does Sora 2 support image-to-video generation and video editing?
Yes. Sora 2 accepts image references for image-guided video generation. Through the Videos API, it also supports video extension, targeted editing, downloading, and management. The supplied model data does not identify audio or video files as supported input modalities.
What is Sora 2 and what can it generate?
Sora 2 is OpenAI’s second-generation video generation model. It creates short video clips from text prompts or image references and can generate synchronized dialogue, sound effects, music, and ambient audio.
How much does the Sora 2 API cost?
OpenAI lists Sora 2 at $0.10 per generated video second for 720x1280 portrait or 1280x720 landscape output. A 10-second video would therefore cost $1.00 at the listed rate, before any account-specific considerations.


Sources 5
Provider

About OpenAI