omni-moderation

omni-moderation

by OpenAI · Current

OpenAI’s omni-moderation model classifies potentially harmful text and images across categories including harassment, hate, illicit content, self-harm, sexual content, and violence. It returns flags and category scores through the free Moderation API, with support for a latest alias and a dated snapshot.

Text Reasoning Coding
Omni moderation is OpenAI’s current multimodal moderation model for detecting potentially harmful content. Available through the Moderation API, it accepts text and images, produces category-specific safety assessments, and can help applications filter submissions, screen AI-generated content, and route uncertain cases to human reviewers. It is free to use through the API, although usage-tier limits still apply.
Outputs

What omni-moderation can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
8/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family omni-moderation
Model type Moderation
Release date 2024-09-26
Status Current
Knowledge cutoff notes

OpenAI's current model documentation and model information PDF do not provide a separate knowledge-cutoff date for omni-moderation.

Model notes

The current alias is omni-moderation-latest and the documented fixed snapshot is omni-moderation-2024-09-26. The model returns moderation flags, per-category scores, and applicable input types rather than ordinary generative responses. It supports text and images but not audio or video. Image classification is available only for documented categories, while several categories remain text-only. The Moderation API is free, subject to usage-tier rate limits. OpenAI introduced the model on September 26, 2024 and describes it as being based on GPT-4o. The model page lists streaming, function calling, structured outputs, and fine-tuning as unsupported.

Cost

Model pricing

Input Free
Output Free
Model guide

Omni Moderation: OpenAI’s Multimodal Safety Model and API Guide

OpenAI’s omni-moderation is a free moderation model that classifies potentially harmful text and images across categories such as harassment, hate, illicit content, self-harm, sexual content, and violence. It returns moderation flags, category-level scores, and applicable input types rather than ordinary generated text.

What is Omni Moderation?

Omni moderation is OpenAI’s dedicated content-safety model. Its purpose is not to write text, answer questions, generate images, or perform general reasoning. Instead, it examines submitted content and estimates whether it falls into one or more potentially harmful categories.

The model is available through OpenAI’s Moderation API under the current alias omni-moderation-latest. OpenAI also documents the fixed snapshot omni-moderation-2024-09-26, which is useful when an application needs a stable model identifier rather than an alias that may be updated over time. OpenAI introduced the model on September 26, 2024 and describes it as being based on GPT-4o.

In OpenAI’s current catalog, omni moderation occupies a specialized safety-classification role. It complements generative models rather than replacing them: an application can use a generative model to produce a response and then submit that response to omni moderation before displaying it or taking an automated action.

What the model returns

Omni moderation produces a structured moderation result rather than a conversational answer. The response includes an overall flagged prediction, individual category flags, calibrated scores for those categories, and, where applicable, information about the input type associated with a result.

A flag is useful for a first-pass policy decision, while the category scores provide more detail for applications that need different thresholds. For example, a platform could automatically reject content above a high-confidence threshold, hold borderline cases for human review, and allow lower-risk content. Those thresholds are application decisions; OpenAI’s documentation does not prescribe one universal value for every product or policy.

The model should therefore be viewed as a safety signal, not as a complete content-governance system. A production application still needs rules for appeals, escalation, repeat offenses, regional requirements, and actions taken after a category is detected.

Supported inputs and outputs

CapabilitySupport
Text inputSupported
Image inputSupported, including image URLs and base64 data URLs
Audio inputNot supported
Video inputNot supported
Textual moderation outputSupported
Image, audio, or video outputNot supported

Image files submitted for moderation can be up to 20 MB according to the supplied API research. Image analysis is not equivalent to unrestricted visual understanding: only documented moderation categories are available for image inputs. Several other categories remain text-only.

Omni moderation does not generate replacement text, captions, images, audio, or video. Its output is a classification record containing safety-related fields. The model also does not provide ordinary tool use, function calling, streaming responses, or structured-output generation in the sense used for generative models.

Which safety categories does it cover?

The documented categories cover a broad range of harmful content:

  • Harassment and threatening harassment
  • Hate and threatening hate
  • Illicit instructions and violent illicit instructions
  • Self-harm, self-harm intent, and self-harm instructions
  • Sexual content
  • Sexual content involving minors
  • Violence and graphic violence

Image moderation is category-specific. Images can be evaluated for violence, graphic violence, self-harm, self-harm intent, self-harm instructions, and sexual content. Categories including harassment, hate, illicit content, and sexual content involving minors are documented as text-only categories in the supplied model information.

This distinction matters when designing a pipeline. An application should not assume that every text category can be applied to an image, or that an image result covers every possible safety concern. Text and image checks may need to be handled as separate policy paths.

Pricing and API access

OpenAI makes omni moderation available at no charge through the Moderation API. “Free” refers to the endpoint’s pricing, not unlimited capacity: requests remain subject to rate limits determined by the account’s usage tier. The research also identifies batch API support, but it does not provide a separate paid or discounted moderation price.

The model can be used to inspect user-submitted material before publication, check generated text or images, or evaluate content produced elsewhere. It can also be placed after a conversational or generative model as a review step. Because the moderation result is separate from generation, applications can choose whether to block, redact, delay, label, or escalate content.

Main strengths and limitations

Strengths

  • Text-and-image coverage: It supports two important content types in a single moderation model, rather than limiting safety checks to text.
  • Broad harm taxonomy: The categories distinguish related risks such as general violence versus graphic violence, and self-harm content versus self-harm intent or instructions.
  • Detailed signals: Per-category flags and scores offer more control than a single yes-or-no moderation result.
  • Low direct cost: The Moderation API is free, making automated first-pass screening practical for many applications.
  • Useful for routing: Scores can help direct uncertain or high-risk cases to human review instead of forcing every decision to be fully automated.

Limitations

  • No audio or video moderation: Applications handling spoken audio or video must use another approach, such as extracting text or frames and evaluating those separately, although the supplied research does not define a complete alternative pipeline.
  • Uneven image coverage: Not every documented category applies to images.
  • No general generation: Omni moderation is not a conversational, coding, reasoning, image-generation, or speech model.
  • No documented context or output limits: The supplied research does not provide a context-window size or maximum output-token figure. Its response is a moderation result rather than a long-form generated answer.
  • Policy decisions remain application-specific: A score does not by itself determine whether content should be removed, reported, or allowed.
  • Special child-safety constraint: OpenAI states that the Moderation API is not designed to handle known or suspected child sexual abuse material and should not replace dedicated child-safety safeguards.

Reasoning, coding, speed, and cost positioning

Omni moderation is optimized for classification rather than open-ended reasoning or code generation. It should not be selected when the central task is solving a complex problem, writing software, or producing a nuanced conversational explanation. A general-purpose language model may be more appropriate for those tasks, with omni moderation added as a safety check before or after generation.

The supplied evaluation fields rate its speed and cost favorably relative to more capable generative systems, but those are editorial scores rather than provider-published benchmark results. The verified cost distinction is simpler: the Moderation API endpoint is free, while the model is specialized and returns a compact moderation result. That combination makes it a practical screening layer even when the application already uses a more expensive model for generation.

Its trade-off is breadth versus specialization. A multimodal generative model may understand a wider range of requests or produce explanations, but it may be more expensive and is not a substitute for a dedicated moderation policy layer. Conversely, omni moderation is inexpensive and focused, but it cannot independently manage a complete audio-video safety workflow or explain a policy decision in ordinary prose.

When to choose Omni Moderation

Choose omni moderation when the primary requirement is automated safety classification for text, images, or both. It is a strong fit for:

  • Pre-moderating user posts, comments, messages, and image uploads
  • Screening text or images generated by another AI system
  • Routing borderline submissions to human reviewers
  • Applying separate thresholds to different harm categories
  • Adding a low-direct-cost safety check to an API-based application
  • Building a moderation layer that needs both text and documented image-category support

It is less suitable when the input is primarily audio or video, when the application needs a generative explanation, or when the core task is general reasoning, coding, transcription, translation, or content creation. It is also not sufficient by itself for high-risk child-safety operations or for designing a complete platform enforcement policy.

Practical integration guidance

A sensible integration pattern is to submit content before publication, inspect the overall flag and relevant category results, and then apply application-defined actions. Low-risk content might continue automatically; high-confidence violations might be blocked; uncertain results can be queued for human review. For generated content, the moderation step can run before the result reaches the user.

Applications should log the model identifier used, including whether they selected omni-moderation-latest or the dated snapshot. They should also preserve enough category information to explain an action internally, while applying appropriate privacy and retention controls to submitted content.

Because image categories differ from text categories, the user interface and policy engine should not present an image result as if every text risk had been evaluated. Similarly, a moderation score should support a broader safety process rather than serve as the only control for sensitive use cases.

Bottom line

Omni moderation is a focused OpenAI safety model for classifying harmful text and images. Its strongest practical advantages are multimodal input support, detailed category-level results, a broad documented harm taxonomy, and free access through the Moderation API. Its boundaries are equally important: it does not moderate audio or video directly, does not generate ordinary content, has category-specific image coverage, and does not replace human review or dedicated safeguards for high-risk material.


Answers to Frequently Asked Questions

What is the difference between omni-moderation-latest and omni-moderation-2024-09-26?
The alias omni-moderation-latest may be updated by OpenAI over time, while omni-moderation-2024-09-26 is a fixed snapshot. Applications that require a stable model identifier can use the dated snapshot.
Can Omni Moderation replace a complete content safety system?
No. Omni Moderation provides safety classifications and scores, but applications still need their own enforcement rules, thresholds, appeals process, escalation paths, privacy controls, and human review. It is also not designed to replace dedicated safeguards for known or suspected child sexual abuse material.
Is OpenAI Omni Moderation free to use?
OpenAI provides Omni Moderation through the Moderation API at no charge. However, requests are still subject to rate limits based on the account’s usage tier, so free access does not mean unlimited capacity.
What is OpenAI Omni Moderation used for?
Omni Moderation is a dedicated OpenAI safety model that classifies potentially harmful text and images. It returns overall and category-level safety signals that applications can use to allow, block, label, delay, or send content for human review.
What types of content does Omni Moderation support?
Omni Moderation supports text and image inputs, including image URLs and base64 data URLs. It does not directly support audio or video moderation. Image checks cover documented categories such as violence, graphic violence, self-harm, and sexual content, while some categories are text-only.


Sources 5
Provider

About OpenAI