omni-moderation

omni-moderation-latest

by OpenAI · Current default moderation model

OpenAI's omni-moderation-latest is a free specialized moderation model for text and images. It returns overall and category-level safety flags plus numerical scores for harms including harassment, hate, sexual content, violence, self-harm, and illicit activity. It is designed for automated filtering and review workflows, but does not directly support audio or video and should be combined with application-specific policies and human review.

Text
omni-moderation-latest is OpenAI's current default moderation model for classifying potentially harmful text and images. It is intended for application safety workflows such as filtering user content, checking generated responses, prioritizing human review, and enforcing product-specific policies. Unlike a general-purpose language model, it does not generate open-ended answers: its role is to return structured moderation classifications and scores.
Outputs

What omni-moderation-latest can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Batch API
Model profile

Performance characteristics

9/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family omni-moderation
Model type Other
Release date 2024-09-26
Status Current default moderation model
Knowledge cutoff notes

OpenAI's current model documentation does not publish a knowledge cutoff for omni-moderation-latest. The model is a classification system whose results are based on the submitted content and its trained moderation classifiers.

Model notes

OpenAI's current default moderation model and successor to the text-only moderation models. It accepts text and images but does not classify audio or video. The API returns an overall flagged decision, category flags, and category scores. Image moderation is available only for selected categories; some categories remain text-only. Image inputs can be up to 20 MB. OpenAI states that the Moderation API is free to use. The model page lists streaming, function calling, structured outputs, and fine-tuning as unsupported.

Cost

Model pricing

Input Free through the Moderation API
Output Free through the Moderation API
Model guide

OpenAI omni-moderation-latest: Multimodal Moderation Model, Features and API Details

OpenAI's omni-moderation-latest is a free moderation model for detecting potentially harmful content in text and images. It returns an overall flagged decision, category-level Boolean results, and numerical scores for harms such as harassment, hate, sexual content, violence, self-harm, and illicit activity.

What is omni-moderation-latest?

omni-moderation-latest is OpenAI's current general-purpose moderation model. It examines submitted content and estimates whether it falls into potentially harmful categories. The model accepts text, images, or a combination of text and image content, then returns a structured moderation result rather than a conversational response.

In practical terms, an application can send a user's post, a message generated by another model, or an image uploaded by a user for screening. The response includes an overall flagged decision, category-specific Boolean flags, and numerical category scores. A product can use the overall decision for a simple first-pass filter, or inspect individual categories and scores when it needs more precise routing.

The model is available through OpenAI's Moderation API and is free to use according to the supplied OpenAI documentation. OpenAI presents it as the successor to earlier text-only moderation models. The related Omni Moderation page provides broader background on this model family, while this page focuses on the current omni-moderation-latest item.

Where it fits in OpenAI's lineup

omni-moderation-latest occupies a specialized safety-classification role in OpenAI's catalog. It is not a general-purpose reasoning model, a text-generation model, an embedding model, or a media-generation system. Its purpose is to help applications make safety decisions about content produced by people, models, or connected tools.

The word “omni” refers to its broader input coverage compared with OpenAI's earlier text-only moderation systems: it can evaluate images as well as text. This does not mean that every moderation category works equally across both modalities. Some categories remain text-only, and image moderation is available for selected violence, self-harm, and sexual-content categories.

It is therefore best understood as an automated classification component inside a larger safety pipeline. It can support filtering and triage, but it does not define a complete content policy for a product. The application owner still needs to decide what happens when content is flagged, which thresholds apply, whether borderline cases go to human reviewers, and how regional or legal requirements are handled.

Supported inputs and outputs

The verified input modalities are text and images. OpenAI's model documentation states that image files sent to the Moderation API can be up to 20 MB. Audio and video are not supported as direct inputs, so an application that needs to moderate those formats would need an additional processing step or a different specialized system. The supplied research does not establish that omni-moderation-latest can interpret audio tracks, video frames, or video streams directly.

CapabilityStatusPractical meaning
Text inputSupportedClassifies text for potentially harmful content.
Image inputSupportedClassifies supported visual safety categories; image files can be up to 20 MB.
Combined text and image inputSupportedAllows applications to assess multimodal content.
Audio inputNot supportedAudio must be handled by another system or converted into a supported representation.
Video inputNot supportedThe model does not directly classify video.
Generated prose outputNot supportedThe response is a structured moderation result, not an open-ended explanation.

The model's output is text-based structured classification data. It does not generate images, audio, or video. The research does not publish a context-window size or maximum output-token limit for this model, so those specifications should not be assumed from the limits of OpenAI's other models.

What the moderation result contains

The result has three useful layers. First, an overall flagged value gives an application a simple indication that content may require action. Second, category-level Boolean fields identify which types of harm may be present. Third, numerical category scores provide probability-like signals that can be used for prioritization or custom thresholds.

Relevant categories include harassment, hate, sexual content, violence, self-harm, and illicit activity. The exact category coverage is not identical for text and images. For example, some hate, harassment, and illicit-activity classifications may require textual input, while image moderation is available for selected violence, self-harm, and sexual-content categories.

A score should not be treated as a universal legal or product decision. A marketplace might block content above one threshold, send borderline material to review, and allow lower-scoring content. A children's product, a private workplace tool, and a public social platform may reasonably choose different enforcement policies. omni-moderation-latest supplies signals for those decisions; it does not automatically know the application's rules or risk tolerance.

Main strengths

  • Text and image coverage: It can screen two important content types through the same moderation model rather than limiting a pipeline to text.
  • No model-use charge: OpenAI documents the Moderation API as free to use, which makes the model suitable for high-volume first-pass screening where infrastructure and review costs still need to be considered separately.
  • Structured results: Overall flags, category flags, and scores support both simple filters and more detailed review workflows.
  • Broad safety categories: The model covers common abuse categories and includes newer illicit-activity categories described in OpenAI's moderation materials.
  • Application workflow support: It can be used for user-generated content, model-generated responses, and moderation of relevant tool-call arguments or tool outputs when they appear as conversation content.
  • Multilingual improvement claim: OpenAI describes the model as offering improved multilingual moderation compared with its previous text-only model. This is a provider claim rather than an independent benchmark conclusion in the supplied research.

Limitations and trade-offs

The model's largest limitation is scope. It is a classifier, not an all-purpose safety analyst or a replacement for human judgment. It cannot directly moderate audio or video, and not every category is available for image input. An application dealing with video may need to extract frames and separately process speech or transcripts, but the supplied research does not specify how such a system should be built or what accuracy it would achieve.

Moderation scores can also produce false positives and false negatives. A score is evidence for a decision, not proof that content violates a particular platform rule. Slang, cultural context, quotation, satire, educational discussion, and ambiguous imagery can all require policy-aware handling beyond a single automated result.

OpenAI lists streaming, function calling, structured outputs, and fine-tuning as unsupported model features. The Moderation API itself returns structured objects, but that should not be confused with a general-purpose structured-output mode for generated text. The model also has no published reasoning or coding capability in the usual sense: it does not solve programming problems or provide visible step-by-step analysis. Its job is to classify submitted content.

The supplied documentation does not publish a context length, maximum output-token limit, or a conventional per-token price. The relevant pricing fact is that moderation requests are free through the Moderation API. Developers should still account for application infrastructure, storage, logging, reviewer time, and any other services used around the model.

How it compares with other option types

Compared with a general-purpose language model, omni-moderation-latest is narrower but more directly suited to safety classification. A general model may be able to explain why content appears risky or help draft a user-facing notice, but using it as the primary moderation decision-maker can add cost, latency, and inconsistency. omni-moderation-latest instead returns a purpose-built classification response.

Compared with a text-only moderation model, omni-moderation-latest is the more appropriate option when images are part of the product's content surface. Its image coverage remains category-specific, so a text-only workflow may still be sufficient for a system that never receives images.

Compared with a human-review process, the model is faster and more scalable for first-pass screening, but it cannot provide the contextual judgment or accountability that difficult cases may require. The practical trade-off is usually automated classification for routine or clearly risky content, combined with escalation for ambiguous or high-impact decisions.

Because the Moderation API is free, cost is less about model-token pricing than about system design. The model can reduce the volume of content sent to reviewers, but a high-volume service still needs to budget for API integration, image handling, storage, monitoring, appeals, and human review. The supplied research does not provide a latency guarantee, so speed should be measured in the target application rather than inferred from the model's cost.

Best use cases

omni-moderation-latest is suited to applications that need automated screening before or after content becomes visible. Common uses include:

  • Checking user posts, comments, messages, marketplace listings, or profile content.
  • Screening images uploaded to community, commerce, productivity, or social products.
  • Running a safety check on responses generated by another AI model before delivery.
  • Prioritizing items for a human moderation queue using category flags and scores.
  • Applying different actions to different categories, such as blocking, warning, delaying publication, or requesting review.
  • Auditing content and identifying recurring safety patterns for policy and trust-and-safety teams.
  • Checking relevant tool-call arguments or tool outputs when they are represented as conversation content.

A basic workflow might first inspect the overall flagged value, then use category flags to choose an action, and finally retain scores and the original content according to the application's privacy and retention policy. High-impact actions should generally include a review or appeal path rather than relying on an opaque automatic block alone.

When to choose omni-moderation-latest

Choose this model when the main requirement is free, automated moderation of text or supported image content through OpenAI's Moderation API. It is especially attractive when a product needs category-level results, wants to screen both user and generated content, or needs a scalable first layer before human review.

Choose another option or add another system when the content is audio or video, when all required categories must work on images, or when the application needs an explanation, policy-specific reasoning, fine-tuning, streaming, or function calling. A general-purpose model may be useful alongside the moderation classifier for drafting explanations or handling workflow automation, but that does not make it a substitute for the model's dedicated safety signals.

For straightforward text-only screening, an earlier or narrower text moderation option may be adequate if it meets the application's category and policy requirements. For mixed text-and-image products, omni-moderation-latest is the more natural fit among the options documented here, provided its category coverage and operational behavior are tested against representative content.

Bottom line

omni-moderation-latest is a specialized, free OpenAI model for classifying potentially harmful text and images. Its useful distinction is not open-ended intelligence but structured safety output: an overall flag, category flags, and numerical scores that applications can turn into filters, queues, audits, and policy actions. It is a strong first-pass component for multimodal content safety, but its image-category restrictions, lack of audio and video support, unsupported advanced model features, and need for application-specific thresholds mean it should be deployed as part of a broader safety process rather than treated as a complete moderation solution.


Answers to Frequently Asked Questions

What are the main limitations of omni-moderation-latest?
The model is a classifier rather than a complete safety solution. It cannot directly moderate audio or video, not every moderation category applies to images, and its scores can produce false positives or false negatives. Streaming, function calling, structured outputs for generated text, and fine-tuning are also unsupported.
Is omni-moderation-latest free to use?
OpenAI documents the Moderation API as free to use. However, applications may still incur costs for infrastructure, image processing, storage, monitoring, human review, and other services connected to the moderation workflow.
What moderation categories and outputs does the model provide?
omni-moderation-latest provides signals for categories such as harassment, hate, sexual content, violence, self-harm, and illicit activity. Its response includes an overall flagged value, category-level flags, and numerical scores, although category coverage differs between text and images.
What is omni-moderation-latest?
omni-moderation-latest is OpenAI's general-purpose moderation model for classifying potentially harmful text and images. It returns a structured result with an overall flagged decision, category-specific Boolean flags, and numerical category scores.
What types of content does omni-moderation-latest support?
The model supports text, images, and combined text-and-image inputs. Image files sent through the Moderation API can be up to 20 MB. Audio and video are not supported as direct inputs.


Sources 4
Provider

About OpenAI