CLIP

CLIP

by OpenAI · Public research release with downloadable weights; not verified as a current OpenAI hosted API model

OpenAI CLIP is a public research model that maps images and text into a shared embedding space. It supports zero-shot classification, image-text similarity, semantic retrieval, ranking, and multimodal indexing, but does not generate text, images, audio, or video and is not verified as a current OpenAI hosted API model.

Embeddings Reasoning Coding
OpenAI CLIP, short for Contrastive Language-Image Pre-Training, connects images and natural-language descriptions in a shared embedding space. Its publicly released checkpoints can turn images and text into comparable vector representations, making CLIP useful for zero-shot classification, image-text similarity, semantic image retrieval, and multimodal indexing. It is an embedding and matching model rather than a conversational or generative AI system.
Outputs

What CLIP can produce

Embeddings
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family CLIP
Model type Multimodal
Context window 77 tokens
Release date 2021-01-05
Status Public research release with downloadable weights; not verified as a current OpenAI hosted API model
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was published for the CLIP release. CLIP is trained for image-text representation learning rather than general factual question answering.

Model notes

CLIP is a family-level public research release rather than one independently deployed checkpoint. The official repository provides RN50, RN101, RN50x4, RN50x16, RN50x64, ViT-B/32, ViT-B/16, ViT-L/14, and ViT-L/14@336px variants. The 77-token context length applies to the released text encoder implementation. CLIP returns image and text representations and similarity scores, not natural-language responses. The model card identifies the release as a research output and warns that users should evaluate bias, robustness, and suitability for their own deployment context.

Model guide

CLIP: OpenAI’s Vision-Language Model for Zero-Shot Image Understanding

CLIP is an OpenAI research model that learns shared representations for images and text, allowing applications to compare them, perform zero-shot image classification, and build cross-modal search or retrieval systems without training a task-specific classifier.

What is CLIP?

CLIP, or Contrastive Language-Image Pre-Training, is an OpenAI vision-language research model released on January 5, 2021. It was trained on large collections of image-text pairs to learn which images and descriptions belong together. The result is a model that represents images and text in a common mathematical space, allowing software to measure how closely an image matches a phrase.

This design makes CLIP different from a conventional image classifier. A traditional classifier is usually trained for a fixed list of categories, such as cats, dogs, and cars. CLIP can instead compare an image with natural-language candidates supplied at inference time. For example, an application can score an image against prompts such as “a photo of a dog” and “a photo of a cat” and select the more compatible description. This is known as zero-shot classification because the model can evaluate the categories without a new task-specific training phase.

CLIP is the primary subject here, not a current OpenAI hosted conversational model. OpenAI released its code and model weights for research and local or compatible third-party deployment. The supplied research does not verify CLIP as a current model available through OpenAI’s hosted API catalog.

How CLIP connects images and text

CLIP contains two main encoders: an image encoder and a text encoder. The image encoder converts an image into a vector, while the text encoder converts a phrase into a vector of the same general representation space. During training, matching image-text pairs are encouraged to have high similarity, while unrelated pairs are pushed farther apart.

At use time, an application can encode one image and several candidate text descriptions, then compare their vectors. The highest similarity score indicates which description is most compatible according to the model. The system therefore does not need to generate an explanation or write a textual answer. Its core output is a set of image and text representations, along with similarity scores that downstream software can use.

The released implementation uses a fixed text context of 77 tokens. In practical terms, this limits the length of the text sequence that the text encoder can process in the published implementation. The exact image resolution and computational requirements depend on the selected checkpoint and its image encoder.

Public CLIP variants and compute trade-offs

OpenAI’s public repository provides several CLIP checkpoints with different image encoders. The listed variants are RN50, RN101, RN50x4, RN50x16, RN50x64, ViT-B/32, ViT-B/16, ViT-L/14, and ViT-L/14@336px.

The names indicate different architecture and scale choices. RN variants use ResNet-based image encoders, while ViT variants use Vision Transformer image encoders. Smaller checkpoints such as RN50 and ViT-B/32 are generally easier to run because they require less memory and computation. Larger variants, including RN50x64 and ViT-L/14@336px, generally offer stronger representation quality at the cost of greater resource requirements. These are practical trade-offs associated with the model family; the supplied research does not provide a single benchmark score that can be used to rank every variant for every task.

Choosing a checkpoint is therefore an engineering decision. A lightweight variant may be preferable for rapid experimentation, high-throughput indexing, or hardware with limited memory. A larger variant may be worth evaluating when representation quality is more important than inference cost. The best choice can also vary with image domain, preprocessing, prompt design, and the intended retrieval or classification task.

What CLIP can do

CLIP’s main capability is cross-modal representation: it places images and text into a space where related items can be compared. Common uses include:

  • Zero-shot image classification: compare an image with natural-language class prompts without training a new classifier for each label set.
  • Image-text similarity: score how well a description corresponds to an image.
  • Semantic image retrieval: search an image collection with text queries or find images related to a reference image.
  • Cross-modal ranking: rank candidate images for a description, or candidate descriptions for an image.
  • Image organization and tagging: use prompt-based similarity to help group, label, or explore visual datasets.
  • Multimodal research: study shared representations between visual and language data.

CLIP outputs image embeddings, text embeddings, and similarity scores. It does not produce natural-language responses, generated images, audio, or video. It also does not provide a built-in conversational interface, general-purpose reasoning answer, or guaranteed structured response format. Those distinctions matter when evaluating CLIP against generative multimodal models.

Using prompts for zero-shot classification

A zero-shot classifier built with CLIP typically defines a set of candidate descriptions, encodes those descriptions, encodes an input image, and compares the resulting vectors. Instead of treating a class as only a bare label such as “dog,” an application may use a phrase such as “a photo of a dog.” The wording can influence the similarity scores, so prompt selection and testing are part of the application design.

This approach is useful when categories change frequently or when collecting labeled training examples is impractical. A retailer, for example, could explore whether product images are closer to descriptions such as “a red running shoe” or “a black hiking boot.” A dataset team could use text queries to inspect a large image collection before building a more specialized pipeline.

However, zero-shot classification should not be treated as a guarantee of reliable classification in every domain. Performance can depend on the wording of prompts, the image distribution, preprocessing, and the selected checkpoint. Domain-specific images, unusual categories, ambiguous scenes, and sensitive decisions require evaluation rather than assuming that a high similarity score is a dependable ground-truth label.

Strengths and limitations

CLIP’s central strength is flexibility without task-specific classifier training. The same image and text encoders can support classification, search, ranking, and dataset exploration. Because text descriptions define the comparison targets, users can change the candidate concepts without rebuilding the entire model.

Its main limitation is that it is a representation model, not a complete application or answer-generating assistant. CLIP does not explain why an image received a score, write a detailed description, follow tool instructions, or produce a conversational response. It is also not a substitute for a managed inference service: deployment generally involves running the released implementation and weights yourself or using a compatible third-party implementation.

The model card identifies the release as primarily intended for research and warns that users should assess bias and suitability for their own context. Because CLIP learns from image-text data, it can reflect limitations and biases in those data. Robustness may also vary across image domains, prompt styles, and categories. Applications used for consequential decisions should perform task-specific testing and establish appropriate human review and safeguards.

Availability, pricing, and API positioning

The supplied research describes CLIP as a public research release with downloadable weights and an official open-source repository. No current OpenAI hosted API pricing is verified for CLIP, and no recurring hosted inference price is supplied. As a result, there is no reliable provider price to quote for per-image or per-token usage on this page.

That does not mean deployment has no cost. A self-hosted implementation can require suitable compute, storage, engineering work, monitoring, and maintenance. A third-party service may add its own usage fees and operational constraints, but those terms are separate from the official OpenAI CLIP release and are not established by the supplied research.

CLIP also lacks the characteristics normally associated with a hosted generative API model in this dataset: tool use, streaming responses, conversational text output, and guaranteed JSON output are not verified capabilities of the released model. Its output is primarily embeddings and similarity information.

When to choose CLIP

Choose CLIP when the core problem is comparing visual content with language or with other visual representations. It is a good candidate for:

  • prototyping text-to-image search over a private collection;
  • exploring and organizing image datasets;
  • testing zero-shot labels before investing in supervised training;
  • building image-text ranking or similarity features;
  • researching vision-language embeddings; and
  • running a downloadable research model where local control is more important than a managed endpoint.

Its family of checkpoints also gives engineers a speed-and-cost choice. Smaller variants are more practical when throughput, latency, or limited hardware matters. Larger variants may be considered when representation quality justifies additional compute, although the appropriate choice should be validated on the target data.

When another option may be more appropriate

A generative multimodal model is more appropriate when the application must answer questions about an image in natural language, follow a conversation, generate text, or combine vision with tool calls. A dedicated image-generation model is appropriate when the goal is to create or edit images. Speech or video tasks require models that support those modalities, which CLIP does not.

A supervised vision classifier may be preferable when the label set is stable, labeled examples are available, and consistent task-specific accuracy is more important than flexible natural-language prompting. A managed embedding or vision service may also be a better operational fit when a team does not want to host model weights and inference infrastructure.

CLIP remains most useful when its actual output—shared image-text representations and similarity scores—is exactly what the system needs. It should be selected for that focused role rather than treated as a general-purpose assistant or a current OpenAI API model.


Answers to Frequently Asked Questions

Can CLIP perform zero-shot image classification?
Yes. CLIP can classify images without task-specific retraining by comparing an image with natural-language prompts such as "a photo of a dog" or "a photo of a cat." The prompt wording, image domain, preprocessing, and selected checkpoint can affect the results.
What is CLIP and how does it work?
CLIP, or Contrastive Language-Image Pre-Training, is an OpenAI vision-language model that maps images and text into a shared mathematical space. It uses an image encoder and a text encoder, then compares their vectors to determine how closely an image matches a description.
What are the main uses of CLIP?
CLIP can support zero-shot image classification, image-text similarity scoring, text-to-image search, semantic image retrieval, cross-modal ranking, image tagging, dataset organization, and vision-language research. Its primary outputs are image embeddings, text embeddings, and similarity scores.
Is CLIP available through a current OpenAI hosted API, and what does it cost?
The supplied research describes CLIP as a public research release with downloadable code and model weights, but it does not verify CLIP as a current model in OpenAI’s hosted API catalog. No official hosted inference pricing is established. Self-hosting can still involve compute, storage, engineering, monitoring, and maintenance costs.
Does CLIP generate text or provide conversational answers?
No. CLIP is a representation model rather than a generative conversational assistant. It does not normally generate natural-language explanations, answer questions about images, create images, or provide built-in tool use. Applications must use its embeddings and similarity scores for downstream tasks.


Sources 4
Provider

About OpenAI