DINOv3

DINOv3

by Meta AI · Current; downloadable open-weight research model suite

Meta AI’s DINOv3 is a downloadable self-supervised vision model suite that produces reusable image embeddings and dense visual features. Its Vision Transformer and ConvNeXt variants support classification, retrieval, segmentation, depth estimation, object discovery, tracking pipelines, and geospatial analysis, but it does not generate text, images, audio, or video directly.

Embeddings Reasoning Coding
DINOv3 is a family of downloadable vision backbones developed by Meta AI Research. Instead of generating text, images, or video, the models convert images into reusable visual representations. These embeddings and feature tokens can support classifiers, retrieval systems, segmentation heads, depth-estimation systems, tracking pipelines, and other computer-vision applications. The suite includes large and distilled Vision Transformer models, ConvNeXt variants, and models pretrained on both web imagery and satellite imagery.
Outputs

What DINOv3 can produce

Embeddings
Inputs

What it can understand

Images
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

0/10 Reasoning
0/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family DINOv3
Model type Multimodal
Context window tokens
Maximum output tokens
Release date 2025-08-14
Status Current; downloadable open-weight research model suite
Knowledge cutoff notes

DINOv3 is an image representation model rather than a knowledge-grounded language model. Meta’s public documentation does not specify a conventional knowledge cutoff date.

Model notes

DINOv3 is a family-level model identity covering multiple Vision Transformer and ConvNeXt backbones rather than one single checkpoint. The official suite includes web-pretrained and satellite-pretrained variants, including models from small architectures through a 7B-parameter ViT. Models return visual representations such as class tokens, patch tokens, register tokens, and pooled features. The weights and code are distributed under the DINOv3 License. No official first-party hosted API pricing, context window, maximum text output, or token limits were identified. Scores for reasoning and coding are not applicable because DINOv3 is a vision representation model rather than a language or code model.

Model guide

DINOv3: Meta AI’s Self-Supervised Vision Foundation Model for Dense Image Features

DINOv3 is Meta AI’s self-supervised computer-vision foundation model suite for producing high-quality image representations and dense visual features. Its Vision Transformer and ConvNeXt backbones are intended for downstream tasks including image classification, retrieval, segmentation, depth estimation, object discovery, video tracking, and geospatial analysis.

What is DINOv3?

DINOv3 is Meta AI’s self-supervised vision foundation model suite. It learns general-purpose visual representations from images without depending primarily on human-labeled examples. The resulting representations are designed to remain useful across multiple computer-vision tasks instead of being tied to one narrow application.

In practical terms, DINOv3 is a visual backbone. You provide an image, and the model produces features that describe visual content in a form that other software can use. Those features can support image classification, visual similarity search, semantic segmentation, depth estimation, object discovery, tracking, and geospatial analysis.

DINOv3 is not a conversational model and does not operate as a conventional text-generation system. It does not directly write captions, answer questions, generate images, or produce video. A separate task-specific model or processing pipeline is needed when the application requires those outputs.

Where DINOv3 fits in Meta AI’s lineup

DINOv3 is a research model family rather than a single hosted endpoint or one fixed-size checkpoint. Meta AI released downloadable reference code and pretrained weights for local or self-managed use. The suite is positioned for researchers and developers who need reusable visual features, including teams building specialized computer-vision systems.

The release includes models trained for different visual domains. Web-image models are trained on the LVD-1689M dataset, while satellite-oriented variants are trained on the SAT-493M dataset. This distinction matters because a model trained on satellite imagery may be a more appropriate starting point for aerial or geospatial work than a general web-image checkpoint.

The model family includes Vision Transformer backbones, ConvNeXt backbones, smaller distilled variants, and a Vision Transformer model with up to 7 billion parameters. These choices allow users to trade representation capacity against memory use, deployment cost, and processing speed.

Architecture and outputs

DINOv3 includes both Transformer-based and convolutional architectures. Vision Transformer models divide an image into patches and process the resulting visual tokens. The Transformer variants use a patch size of 16 and can return class tokens, patch tokens, and register tokens. These outputs provide different levels of information: a class token can summarize an image, while patch-level tokens preserve more localized visual detail.

Dense features are particularly important for tasks that need to understand where something appears in an image. For example, a segmentation system can use patch-level representations to assign labels to image regions, while a retrieval system may use pooled or global features to compare complete images.

The models support larger image resolutions when the dimensions are compatible with the model’s patch size. However, the supplied documentation does not specify one universal context window, maximum image size, maximum output-token limit, or fixed inference resolution for the entire DINOv3 family. Those details depend on the selected checkpoint and deployment configuration rather than on a single family-wide value.

How DINOv3 learns visual representations

DINOv3 uses self-supervised learning. Instead of requiring a human to label every training image, the training process creates different augmented views of the same image and teaches the model to produce consistent representations for those views.

The reported training approach combines several techniques: DINO self-distillation, iBOT masked-image modeling, KoLeo regularization, and Gram anchoring. Gram anchoring is intended to help preserve the quality of dense feature maps during long training schedules. Together, these methods are designed to produce representations that remain useful when transferred to tasks for which the original training data did not provide direct labels.

This design makes DINOv3 different from a task-specific classifier. A classifier is usually trained to predict a predefined set of labels, while DINOv3 produces general visual information that can be connected to a classifier, retrieval index, segmentation head, or another downstream component.

Main capabilities and use cases

DINOv3 is most useful when an application needs a strong image representation rather than a complete end-user AI interaction. Common uses include:

  • Image classification: DINOv3 features can be paired with k-nearest-neighbor search, logistic regression, or lightweight linear classifiers.
  • Image retrieval: Global or pooled embeddings can help find visually similar images in a collection.
  • Semantic segmentation: Dense patch features can provide inputs to systems that label individual regions or pixels.
  • Depth estimation: Visual representations can be connected to a depth-estimation head for predicting scene structure.
  • Object discovery: Patch-level features can help identify recurring objects or regions without starting from a fully labeled detection dataset.
  • Video pipelines: Frame-level features can be used by separate temporal systems for video segmentation, tracking, or video classification.
  • Geospatial analysis: Satellite-pretrained variants are intended for aerial and satellite imagery tasks, including canopy-height estimation and aerial-object analysis.
  • Specialized research: The representations may serve as a starting point for medical imaging, robotics, and scientific vision experiments.

A common deployment pattern is to keep the DINOv3 backbone frozen and train a lightweight adapter or task-specific head on top of it. This can reduce training requirements and make experimentation easier. Full fine-tuning remains possible, but it may require substantially more compute and can make results more dependent on the quality and representativeness of the adaptation data.

Strengths and practical trade-offs

DINOv3’s main strength is reuse. One downloaded backbone can provide features for several visual tasks, reducing the need to train a separate representation model from scratch for every application. The availability of both dense patch features and global representations also makes the family suitable for tasks ranging from image-level search to region-level analysis.

The range of checkpoints is another practical advantage. Smaller Vision Transformer and ConvNeXt variants are more suitable for constrained environments, while larger models may offer greater representation capacity at the cost of memory and processing requirements. Distilled models provide another way to reduce deployment cost compared with the largest teacher models.

These benefits involve trade-offs. Larger checkpoints require more substantial hardware, and the family does not provide a single speed or memory profile. A model that is appropriate for offline research may be too expensive for high-throughput or edge deployment. Conversely, a smaller model may be easier to operate but could provide a different accuracy or feature-quality trade-off for a particular task. The supplied research does not provide a standardized benchmark table that would justify a universal ranking of the variants.

Pricing and deployment

DINOv3 is distributed as downloadable model code and pretrained weights through Meta AI’s research repositories and associated model collections. It is intended for local or self-managed inference rather than a metered, first-party token-based inference API.

No official hosted API pricing, recurring subscription price, context limit, maximum text-output limit, or token-based billing schedule was identified in the supplied research. Users therefore need to account for their own infrastructure, storage, memory, and inference costs. The largest checkpoints can require substantial compute, while smaller Vision Transformer, ConvNeXt, and distilled variants are more practical for resource-constrained systems.

The weights and code are distributed under the DINOv3 License. Before commercial deployment, redistribution, modification, or creation of derivative model materials, users should review the license and confirm that their intended use is permitted.

Supported inputs and outputs

CategoryDINOv3 support
Primary inputImages
Primary outputImage embeddings, class tokens, patch tokens, register tokens, and related visual features
Text input or outputNot supported as a native model function
Audio or speechNot supported
Direct image or video generationNot supported
VideoFrame-level features can support separate video pipelines; DINOv3 is not itself a video-generation or complete video-understanding service
Tool or function callingNot provided
Structured text outputNot provided

DINOv3 should therefore be evaluated as a representation model, not as a multimodal assistant. Although it processes visual input, its direct output is visual feature data rather than a user-facing caption, explanation, or generated media file.

Reasoning, coding, and API limitations

Reasoning and coding scores are not applicable in the usual language-model sense. DINOv3 does not generate a chain of reasoning, write software, execute code, or answer programming questions. It can contribute visual features to a larger system that performs those functions, but those capabilities would come from the additional system components rather than from DINOv3 itself.

The research also does not identify a built-in web-search tool, function-calling interface, structured-output mode, hosted batch API, or streaming text-generation interface. There is no conventional maximum output-token limit because the model does not produce text completions. Output size instead depends on the selected architecture, image resolution, and the feature tensors returned by the implementation.

Limitations and risks

DINOv3 does not replace downstream task design. A raw embedding is not automatically a classification result, segmentation mask, depth map, or tracking decision. Developers must select or train an appropriate head, adapter, index, or temporal processing system and validate it on the target domain.

Domain shift is another limitation. Features learned from web imagery may not behave like features learned from satellite imagery, medical images, industrial cameras, or specialized scientific data. Even when a model transfers successfully, performance can vary across geographic regions and income categories. The supplied model documentation reports remaining differences in fairness evaluations, including a notable gap between lower-income and higher-income groups.

Fine-tuning can improve performance on a specific task, but it can also amplify bias when adaptation labels are incomplete or unrepresentative. Evaluation should therefore include the actual environments, populations, image conditions, and failure cases expected in deployment.

When to choose DINOv3

Choose DINOv3 when you need downloadable visual features and want control over deployment. It is a good fit for image retrieval, reusable computer-vision backbones, dense image analysis, segmentation experiments, depth estimation, object discovery, and satellite or aerial imagery research. It is also attractive when a team wants to freeze a backbone and train a relatively small task-specific component.

A smaller or distilled DINOv3 checkpoint may be more appropriate when memory, throughput, or operational cost matters more than the maximum available model capacity. A satellite-pretrained variant is the more relevant starting point for geospatial imagery when the application matches that domain.

Another type of model is more appropriate when the requirement is conversational assistance, caption generation, visual question answering, code generation, direct image or video generation, speech, or a managed API with usage-based billing. DINOv3 can be one component in such a system, but it is not itself a complete assistant or hosted application.

Bottom line

DINOv3 is a family of Meta AI vision backbones for turning images into reusable global and dense visual representations. Its value lies in transfer: the same general-purpose features can support retrieval, classification, segmentation, depth, tracking, and geospatial workflows. The main costs are operational complexity, hardware requirements for larger checkpoints, the need for downstream task components, and the absence of a conventional hosted text or multimodal assistant API.


Answers to Frequently Asked Questions

Which DINOv3 model should be used for satellite or resource-constrained applications?
Satellite-pretrained variants are the most relevant starting point for aerial and geospatial imagery. Smaller Vision Transformer, ConvNeXt, and distilled variants are better suited to environments where memory, processing speed, throughput, or deployment cost is more important than maximum model capacity.
How is DINOv3 deployed and priced?
DINOv3 is distributed as downloadable model code and pretrained weights for local or self-managed inference. It is not presented as a metered, first-party hosted API, and no official token-based pricing or subscription schedule is identified. Users must cover their own infrastructure, storage, memory, and inference costs.
What kinds of outputs does DINOv3 provide?
Depending on the architecture and configuration, DINOv3 can return image embeddings, class tokens, patch tokens, register tokens, and other visual features. Global features are useful for image-level tasks such as retrieval, while patch-level features preserve localized information for tasks such as segmentation.
What is DINOv3 and what is it used for?
DINOv3 is Meta AI’s self-supervised vision foundation model suite. It converts images into reusable visual features that can support image classification, visual similarity search, semantic segmentation, depth estimation, object discovery, tracking, and geospatial analysis.
Does DINOv3 generate text, captions, images, or video?
No. DINOv3 is a visual representation model, not a conversational or generative AI system. It produces image embeddings and visual feature tensors, while captioning, question answering, image generation, and video generation require additional models or processing pipelines.


Sources 7
Provider

About Meta AI