WildDet3D

WildDet3D

by Allen Institute for Artificial Intelligence (Ai2) · Current open-weight research model

Open-weight monocular 3D detection model that uses text, point, and 2D box prompts to predict 3D object boxes, depth maps, and camera intrinsics from RGB images, with optional depth input for improved localization.

Reasoning Coding
WildDet3D brings open-vocabulary prompting to monocular 3D perception. Instead of detecting only a fixed set of categories, it can use a text query, point click, or 2D box to identify an object and estimate its 3D position, dimensions, and orientation. The downloadable research model is aimed at spatial perception, robotics, augmented reality, and experiments that need 3D understanding from ordinary images.
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
2/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family WildDet3D
Model type Other
Context window tokens
Maximum output tokens
Release date 2026-04-07
Status Current open-weight research model
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was published. WildDet3D is a vision detection model whose behavior depends primarily on its trained visual and geometric representations rather than a stated language-model knowledge date.

Model notes

WildDet3D is a specialized vision model rather than a language model. The final released paper checkpoint is wilddet3d_alldata_all_prompt_v1.0.pt and is approximately 4.7 GB. It combines a SAM 3 ViT backbone with a LingBot-Depth DINOv2 ViT-L/14 geometry backend and has approximately 1.2 billion parameters. Inputs include an RGB image, optional camera intrinsics, optional sparse or dense depth, and text, point, or box prompts. Outputs include 2D boxes, 3D boxes, depth maps, scores, class identifiers, and predicted intrinsics. It is released for research and educational use under the SAM License and is not deployed through an official hosted inference provider.

Cost

Model pricing

Input No official hosted API pricing; downloadable checkpoint
Output No official hosted API pricing; downloadable checkpoint
Model guide

WildDet3D: Promptable Open-World 3D Detection from a Single Image

WildDet3D is an open-weight monocular 3D object detection model from the Allen Institute for AI. It uses text, point, or 2D box prompts to locate selected objects in three dimensions from a single RGB image, with optional camera intrinsics and sparse or dense depth to improve localization.

What Is WildDet3D?

WildDet3D is an open-weight model for promptable monocular 3D object detection, developed by the Allen Institute for AI and academic collaborators. Its main task is to infer where an object is in three-dimensional space from a single RGB image. The output can include a 2D bounding box, a 3D bounding box, depth information, confidence scores, class identifiers, and predicted camera intrinsics.

“Monocular” means that the model starts with one camera view rather than requiring a stereo camera pair or a complete 3D sensor setup. This is useful for applications that have ordinary photographs or video frames but do not have a fully calibrated perception system. The model can also accept optional geometric information, including camera intrinsics and sparse or dense depth, when that information is available.

WildDet3D is not a conversational assistant or a general-purpose generative model. It is a specialized vision system for identifying prompted objects and estimating their 3D geometry.

How Prompting Works

The model supports several ways to specify what should be detected. A text prompt can request objects by category, while a point or box can identify a particular object already visible in the image.

  • Text prompts: Ask the model to find instances of a named object category.
  • Point prompts: Use positive or negative image points to indicate the object of interest or refine the selection.
  • Geometric box prompts: Provide a 2D bounding box and ask the model to lift it into 3D.
  • Visual exemplar boxes: Use a box around an example object to help locate visually similar objects.

This promptable design separates WildDet3D from a conventional detector that always runs against a predefined label list. It can be used to investigate a specific object, reuse an existing 2D detection, or explore categories that are not part of a narrowly fixed detection interface. The supplied research describes evaluation across more than 700 object categories.

Inputs and Outputs

The basic input is an RGB image. Camera intrinsics, which describe properties such as focal length and the camera's optical center, can be supplied when known. If they are not available, WildDet3D can estimate them internally. Optional sparse or dense depth can also be provided. Examples include LiDAR, time-of-flight, RGB-D, or stereo-derived depth.

Additional depth information can improve the model's 3D localization, but it is not required for the core single-image use case. The model's outputs are perception results rather than generated prose or media:

  • 2D bounding boxes
  • 3D bounding boxes
  • Depth maps
  • Detection scores
  • Class identifiers
  • Predicted camera intrinsics

WildDet3D accepts visual inputs and text or spatial prompts, but it does not produce images, video, audio, or natural-language responses as its primary output. There is no published language-model context window or maximum output-token limit because this is not a text-generation model.

Architecture and Released Checkpoints

The released WildDet3D system combines a SAM 3 Vision Transformer backbone with a LingBot-Depth geometry backend based on DINOv2 ViT-L/14. The supplied specifications describe the model as having approximately 1.2 billion parameters. In practical terms, the architecture combines visual-semantic features for recognizing prompted objects with geometry estimation for recovering their spatial properties.

Ai2 released three checkpoint stages. The final paper checkpoint is wilddet3d_alldata_all_prompt_v1.0.pt. It was trained using Omni3D data together with in-the-wild, human-verified prompt mixtures. Earlier Stage 1 and Stage 2 checkpoints are available for users interested in partial training reproduction or comparing development stages.

The final checkpoint is approximately 4.7 GB, so local use requires suitable storage and a compatible software and hardware environment. The supplied research specifically identifies a PyTorch and CUDA setup as part of the expected local deployment environment.

Reported Performance

According to the supplied Ai2 research, WildDet3D achieves 34.2 AP3D with text prompts and 36.4 AP3D with oracle box prompts on Omni3D. With sparse depth supplied, the reported results increase to 41.6 AP3D with text prompts and 45.8 AP3D with oracle prompts.

For zero-shot evaluations, Ai2 reports 40.3 ODS on Argoverse 2 and 48.9 ODS on ScanNet. These figures are provider-reported benchmark results, not independent guarantees of performance on every image or deployment environment. Results can vary with camera viewpoint, image quality, object category, prompt quality, available depth, and the target dataset.

The distinction between text prompts and oracle box prompts is important. An oracle box gives the model a highly informative 2D localization supplied by the evaluation setup, while a text prompt requires the system to find the relevant object from the image. As a result, the two numbers represent different operating conditions rather than interchangeable measurements.

Main Strengths and Trade-offs

WildDet3D's central strength is the combination of open-world prompting and 3D localization. A user does not have to rely only on a closed detector with a fixed category list, and the system can work from one RGB image. The ability to combine text, point, and box prompts also makes it suitable for interactive perception workflows in which a human or another vision system specifies the object of interest.

  • Open-weight research access: Checkpoints, code, evaluation tools, training materials, an interactive demo, and an iOS application are available according to the supplied research.
  • Multiple prompt types: Text and spatial prompts support both category-level search and instance-level selection.
  • 3D outputs from images: The model estimates spatial boxes and depth-related information instead of stopping at 2D detection.
  • Optional geometric assistance: Camera calibration and depth can be used when available to improve localization.
  • Research flexibility: Local checkpoints allow researchers to inspect and adapt the system rather than depending on a proprietary hosted endpoint.

The trade-offs are equally important. WildDet3D is a large, specialized checkpoint rather than a lightweight, low-latency service. Running it locally involves downloading a roughly 4.7 GB model and configuring the required PyTorch and CUDA environment. The supplied information does not provide a standardized latency figure, hardware recommendation, hosted uptime guarantee, or official API endpoint.

Its capabilities are also narrow by design. It does not provide conversational generation, code generation, image generation, audio processing, video synthesis, function calling, web search, streaming responses, or structured text generation. Those omissions are not shortcomings for a 3D detector, but they make other model types more appropriate for assistant, content-generation, or API-automation tasks.

Pricing and Availability

There is no official hosted API pricing listed for WildDet3D. The model is provided as a downloadable checkpoint rather than as a recurring consumer subscription or metered inference API. Users should therefore think about local infrastructure, storage, GPU availability, and engineering time as the main operational costs.

The model is intended primarily for research and educational use under the SAM License and Ai2 responsible-use guidance. License conditions and any restrictions should be reviewed before commercial deployment or redistribution. Availability may also change as the project develops, because research models, demos, and downloadable artifacts can move between experimental and more stable stages.

Reasoning, Coding, and Tool Support

WildDet3D should not be evaluated using the same criteria as a language model. It does not perform general textual reasoning, write code, call external tools, browse the web, or return JSON through a model-native structured-output mode. Its useful “reasoning” is specialized visual and geometric inference: interpreting prompts, associating them with image regions, estimating depth, and producing 3D boxes.

External software can of course use the model's outputs inside a larger robotics, AR, or computer-vision pipeline. That is application-level integration, not built-in tool or function support supplied by the checkpoint itself.

Best Use Cases

WildDet3D is a strong candidate when the central problem is open-vocabulary 3D perception from images. Suitable uses include:

  • Research into open-world object detection and spatial intelligence
  • Robotics experiments that need approximate 3D object locations
  • Augmented-reality prototypes using monocular camera input
  • Converting text, point, or 2D detection prompts into 3D boxes
  • Benchmarking promptable detection across varied object categories
  • Studying how optional depth and camera calibration affect monocular 3D localization

It is less suitable for a production application that needs a small, fast, managed inference API, predictable latency, or guaranteed availability. A conventional specialized detector may be preferable when the object categories are fixed and throughput is more important than open-world prompting. A hosted computer-vision service may be preferable when the team does not want to manage CUDA dependencies or GPU infrastructure. A multimodal language model is a better fit when the required result is an explanation, code, document analysis, or a conversational workflow rather than metric 3D boxes.

When to Choose WildDet3D

Choose WildDet3D when you need an inspectable research model that can connect flexible prompts with 3D object localization from a single image. Its most distinctive value is the combination of open-world prompting, monocular input, and downloadable weights. It becomes more attractive when you can provide depth or camera information and have the hardware and engineering resources to run a large checkpoint locally.

Choose another option when the priority is low-cost hosted inference, high and predictable speed, a compact deployment footprint, language interaction, or media generation. WildDet3D is best understood as a specialized spatial-perception component, not as a replacement for a general AI assistant or an all-purpose vision API.


Answers to Frequently Asked Questions

What types of prompts does WildDet3D support?
WildDet3D supports text prompts, positive or negative point prompts, 2D geometric box prompts, and visual exemplar boxes. These prompts can specify an object category, select a particular instance, or help locate visually similar objects.
What is WildDet3D used for?
WildDet3D is an open-weight model for promptable monocular 3D object detection. It identifies objects from a single RGB image and estimates their 2D and 3D bounding boxes, depth-related information, confidence scores, class identifiers, and camera intrinsics.
What inputs and outputs does WildDet3D support?
The primary input is an RGB image. Users can optionally provide camera intrinsics and sparse or dense depth from sources such as LiDAR, RGB-D, time-of-flight, or stereo systems. Outputs include 2D and 3D bounding boxes, depth maps, detection scores, class identifiers, and predicted camera intrinsics.
Does WildDet3D offer a hosted API or commercial pricing?
No official hosted API pricing is listed for WildDet3D. It is provided as a downloadable checkpoint, so users must account for local storage, GPU infrastructure, software setup, and engineering costs. The model is primarily intended for research and educational use under the SAM License and Ai2 responsible-use guidance.
What are the WildDet3D checkpoints and hardware requirements?
The final paper checkpoint is wilddet3d_alldata_all_prompt_v1.0.pt and is approximately 4.7 GB. WildDet3D combines a SAM 3 Vision Transformer backbone with a LingBot-Depth geometry backend based on DINOv2 ViT-L/14, and local use requires a compatible PyTorch and CUDA environment.


Sources 4
Provider

About Allen Institute for Artificial Intelligence (Ai2)