What Is WildDet3D?
WildDet3D is an open-weight model for promptable monocular 3D object detection, developed by the Allen Institute for AI and academic collaborators. Its main task is to infer where an object is in three-dimensional space from a single RGB image. The output can include a 2D bounding box, a 3D bounding box, depth information, confidence scores, class identifiers, and predicted camera intrinsics.
“Monocular” means that the model starts with one camera view rather than requiring a stereo camera pair or a complete 3D sensor setup. This is useful for applications that have ordinary photographs or video frames but do not have a fully calibrated perception system. The model can also accept optional geometric information, including camera intrinsics and sparse or dense depth, when that information is available.
WildDet3D is not a conversational assistant or a general-purpose generative model. It is a specialized vision system for identifying prompted objects and estimating their 3D geometry.
How Prompting Works
The model supports several ways to specify what should be detected. A text prompt can request objects by category, while a point or box can identify a particular object already visible in the image.
- Text prompts: Ask the model to find instances of a named object category.
- Point prompts: Use positive or negative image points to indicate the object of interest or refine the selection.
- Geometric box prompts: Provide a 2D bounding box and ask the model to lift it into 3D.
- Visual exemplar boxes: Use a box around an example object to help locate visually similar objects.
This promptable design separates WildDet3D from a conventional detector that always runs against a predefined label list. It can be used to investigate a specific object, reuse an existing 2D detection, or explore categories that are not part of a narrowly fixed detection interface. The supplied research describes evaluation across more than 700 object categories.
Inputs and Outputs
The basic input is an RGB image. Camera intrinsics, which describe properties such as focal length and the camera's optical center, can be supplied when known. If they are not available, WildDet3D can estimate them internally. Optional sparse or dense depth can also be provided. Examples include LiDAR, time-of-flight, RGB-D, or stereo-derived depth.
Additional depth information can improve the model's 3D localization, but it is not required for the core single-image use case. The model's outputs are perception results rather than generated prose or media:
- 2D bounding boxes
- 3D bounding boxes
- Depth maps
- Detection scores
- Class identifiers
- Predicted camera intrinsics
WildDet3D accepts visual inputs and text or spatial prompts, but it does not produce images, video, audio, or natural-language responses as its primary output. There is no published language-model context window or maximum output-token limit because this is not a text-generation model.
Architecture and Released Checkpoints
The released WildDet3D system combines a SAM 3 Vision Transformer backbone with a LingBot-Depth geometry backend based on DINOv2 ViT-L/14. The supplied specifications describe the model as having approximately 1.2 billion parameters. In practical terms, the architecture combines visual-semantic features for recognizing prompted objects with geometry estimation for recovering their spatial properties.
Ai2 released three checkpoint stages. The final paper checkpoint is wilddet3d_alldata_all_prompt_v1.0.pt. It was trained using Omni3D data together with in-the-wild, human-verified prompt mixtures. Earlier Stage 1 and Stage 2 checkpoints are available for users interested in partial training reproduction or comparing development stages.
The final checkpoint is approximately 4.7 GB, so local use requires suitable storage and a compatible software and hardware environment. The supplied research specifically identifies a PyTorch and CUDA setup as part of the expected local deployment environment.
Reported Performance
According to the supplied Ai2 research, WildDet3D achieves 34.2 AP3D with text prompts and 36.4 AP3D with oracle box prompts on Omni3D. With sparse depth supplied, the reported results increase to 41.6 AP3D with text prompts and 45.8 AP3D with oracle prompts.
For zero-shot evaluations, Ai2 reports 40.3 ODS on Argoverse 2 and 48.9 ODS on ScanNet. These figures are provider-reported benchmark results, not independent guarantees of performance on every image or deployment environment. Results can vary with camera viewpoint, image quality, object category, prompt quality, available depth, and the target dataset.
The distinction between text prompts and oracle box prompts is important. An oracle box gives the model a highly informative 2D localization supplied by the evaluation setup, while a text prompt requires the system to find the relevant object from the image. As a result, the two numbers represent different operating conditions rather than interchangeable measurements.
Main Strengths and Trade-offs
WildDet3D's central strength is the combination of open-world prompting and 3D localization. A user does not have to rely only on a closed detector with a fixed category list, and the system can work from one RGB image. The ability to combine text, point, and box prompts also makes it suitable for interactive perception workflows in which a human or another vision system specifies the object of interest.
- Open-weight research access: Checkpoints, code, evaluation tools, training materials, an interactive demo, and an iOS application are available according to the supplied research.
- Multiple prompt types: Text and spatial prompts support both category-level search and instance-level selection.
- 3D outputs from images: The model estimates spatial boxes and depth-related information instead of stopping at 2D detection.
- Optional geometric assistance: Camera calibration and depth can be used when available to improve localization.
- Research flexibility: Local checkpoints allow researchers to inspect and adapt the system rather than depending on a proprietary hosted endpoint.
The trade-offs are equally important. WildDet3D is a large, specialized checkpoint rather than a lightweight, low-latency service. Running it locally involves downloading a roughly 4.7 GB model and configuring the required PyTorch and CUDA environment. The supplied information does not provide a standardized latency figure, hardware recommendation, hosted uptime guarantee, or official API endpoint.
Its capabilities are also narrow by design. It does not provide conversational generation, code generation, image generation, audio processing, video synthesis, function calling, web search, streaming responses, or structured text generation. Those omissions are not shortcomings for a 3D detector, but they make other model types more appropriate for assistant, content-generation, or API-automation tasks.
Pricing and Availability
There is no official hosted API pricing listed for WildDet3D. The model is provided as a downloadable checkpoint rather than as a recurring consumer subscription or metered inference API. Users should therefore think about local infrastructure, storage, GPU availability, and engineering time as the main operational costs.
The model is intended primarily for research and educational use under the SAM License and Ai2 responsible-use guidance. License conditions and any restrictions should be reviewed before commercial deployment or redistribution. Availability may also change as the project develops, because research models, demos, and downloadable artifacts can move between experimental and more stable stages.
Reasoning, Coding, and Tool Support
WildDet3D should not be evaluated using the same criteria as a language model. It does not perform general textual reasoning, write code, call external tools, browse the web, or return JSON through a model-native structured-output mode. Its useful “reasoning” is specialized visual and geometric inference: interpreting prompts, associating them with image regions, estimating depth, and producing 3D boxes.
External software can of course use the model's outputs inside a larger robotics, AR, or computer-vision pipeline. That is application-level integration, not built-in tool or function support supplied by the checkpoint itself.
Best Use Cases
WildDet3D is a strong candidate when the central problem is open-vocabulary 3D perception from images. Suitable uses include:
- Research into open-world object detection and spatial intelligence
- Robotics experiments that need approximate 3D object locations
- Augmented-reality prototypes using monocular camera input
- Converting text, point, or 2D detection prompts into 3D boxes
- Benchmarking promptable detection across varied object categories
- Studying how optional depth and camera calibration affect monocular 3D localization
It is less suitable for a production application that needs a small, fast, managed inference API, predictable latency, or guaranteed availability. A conventional specialized detector may be preferable when the object categories are fixed and throughput is more important than open-world prompting. A hosted computer-vision service may be preferable when the team does not want to manage CUDA dependencies or GPU infrastructure. A multimodal language model is a better fit when the required result is an explanation, code, document analysis, or a conversational workflow rather than metric 3D boxes.
When to Choose WildDet3D
Choose WildDet3D when you need an inspectable research model that can connect flexible prompts with 3D object localization from a single image. Its most distinctive value is the combination of open-world prompting, monocular input, and downloadable weights. It becomes more attractive when you can provide depth or camera information and have the hardware and engineering resources to run a large checkpoint locally.
Choose another option when the priority is low-cost hosted inference, high and predictable speed, a compact deployment footprint, language interaction, or media generation. WildDet3D is best understood as a specialized spatial-perception component, not as a replacement for a general AI assistant or an all-purpose vision API.

