StreamPETR

StreamPETR

by NVIDIA AI · Current NVIDIA NIM endpoint; free endpoint access requires an NVIDIA API key

A specialized NVIDIA model for camera-only multi-view 3D object detection and tracking in autonomous-driving scenes. StreamPETR propagates sparse object queries across frames and produces annotated camera and bird's-eye-view videos through a free, scene-based NIM endpoint requiring an NVIDIA API key.

Video generation
StreamPETR is designed for a specific computer-vision problem: understanding road scenes from multiple camera views over time. Rather than serving as a general-purpose chatbot or image generator, it detects and tracks objects in autonomous-driving footage, estimates their 3D positions, and presents the results through annotated camera-view and bird's-eye-view videos. NVIDIA currently exposes the model through a free NIM endpoint that requires an NVIDIA API key.
Outputs

What StreamPETR can produce

Video generation
Inputs

What it can understand

Images Video Multimodal input
Capabilities

Supported features

Streaming Structured output Multimodal output
Model profile

Performance characteristics

8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family StreamPETR
Model type Other
Release date 2023-03-21
Status Current NVIDIA NIM endpoint; free endpoint access requires an NVIDIA API key
Knowledge cutoff notes

No authoritative knowledge-cutoff date was found for this perception model.

Model notes

StreamPETR was introduced in the 2023 research paper 'Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection' and has an official research implementation. NVIDIA currently exposes a model named streampetr through its NIM catalog. The documented API accepts a scene_id and optional configuration, then returns annotated camera and bird's-eye-view MP4 videos plus inference metadata in a ZIP or redirected object-storage response. The NVIDIA endpoint is specialized and scene-based; its documented interface does not establish a general raw-frame upload contract. The original research repository reports nuScenes benchmark results and supports TensorRT inference, streaming-video processing, and 3D object tracking.

Cost

Model pricing

Input Free endpoint
Model guide

NVIDIA StreamPETR: Camera-Based 3D Perception for Autonomous Driving

NVIDIA StreamPETR is a specialized camera-only, multi-view 3D object-detection and tracking model for autonomous-driving perception. It uses temporal object-query propagation to analyze driving scenes efficiently and is available through a free, scene-based NVIDIA NIM endpoint that returns annotated camera and bird's-eye-view videos.

What is NVIDIA StreamPETR?

StreamPETR is a camera-based, multi-view 3D object-detection model for autonomous-driving perception. Its job is to identify objects such as vehicles and other road users in surrounding camera feeds, estimate their positions in three-dimensional space, and track those objects as the scene changes over time.

The name reflects the model's approach: it uses object-centric temporal modeling based on sparse object queries. In practical terms, StreamPETR does not need to rebuild its entire understanding of the scene from scratch for every frame. Instead, it propagates compact representations of detected objects across consecutive frames. This design is intended to make streaming multi-camera perception more efficient while preserving information about object motion and identity.

StreamPETR was introduced in the 2023 research paper Exploring Object-Centric Temporal Modeling for Efficient Multi-View 3D Object Detection. NVIDIA now lists it in its NIM catalog as a specialized inference endpoint under the model name streampetr.

Primary purpose and positioning

StreamPETR belongs to the perception layer of an autonomous-driving system. Perception models turn sensor observations into structured information about the surrounding environment. For StreamPETR, the sensor input is camera video from multiple views rather than lidar or language prompts.

This makes the model substantially narrower than a general computer-vision or multimodal model. It is intended for driving-scene analysis, 3D object detection, multi-object tracking, and bird's-eye-view visualization. It is not a conversational model, a coding model, an image-generation system, or a general video-understanding assistant.

NVIDIA's current hosted offering is also more constrained than a typical arbitrary-video API. The documented interface accepts a predefined scene_id with optional configuration and returns processed results for that scene. The supplied research does not establish a general contract for uploading arbitrary raw frames or any user-selected video directly to the public endpoint.

How StreamPETR processes driving scenes

Multi-view 3D detection is difficult because each camera sees only a two-dimensional projection of the world. A perception system must combine views, infer depth and 3D position, and maintain consistent object identities as vehicles and other objects move through the scene.

StreamPETR addresses the temporal part of this problem by carrying sparse object queries from one frame to the next. An object query can be understood as a compact internal representation associated with a possible detected object. Propagating these queries allows the model to reuse information from earlier frames instead of repeatedly processing every historical image feature.

The result is a streaming-oriented approach to multi-view perception. The original research implementation supports 3D object detection and tracking, and the NVIDIA endpoint produces annotated camera-view and bird's-eye-view MP4 videos alongside inference metadata. The bird's-eye-view representation is useful because it presents the scene from above, making the estimated locations and movement of detected objects easier to inspect.

Capabilities and supported modalities

StreamPETR's supported input and output types are specialized:

AreaWhat is supported or documented
InputCamera-based multi-view driving-scene data through the documented scene-based NIM interface
VideoStreaming-video processing and driving-scene analysis
Perception3D object detection and multi-object tracking
OutputAnnotated camera-view videos, annotated bird's-eye-view videos, and inference metadata
AudioNot supported
Text generationNot supported
Image generationNot supported
Tool or function callingNot supported as a model capability

The model therefore has multimodal behavior in the computer-vision sense, but it should not be confused with a general-purpose multimodal assistant. Its output is visual annotation and perception metadata rather than natural-language answers. The supplied specifications identify video output, but they do not provide a general maximum output duration, frame limit, context window, or maximum-token value.

Research performance

The original StreamPETR research implementation reports results on the nuScenes autonomous-driving benchmark. For a ViT-Large configuration on the test set, the paper reports 62.0 mAP and 67.6 NDS. These are research benchmark results rather than a guarantee of identical performance for every NVIDIA-hosted configuration.

Throughput depends on factors such as the visual backbone, number of object queries, hardware, and inference configuration. The research materials support the model's focus on efficient temporal processing, but they do not establish one universal frames-per-second figure for the current public endpoint. Any deployment comparison should therefore measure the exact model configuration and hardware being used.

NVIDIA endpoint and pricing

NVIDIA lists StreamPETR as a free endpoint in its NIM catalog. Access requires an NVIDIA API key. The supplied research does not identify a recurring subscription price, per-request charge, or separate output fee for this endpoint.

The documented request is scene-based: a client provides a scene_id and optional configuration, and the service returns annotated camera and bird's-eye-view MP4 videos with inference metadata. The response may be delivered as a ZIP file or through a redirected object-storage response, according to the API reference.

“Free endpoint” should not be interpreted as a complete statement about every possible deployment cost. A hosted request may still require an NVIDIA account and API access, while running or adapting the research implementation locally can involve GPU, storage, engineering, and infrastructure costs. The supplied sources do not provide a separate commercial license schedule for all self-hosted uses.

Main strengths and limitations

Strengths

  • Focused design: StreamPETR is built specifically for multi-view 3D driving perception rather than adapted from a general-purpose vision model.
  • Temporal efficiency: Sparse object-query propagation can reuse information across frames, supporting a streaming approach to detection and tracking.
  • Camera-only operation: The model targets perception from multiple camera views, which is useful for systems centered on visual sensors.
  • Inspectable results: Annotated camera and bird's-eye-view videos make the detections easier to review than raw numerical output alone.
  • Accessible evaluation: NVIDIA provides a free catalog endpoint, allowing users with an API key to try the hosted interface without a separately listed model fee.

Limitations

  • Narrow domain: It is specialized for autonomous-driving scenes and is not suitable as a general image, video, language, or audio model.
  • Scene-based public interface: The documented endpoint does not establish unrestricted arbitrary-video or raw-frame uploads.
  • Undisclosed general limits: No authoritative context length, maximum input duration, maximum output duration, or token-style output limit is supplied.
  • No language or developer reasoning: StreamPETR does not provide conversational reasoning, code generation, function calling, or web search.
  • Configuration-dependent performance: Speed and accuracy vary with the backbone, query count, GPU hardware, and inference settings.
  • Research-to-service distinction: Published nuScenes results describe the research implementation and should not automatically be treated as guaranteed hosted-service results.

Speed, cost, and capability trade-offs

StreamPETR's main trade-off is specialization. Compared with a broad video-understanding model, it offers a more directly relevant architecture and output format for 3D road-scene perception. It can produce detections, tracks, and bird's-eye-view visualizations rather than requiring a general model to infer those structures indirectly from video.

Its cost position is also attractive for experimentation because NVIDIA lists the NIM endpoint as free. However, free access does not make it a universal low-cost replacement for every perception pipeline. The endpoint is limited to its documented scene workflow, and production deployments may require dedicated hardware, integration work, monitoring, and validation.

Speed should be evaluated empirically. The model's temporal-query design is intended to reduce unnecessary repeated processing, but actual throughput depends on the selected backbone and deployment configuration. A smaller or more heavily optimized perception model may be preferable when latency or hardware budget is the overriding requirement; a broader video model may be preferable when the task includes natural-language interpretation or open-ended video questions.

When to choose StreamPETR

Choose StreamPETR when the main task is camera-based autonomous-driving perception and the desired result is structured 3D detection or tracking over a driving scene. It is a particularly reasonable option for:

  • evaluating multi-view 3D object detection on driving data;
  • visualizing detected objects in camera and bird's-eye-view formats;
  • prototyping camera-only perception and temporal tracking workflows;
  • studying object-centric temporal modeling;
  • testing an NVIDIA-hosted NIM endpoint before building a more controlled deployment.

Another type of model may be more appropriate if the requirement is arbitrary user-video analysis, video captioning, question answering, natural-language reasoning, image generation, audio processing, or code generation. StreamPETR is also not the right choice when a system needs a documented general-purpose raw-video upload API, because the supplied endpoint documentation describes a predefined scene-based workflow instead.

Overall assessment

StreamPETR is best understood as a focused autonomous-driving perception model, not as a general AI assistant. Its distinguishing value comes from combining multi-view camera input with temporal object-query propagation for 3D detection and tracking. NVIDIA's free NIM endpoint makes the model practical to evaluate, while its scene-based interface and narrow domain define the boundaries of that convenience.

For readers who need driving-scene detections, object tracks, and bird's-eye-view outputs, StreamPETR is a relevant specialized tool. For broader video understanding or language-based interaction, its lack of text generation, reasoning, tools, and arbitrary-video coverage makes a different model category more suitable.


Answers to Frequently Asked Questions

What are the main limitations of StreamPETR?
StreamPETR is specialized for camera-based autonomous-driving perception and is not a general video, language, audio, image-generation, or conversational model. Its documented public interface is scene-based rather than an unrestricted arbitrary-video upload API, and performance depends on the model configuration and hardware.
Is NVIDIA StreamPETR free to use?
NVIDIA lists StreamPETR as a free NIM catalog endpoint, but access requires an NVIDIA API key. The supplied information does not specify recurring subscription, per-request, or separate output fees; local deployment may still incur GPU, storage, infrastructure, and engineering costs.
What input and output formats does the StreamPETR NVIDIA NIM endpoint support?
The documented endpoint uses a predefined scene-based workflow with a required scene_id and optional configuration. It returns annotated camera-view and bird's-eye-view MP4 videos, along with inference metadata, potentially as a ZIP file or redirected object-storage response.
What is NVIDIA StreamPETR used for?
NVIDIA StreamPETR is a camera-based, multi-view 3D perception model for autonomous driving. It detects and tracks vehicles and other road users, estimates their 3D positions, and produces annotated camera-view and bird's-eye-view outputs.
How does StreamPETR process autonomous-driving video?
StreamPETR propagates sparse object queries between consecutive frames. This object-centric temporal modeling lets it reuse information about detected objects over time instead of rebuilding the entire scene representation for every frame.


Sources 4
Provider

About NVIDIA AI