What the NVIDIA API Catalog and NIM API are
The NVIDIA API Catalog is a developer portal for discovering and testing NVIDIA-hosted AI endpoints. At build.nvidia.com, you can select a supported model or NIM microservice, try it in a browser playground, create an API key, and copy request examples.
NVIDIA NIM is the deployment layer behind many of these services. A NIM microservice packages an inference runtime, model-serving components, and an API so that a compatible AI workload can run on NVIDIA GPU infrastructure. You can call hosted endpoints for prototyping or deploy NIM containers yourself when you need more control over networking, data locality, capacity, or customization.
The platform is therefore not one single model API. It is a collection of model- and workload-specific services that commonly use OpenAI-compatible HTTP patterns. Available services can cover language, vision-language, embeddings, retrieval, image generation, speech, text-to-speech, healthcare, robotics, and other specialized workloads.
Who should use it?
NVIDIA’s API is aimed primarily at developers, application teams, and organizations building AI features on NVIDIA-accelerated infrastructure. It is a good fit when you need to evaluate several models, add inference to an application, use multimodal models, or move from a hosted prototype to a private GPU deployment.
It is less suitable if you want a single consumer chatbot, a universal file-analysis service, or a fully managed assistant product with persistent state and built-in web search. NIM supplies model inference; application state, authentication, memory, tool execution, and agent behavior generally remain responsibilities of your application or an agent framework such as NVIDIA NeMo Agent Toolkit.
Getting access and obtaining an API key
- Open the NVIDIA API Catalog at
build.nvidia.com. - Select a model or NIM microservice and review its model-specific documentation.
- Use the browser playground to test an example request where available.
- Join or sign in to the NVIDIA Developer Program if the selected hosted endpoint requires it.
- Create or retrieve an NVIDIA API key.
- Store the key in an environment variable such as
NVIDIA_API_KEY; do not place it in browser code, source control, or public logs.
Hosted access is intended for experimentation and evaluation and is subject to endpoint-, account-, model-, and service-specific limits. The exact model identifier, supported parameters, and commercial terms should be checked on the selected catalog page.
Choosing an API and model
Start with the workload rather than with a generic model name. Use a language model for text generation and chat, a vision-language model when the request includes an image, an embedding service for semantic search, or a specialized image, speech, retrieval, healthcare, or robotics service when that matches your application.
For language and vision-language inference, the main OpenAI-compatible paths include /v1/chat/completions, /v1/completions, and, where supported, /v1/responses. Some NIM versions also expose an Anthropic-compatible /v1/messages endpoint. Support is not identical across models, so confirm the selected model’s API reference before depending on a feature.
The hosted base URL documented for NVIDIA API Catalog integrations is:
https://integrate.api.nvidia.com/v1Making a first request
The following example uses the current OpenAI Python client with NVIDIA’s OpenAI-compatible endpoint. Install the package with pip install openai, set NVIDIA_API_KEY, and replace the model ID if the catalog lists a different current identifier.
import os
from openai import OpenAI
api_key = os.environ.get("NVIDIA_API_KEY")
if not api_key:
raise RuntimeError("Set NVIDIA_API_KEY before running this example")
client = OpenAI(
api_key=api_key,
base_url="https://integrate.api.nvidia.com/v1",
)
response = client.chat.completions.create(
model="meta/llama-3.1-8b-instruct",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain GPU acceleration in one sentence."},
],
max_tokens=128,
temperature=0.2,
)
print(response.choices[0].message.content)The equivalent raw HTTP request is useful when you are not using an SDK:
curl -sS https://integrate.api.nvidia.com/v1/chat/completions
-H "Authorization: Bearer ${NVIDIA_API_KEY}"
-H "Content-Type: application/json"
-d '{
"model": "meta/llama-3.1-8b-instruct",
"messages": [
{"role": "user", "content": "Explain GPU acceleration in one sentence."}
],
"max_tokens": 128,
"temperature": 0.2
}'Understanding the response
A Chat Completions response normally contains a choices array. The generated assistant text is commonly found at choices[0].message.content. Depending on the model and request, the message may instead contain tool calls or other structured fields. Applications should handle error responses, empty content, and model-specific response details rather than assuming every successful response contains plain text.
For multi-turn conversations, send the relevant earlier messages in the messages array. NIM does not automatically provide universal persistent conversation memory; your application must store and manage conversation state when it is needed.
How pricing generally works
NVIDIA Developer Program members can receive free access to hosted NIM API endpoints for prototyping, subject to applicable service limits and terms. Hosted endpoint pricing and quotas can vary by endpoint and are not necessarily uniform across the catalog.
Self-hosted NIM is a different cost model. Production use of downloadable NIM generally requires NVIDIA AI Enterprise licensing. The supplied NVIDIA pricing documentation describes licensing beginning at $4,500 per GPU per year, or approximately $1 per GPU per hour in cloud environments. Treat these as documented starting figures rather than a universal price for every deployment; licensing, infrastructure, cloud GPU charges, support, and capacity planning can affect the total cost.
Core capabilities
- Streaming: Applicable hosted and self-hosted models can stream generated output, commonly by setting
stream: true. This lets an interface display partial output before the full response is complete. - Multi-turn chat: Send previous messages with the next request. Persistent storage and conversation management are application responsibilities.
- Tool and function calling: Compatible models can return structured requests for application-defined functions using OpenAI-style
toolsandtool_choice. Your application must execute the function and send the result back; the model does not automatically perform arbitrary external actions. - Structured output: Compatible NIM deployments can support guided JSON, JSON Schema, grammars, regular expressions, or constrained choices. Verify the exact method supported by the selected model and NIM version.
- Image input: Selected vision-language models accept image content or image URLs according to their individual request format.
- Image generation and other media: NVIDIA provides specialized NIM services for supported image, audio, speech, video, and other workloads. These are not guaranteed capabilities of every language endpoint.
- Embeddings and retrieval: Catalog services can support embedding and retrieval workflows, but the request format and deployment requirements depend on the selected service.
Streaming, tools, and structured data
A streaming request uses the same general endpoint but sets stream to true:
stream = client.chat.completions.create(
model="meta/llama-3.1-8b-instruct",
messages=[
{"role": "user", "content": "Explain CUDA in three short paragraphs."},
],
max_tokens=256,
stream=True,
)
for chunk in stream:
text = chunk.choices[0].delta.content
if text:
print(text, end="", flush=True)For tool calling, define the function schema in the request. A compatible model may return a tool call containing a function name and JSON arguments. Your server should validate those arguments, apply authorization and business rules, execute the function, and then continue the conversation with the tool result.
Structured output is useful when downstream code needs predictable JSON rather than prose. However, support varies by model and runtime. Do not assume that a structured-output option available in one NIM deployment is available for every model in the catalog.
Files and file handling
File upload is not a universal feature of NVIDIA’s core chat inference API. Ordinary chat requests generally use text, image URLs, inline content parts, or encoded media according to the selected model’s interface.
Some specialized NVIDIA services may accept uploaded content or provide their own media-input mechanism. If your application needs PDFs, spreadsheets, or other documents, you may need to extract and chunk the content in your own system, use a retrieval pipeline, or select a specialized service. Confirm the exact documentation before sending files or assuming that a file-processing workflow is built in.
SDKs, playground, and developer tools
The easiest starting point is the API Catalog playground at build.nvidia.com. It provides model discovery, interactive testing, API-key access, generated examples, and links to model-specific references.
NVIDIA’s hosted endpoints use OpenAI-compatible HTTP semantics. The official quickstarts prominently document Python, and compatible OpenAI client libraries can be configured with NVIDIA’s base URL. JavaScript, TypeScript, PHP, and other languages can call the HTTPS API directly. SDK compatibility does not guarantee that every model-specific parameter or feature is supported, so use the selected NIM reference as the authority.
When to use self-hosted NIM
Self-hosted NIM is appropriate when an organization needs private networking, control over inference data, predictable GPU capacity, custom deployment operations, or supported LoRA and PEFT adapter serving. NIM services normally expose API paths on a local or private service URL, commonly using port 8000, along with operational endpoints for models, health, metadata, version information, manifests, licenses, and metrics.
Self-hosting transfers more operational responsibility to your team. You must provide compatible NVIDIA GPU infrastructure, deploy and monitor containers, manage capacity and scaling, secure the endpoint, and understand the license for the selected service.
Limits and production considerations
Hosted preview endpoints have service-specific limits. Actual latency and throughput depend on the model, GPU, input and output sizes, concurrency, batching, region, queueing, and whether the service is hosted or self-managed. Do not promise a particular response time or request quota without checking the applicable endpoint documentation.
Before production use, test the exact model and NIM version for:
- maximum context and output sizes;
- tool-calling behavior and argument validation;
- structured-output support;
- image or other media formats;
- streaming behavior and error handling;
- parallel requests, throughput, and concurrency;
- LoRA or adapter support;
- commercial licensing, rate limits, support, and availability.
Hosted API trial terms and individual model disclosures may address how inputs and outputs are recorded, used, or retained. Some catalog experiences may record content to provide the trial, improve products and models, or monitor security, fraud, and abuse. Avoid confidential, regulated, or personal data in trial endpoints unless the applicable terms clearly permit it. With self-hosted NIM, retention and logging are mainly determined by your infrastructure, application, and operational policies, although licensing and enterprise terms still apply.
Advantages and limitations
NVIDIA’s main advantage is the path from an interactive hosted prototype to a deployable GPU-backed inference service. The API Catalog reduces initial setup, while NIM offers more control over deployment, data locality, capacity, and selected customization workflows. OpenAI-compatible request patterns also make it easier to adapt existing application code.
The main limitation is variation. NIM capabilities differ substantially by model, service, and version. The platform is not a single uniform API with identical behavior everywhere. File uploads, structured outputs, tool calling, image input, LoRA serving, and other features must be verified individually.
NVIDIA API Catalog and NIM are a strong choice for teams building on NVIDIA infrastructure, evaluating multiple AI workloads, or planning private inference. They are a poorer choice for someone seeking a turnkey consumer assistant, universal document upload, built-in web search, or fully managed persistent agent state.
