Developer platform

NVIDIA API Catalog and NIM

Developer overview for NVIDIA AI, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API NVIDIA NIM OpenAI-compatible inference API
SDK support Official Python quickstarts and broad raw HTTP support. OpenAI-compatible Python and JavaScript clients can be configured with NVIDIA’s base URL. NVIDIA also provides first-party SDKs and toolkits across the broader CUDA, NeMo, NIM, Triton, and AI Enter
Rate limits Rate limits are endpoint-, account-, model-, and service-dependent. Hosted preview endpoints operate within NVIDIA Developer Program and service-specific limits. Self-hosted limits are controlled by deployment capacity, runtime configuration, and infrastr
Platform

API overview

Endpoints

API access

Base URL https://integrate.api.nvidia.com/v1
Primary API NVIDIA NIM OpenAI-compatible inference API
Pricing

API pricing

Pricing model Hosted API Catalog endpoints are free for NVIDIA Developer Program prototyping within applicable limits; production self-hosted NIM generally requires NVIDIA AI Enterprise licensing priced by GPU.

NVIDIA Developer Program access is free for prototyping. NVIDIA AI Enterprise production licensing starts at $4,500 per GPU per year or approximately $1 per GPU per hour in cloud environments. Hosted endpoint pricing and limits can vary by endpoint.

Developer experience

SDKs & usability

SDKs Official Python quickstarts and broad raw HTTP support. OpenAI-compatible Python and JavaScript clients can be configured with NVIDIA’s base URL. NVIDIA also provides first-party SDKs and toolkits across the broader CUDA, NeMo, NIM, Triton, and AI Enter
Ease of use Easy for prototyping through build.nvidia.com, API-key generation, browser previews, generated code, and OpenAI-compatible HTTP patterns. Self-hosting requires NVIDIA GPU infrastructure, containers, model selection, and operational deployment work.
Documentation Strong and extensive, with centralized API documentation, model-specific references, NIM deployment guides, support matrices, quickstarts, release notes, and NeMo ecosystem documentation. Details can vary by NIM version and model.
Latency NVIDIA positions NIM for low-latency, high-throughput inference on NVIDIA-accelerated infrastructure. Actual latency depends on model, GPU, input and output token counts, concurrency, batching, region, queueing, and whether the endpoint is hosted or self-
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ Fine-tuning
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

NVIDIA’s developer platform includes hosted API Catalog endpoints and downloadable NIM microservices. Hosted endpoints are accessed through build.nvidia.com and the integrate.api.nvidia.com/v1 base URL. Current NIM API references document chat completions, completions, Responses API, Anthropic-compatible Messages, model listing, health, metadata, metrics, and management endpoints. Streaming is supported by applicable endpoints. Tool calling is supported by compatible models and may require model-specific runtime configuration for self-hosted NIM. Structured outputs are supported by compatible NIM deployments using guided JSON, JSON schema, grammars, regular expressions, or constrained choices. Image input is supported by selected vision-language models. File upload is not a universal capability of the core chat API; specialized services may provide their own upload or media-input mechanisms. NVIDIA supports LoRA and PEFT adapter serving in self-hosted NIM, while hosted API Catalog access should not be assumed to provide universal fine-tuning. NVIDIA’s NeMo Agent Toolkit and related NeMo platform components provide first-party agent workflow, runtime, deployment, evaluation, memory, observability, REST, MCP, and tool-integration capabilities, but NVIDIA does not primarily expose these as a universal hosted Assistants object API.

Examples

API examples

curl -sS https://integrate.api.nvidia.com/v1/chat/completions \
  -H "Authorization: Bearer ${NVIDIA_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta/llama-3.1-8b-instruct",
    "messages": [
      {"role": "system", "content": "You are a concise technical assistant."},
      {"role": "user", "content": "Explain GPU acceleration in one sentence."}
    ],
    "max_tokens": 128,
    "temperature": 0.2
  }' | jq -r '.choices[0].message.content'

curl -N -sS https://integrate.api.nvidia.com/v1/chat/completions \
  -H "Authorization: Bearer ${NVIDIA_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "meta/llama-3.1-8b-instruct",
    "messages": [{"role": "user", "content": "Stream a short explanation of CUDA."}],
    "max_tokens": 256,
    "stream": true
  }'
# Install: pip install openai
import json
import os
from openai import OpenAI

api_key = os.environ.get("NVIDIA_API_KEY")
if not api_key:
    raise RuntimeError("Set NVIDIA_API_KEY before running this example")

client = OpenAI(
    api_key=api_key,
    base_url="https://integrate.api.nvidia.com/v1",
)

try:
    response = client.chat.completions.create(
        model="meta/llama-3.1-8b-instruct",
        messages=[
            {"role": "system", "content": "You are a concise technical assistant."},
            {"role": "user", "content": "What is GPU acceleration?"},
        ],
        max_tokens=128,
        temperature=0.2,
    )
    print(response.choices[0].message.content)

    stream = client.chat.completions.create(
        model="meta/llama-3.1-8b-instruct",
        messages=[
            {"role": "user", "content": "Explain CUDA in three short paragraphs."},
        ],
        max_tokens=256,
        stream=True,
    )
    for chunk in stream:
        text = chunk.choices[0].delta.content
        if text:
            print(text, end="", flush=True)
    print()

    tools = [
        {
            "type": "function",
            "function": {
                "name": "get_weather",
                "description": "Get the current weather for a city",
                "parameters": {
                    "type": "object",
                    "properties": {"city": {"type": "string"}},
                    "required": ["city"],
                },
            },
        }
    ]
    tool_response = client.chat.completions.create(
        model="meta/llama-3.1-8b-instruct",
        messages=[{"role": "user", "content": "What is the weather in Austin?"}],
        tools=tools,
        tool_choice="auto",
        max_tokens=256,
    )
    message = tool_response.choices[0].message
    if message.tool_calls:
        for call in message.tool_calls:
            print(json.dumps({"name": call.function.name, "arguments": call.function.arguments}))
    else:
        print(message.content)
except Exception as exc:
    print(f"NVIDIA API request failed: {exc}")
    raise
// Install: npm install openai
import OpenAI from "openai";

const apiKey = process.env.NVIDIA_API_KEY;
if (!apiKey) {
  throw new Error("Set NVIDIA_API_KEY before running this example");
}

const client = new OpenAI({
  apiKey,
  baseURL: "https://integrate.api.nvidia.com/v1",
});

try {
  const response = await client.chat.completions.create({
    model: "meta/llama-3.1-8b-instruct",
    messages: [
      { role: "system", content: "You are a concise technical assistant." },
      { role: "user", content: "Explain GPU acceleration in one sentence." },
    ],
    max_tokens: 128,
    temperature: 0.2,
  });
  console.log(response.choices[0].message.content);

  const stream = await client.chat.completions.create({
    model: "meta/llama-3.1-8b-instruct",
    messages: [{ role: "user", content: "Explain CUDA briefly." }],
    max_tokens: 256,
    stream: true,
  });
  for await (const chunk of stream) {
    const text = chunk.choices[0]?.delta?.content;
    if (text) process.stdout.write(text);
  }
  process.stdout.write("\n");

  const toolResponse = await client.chat.completions.create({
    model: "meta/llama-3.1-8b-instruct",
    messages: [{ role: "user", content: "What is the weather in Austin?" }],
    tools: [
      {
        type: "function",
        function: {
          name: "get_weather",
          description: "Get the current weather for a city",
          parameters: {
            type: "object",
            properties: { city: { type: "string" } },
            required: ["city"],
          },
        },
      },
    ],
    tool_choice: "auto",
    max_tokens: 256,
  });
  console.log(JSON.stringify(toolResponse.choices[0].message));
} catch (error) {
  console.error("NVIDIA API request failed:", error);
  process.exitCode = 1;
}
<?php
$apiKey = getenv('NVIDIA_API_KEY');
if (!$apiKey) {
    throw new RuntimeException('Set NVIDIA_API_KEY before running this example');
}

$url = 'https://integrate.api.nvidia.com/v1/chat/completions';
$payload = [
    'model' => 'meta/llama-3.1-8b-instruct',
    'messages' => [
        ['role' => 'system', 'content' => 'You are a concise technical assistant.'],
        ['role' => 'user', 'content' => 'Explain GPU acceleration in one sentence.'],
    ],
    'max_tokens' => 128,
    'temperature' => 0.2,
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 60,
]);

$responseBody = curl_exec($ch);
if ($responseBody === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('cURL error: ' . $error);
}

$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($responseBody, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? $responseBody;
    throw new RuntimeException('NVIDIA API error (' . $status . '): ' . $message);
}

echo $data['choices'][0]['message']['content'] . PHP_EOL;

$toolPayload = [
    'model' => 'meta/llama-3.1-8b-instruct',
    'messages' => [['role' => 'user', 'content' => 'What is the weather in Austin?']],
    'tools' => [[
        'type' => 'function',
        'function' => [
            'name' => 'get_weather',
            'description' => 'Get the current weather for a city',
            'parameters' => [
                'type' => 'object',
                'properties' => ['city' => ['type' => 'string']],
                'required' => ['city'],
            ],
        ],
    ]],
    'tool_choice' => 'auto',
    'max_tokens' => 256,
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
    ],
    CURLOPT_POSTFIELDS => json_encode($toolPayload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 60,
]);

$toolBody = curl_exec($ch);
if ($toolBody === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('cURL error: ' . $error);
}

$toolStatus = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
$toolData = json_decode($toolBody, true, 512, JSON_THROW_ON_ERROR);

if ($toolStatus < 200 || $toolStatus >= 300) {
    $message = $toolData['error']['message'] ?? $toolBody;
    throw new RuntimeException('NVIDIA API error (' . $toolStatus . '): ' . $message);
}

echo json_encode($toolData['choices'][0]['message'], JSON_PRETTY_PRINT) . PHP_EOL;
?>
Policies

Data & usage

Data training

NVIDIA API Trial Terms state that user content and generated content are generally used during the session to provide the API service and are not stored or used after the session unless a specific service or catalog disclosure says otherwise. However, individual API Catalog model experience pages may state that inputs and outputs are recorded to provide the trial experience, improve NVIDIA products and AI models, and monitor security, fraud, or abuse. Developers should review the applicable model disclosure and avoid confidential or regulated data in trial endpoints. No general claim that all hosted API traffic is excluded from product improvement should be made.

Data retention

General API Trial Terms state that user and generated content is not stored after the session except for disclosed services or purposes. Certain services, including fine-tuning-related services, may retain user content for 30 days after first upload and generated fine-tuning content for 90 days after availability unless different subscription terms apply. Security, fraud, and abuse monitoring may involve logging or storage. Self-hosted NIM retention is primarily controlled by the customer’s infrastructure, logging, and operational policies.

Rate limits

Rate limits are endpoint-, account-, model-, and service-dependent. Hosted preview endpoints operate within NVIDIA Developer Program and service-specific limits. Self-hosted limits are controlled by deployment capacity, runtime configuration, and infrastr

Developer guide

NVIDIA API Catalog and NIM: A Beginner’s Guide to the Developer API

NVIDIA’s developer API platform combines hosted inference endpoints in the NVIDIA API Catalog with downloadable NVIDIA NIM microservices. Developers can test models through build.nvidia.com, authenticate with an NVIDIA API key, use OpenAI-compatible request formats, and later deploy NIM on supported NVIDIA GPU infrastructure for greater control over data, capacity, and customization.
NVIDIA API Catalog and NVIDIA NIM provide access to AI inference without requiring every developer to build a model-serving system from scratch. The hosted catalog is useful for experimenting with supported language, vision-language, embedding, image, speech, retrieval, and other services. NIM extends the same general approach to self-hosted deployments, where organizations can run compatible microservices on their own NVIDIA GPU infrastructure. This guide explains how access works, how to make a first request, which features are available, and what to check before using the platform in production.
NVIDIA API Catalog provides hosted access to selected NIM inference services through build.nvidia.com and an OpenAI-compatible API. Developers can use API keys, streaming, tool calling, structured outputs, selected multimodal models, and SDK-compatible HTTP patterns. Self-hosted NIM adds control over GPU infrastructure, data locality, capacity, and supported customization, but requires more operational work and generally NVIDIA AI Enterprise licensing.

What the NVIDIA API Catalog and NIM API are

The NVIDIA API Catalog is a developer portal for discovering and testing NVIDIA-hosted AI endpoints. At build.nvidia.com, you can select a supported model or NIM microservice, try it in a browser playground, create an API key, and copy request examples.

NVIDIA NIM is the deployment layer behind many of these services. A NIM microservice packages an inference runtime, model-serving components, and an API so that a compatible AI workload can run on NVIDIA GPU infrastructure. You can call hosted endpoints for prototyping or deploy NIM containers yourself when you need more control over networking, data locality, capacity, or customization.

The platform is therefore not one single model API. It is a collection of model- and workload-specific services that commonly use OpenAI-compatible HTTP patterns. Available services can cover language, vision-language, embeddings, retrieval, image generation, speech, text-to-speech, healthcare, robotics, and other specialized workloads.

Who should use it?

NVIDIA’s API is aimed primarily at developers, application teams, and organizations building AI features on NVIDIA-accelerated infrastructure. It is a good fit when you need to evaluate several models, add inference to an application, use multimodal models, or move from a hosted prototype to a private GPU deployment.

It is less suitable if you want a single consumer chatbot, a universal file-analysis service, or a fully managed assistant product with persistent state and built-in web search. NIM supplies model inference; application state, authentication, memory, tool execution, and agent behavior generally remain responsibilities of your application or an agent framework such as NVIDIA NeMo Agent Toolkit.

Getting access and obtaining an API key

  1. Open the NVIDIA API Catalog at build.nvidia.com.
  2. Select a model or NIM microservice and review its model-specific documentation.
  3. Use the browser playground to test an example request where available.
  4. Join or sign in to the NVIDIA Developer Program if the selected hosted endpoint requires it.
  5. Create or retrieve an NVIDIA API key.
  6. Store the key in an environment variable such as NVIDIA_API_KEY; do not place it in browser code, source control, or public logs.

Hosted access is intended for experimentation and evaluation and is subject to endpoint-, account-, model-, and service-specific limits. The exact model identifier, supported parameters, and commercial terms should be checked on the selected catalog page.

Choosing an API and model

Start with the workload rather than with a generic model name. Use a language model for text generation and chat, a vision-language model when the request includes an image, an embedding service for semantic search, or a specialized image, speech, retrieval, healthcare, or robotics service when that matches your application.

For language and vision-language inference, the main OpenAI-compatible paths include /v1/chat/completions, /v1/completions, and, where supported, /v1/responses. Some NIM versions also expose an Anthropic-compatible /v1/messages endpoint. Support is not identical across models, so confirm the selected model’s API reference before depending on a feature.

The hosted base URL documented for NVIDIA API Catalog integrations is:

https://integrate.api.nvidia.com/v1

Making a first request

The following example uses the current OpenAI Python client with NVIDIA’s OpenAI-compatible endpoint. Install the package with pip install openai, set NVIDIA_API_KEY, and replace the model ID if the catalog lists a different current identifier.

import os
from openai import OpenAI

api_key = os.environ.get("NVIDIA_API_KEY")
if not api_key:
    raise RuntimeError("Set NVIDIA_API_KEY before running this example")

client = OpenAI(
    api_key=api_key,
    base_url="https://integrate.api.nvidia.com/v1",
)

response = client.chat.completions.create(
    model="meta/llama-3.1-8b-instruct",
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "Explain GPU acceleration in one sentence."},
    ],
    max_tokens=128,
    temperature=0.2,
)

print(response.choices[0].message.content)

The equivalent raw HTTP request is useful when you are not using an SDK:

curl -sS https://integrate.api.nvidia.com/v1/chat/completions 
  -H "Authorization: Bearer ${NVIDIA_API_KEY}" 
  -H "Content-Type: application/json" 
  -d '{
    "model": "meta/llama-3.1-8b-instruct",
    "messages": [
      {"role": "user", "content": "Explain GPU acceleration in one sentence."}
    ],
    "max_tokens": 128,
    "temperature": 0.2
  }'

Understanding the response

A Chat Completions response normally contains a choices array. The generated assistant text is commonly found at choices[0].message.content. Depending on the model and request, the message may instead contain tool calls or other structured fields. Applications should handle error responses, empty content, and model-specific response details rather than assuming every successful response contains plain text.

For multi-turn conversations, send the relevant earlier messages in the messages array. NIM does not automatically provide universal persistent conversation memory; your application must store and manage conversation state when it is needed.

How pricing generally works

NVIDIA Developer Program members can receive free access to hosted NIM API endpoints for prototyping, subject to applicable service limits and terms. Hosted endpoint pricing and quotas can vary by endpoint and are not necessarily uniform across the catalog.

Self-hosted NIM is a different cost model. Production use of downloadable NIM generally requires NVIDIA AI Enterprise licensing. The supplied NVIDIA pricing documentation describes licensing beginning at $4,500 per GPU per year, or approximately $1 per GPU per hour in cloud environments. Treat these as documented starting figures rather than a universal price for every deployment; licensing, infrastructure, cloud GPU charges, support, and capacity planning can affect the total cost.

Core capabilities

  • Streaming: Applicable hosted and self-hosted models can stream generated output, commonly by setting stream: true. This lets an interface display partial output before the full response is complete.
  • Multi-turn chat: Send previous messages with the next request. Persistent storage and conversation management are application responsibilities.
  • Tool and function calling: Compatible models can return structured requests for application-defined functions using OpenAI-style tools and tool_choice. Your application must execute the function and send the result back; the model does not automatically perform arbitrary external actions.
  • Structured output: Compatible NIM deployments can support guided JSON, JSON Schema, grammars, regular expressions, or constrained choices. Verify the exact method supported by the selected model and NIM version.
  • Image input: Selected vision-language models accept image content or image URLs according to their individual request format.
  • Image generation and other media: NVIDIA provides specialized NIM services for supported image, audio, speech, video, and other workloads. These are not guaranteed capabilities of every language endpoint.
  • Embeddings and retrieval: Catalog services can support embedding and retrieval workflows, but the request format and deployment requirements depend on the selected service.

Streaming, tools, and structured data

A streaming request uses the same general endpoint but sets stream to true:

stream = client.chat.completions.create(
    model="meta/llama-3.1-8b-instruct",
    messages=[
        {"role": "user", "content": "Explain CUDA in three short paragraphs."},
    ],
    max_tokens=256,
    stream=True,
)

for chunk in stream:
    text = chunk.choices[0].delta.content
    if text:
        print(text, end="", flush=True)

For tool calling, define the function schema in the request. A compatible model may return a tool call containing a function name and JSON arguments. Your server should validate those arguments, apply authorization and business rules, execute the function, and then continue the conversation with the tool result.

Structured output is useful when downstream code needs predictable JSON rather than prose. However, support varies by model and runtime. Do not assume that a structured-output option available in one NIM deployment is available for every model in the catalog.

Files and file handling

File upload is not a universal feature of NVIDIA’s core chat inference API. Ordinary chat requests generally use text, image URLs, inline content parts, or encoded media according to the selected model’s interface.

Some specialized NVIDIA services may accept uploaded content or provide their own media-input mechanism. If your application needs PDFs, spreadsheets, or other documents, you may need to extract and chunk the content in your own system, use a retrieval pipeline, or select a specialized service. Confirm the exact documentation before sending files or assuming that a file-processing workflow is built in.

SDKs, playground, and developer tools

The easiest starting point is the API Catalog playground at build.nvidia.com. It provides model discovery, interactive testing, API-key access, generated examples, and links to model-specific references.

NVIDIA’s hosted endpoints use OpenAI-compatible HTTP semantics. The official quickstarts prominently document Python, and compatible OpenAI client libraries can be configured with NVIDIA’s base URL. JavaScript, TypeScript, PHP, and other languages can call the HTTPS API directly. SDK compatibility does not guarantee that every model-specific parameter or feature is supported, so use the selected NIM reference as the authority.

When to use self-hosted NIM

Self-hosted NIM is appropriate when an organization needs private networking, control over inference data, predictable GPU capacity, custom deployment operations, or supported LoRA and PEFT adapter serving. NIM services normally expose API paths on a local or private service URL, commonly using port 8000, along with operational endpoints for models, health, metadata, version information, manifests, licenses, and metrics.

Self-hosting transfers more operational responsibility to your team. You must provide compatible NVIDIA GPU infrastructure, deploy and monitor containers, manage capacity and scaling, secure the endpoint, and understand the license for the selected service.

Limits and production considerations

Hosted preview endpoints have service-specific limits. Actual latency and throughput depend on the model, GPU, input and output sizes, concurrency, batching, region, queueing, and whether the service is hosted or self-managed. Do not promise a particular response time or request quota without checking the applicable endpoint documentation.

Before production use, test the exact model and NIM version for:

  • maximum context and output sizes;
  • tool-calling behavior and argument validation;
  • structured-output support;
  • image or other media formats;
  • streaming behavior and error handling;
  • parallel requests, throughput, and concurrency;
  • LoRA or adapter support;
  • commercial licensing, rate limits, support, and availability.

Hosted API trial terms and individual model disclosures may address how inputs and outputs are recorded, used, or retained. Some catalog experiences may record content to provide the trial, improve products and models, or monitor security, fraud, and abuse. Avoid confidential, regulated, or personal data in trial endpoints unless the applicable terms clearly permit it. With self-hosted NIM, retention and logging are mainly determined by your infrastructure, application, and operational policies, although licensing and enterprise terms still apply.

Advantages and limitations

NVIDIA’s main advantage is the path from an interactive hosted prototype to a deployable GPU-backed inference service. The API Catalog reduces initial setup, while NIM offers more control over deployment, data locality, capacity, and selected customization workflows. OpenAI-compatible request patterns also make it easier to adapt existing application code.

The main limitation is variation. NIM capabilities differ substantially by model, service, and version. The platform is not a single uniform API with identical behavior everywhere. File uploads, structured outputs, tool calling, image input, LoRA serving, and other features must be verified individually.

NVIDIA API Catalog and NIM are a strong choice for teams building on NVIDIA infrastructure, evaluating multiple AI workloads, or planning private inference. They are a poorer choice for someone seeking a turnkey consumer assistant, universal document upload, built-in web search, or fully managed persistent agent state.

Sources 12