What Tencent Cloud TokenHub is and when to use it
Tencent Cloud TokenHub is Tencent’s current developer platform for large-model inference. It provides a centralized gateway for Tencent Hunyuan models and selected models from other providers. Instead of integrating with a separate endpoint for every model family, an application can use TokenHub’s standardized API protocols and choose a model available to its account and region.
TokenHub supports both OpenAI-compatible and Anthropic-compatible API patterns. Developers familiar with those ecosystems can usually reuse established clients and tools by changing the API key, base URL and model identifier.
Use TokenHub when you need programmatic access to supported language or multimodal models, streaming responses, function calling, structured JSON output, embeddings or selected search and file workflows. It is primarily an inference gateway. The current documentation does not establish a universal persistent assistant runtime or a universal fine-tuning service.
New projects should use TokenHub rather than the older standalone Hunyuan API. Tencent’s migration documentation states that the legacy Hunyuan platform is scheduled to shut down on September 30, 2026.
Getting access and obtaining an API key
Start in the TokenHub console and create an API key for the region where your application will run. Keys and services are associated with a regional site, and Tencent warns that keys from different sites are not interchangeable.
- Open the TokenHub console.
- Select the required regional site.
- Create an API key.
- Restrict the key to the models or inference services that the application actually needs.
- Check the available models with
/v1/modelsbefore choosing a model for production.
Send the key as a bearer token in the HTTP Authorization header:
Authorization: Bearer YOUR_API_KEYStore the key in an environment variable or secret manager. Do not place it in browser code, a public repository or client-side application code where users can extract it.
Choose a regional endpoint and model
Your base URL depends on the service region. The mainland China endpoint is https://tokenhub.tencentmaas.com/v1. Tencent also documents an international endpoint at https://tokenhub-intl.tencentmaas.com/v1 and a US regional option at https://tokenhub-us.tencentcloudmaas.com/v1 where available.
| Access scope | Base URL | Typical use |
|---|---|---|
| Chinese mainland | https://tokenhub.tencentmaas.com/v1 | Mainland China resource scheduling |
| Global | https://tokenhub-intl.tencentmaas.com/v1 | International access and global resource scheduling |
| United States regional access | https://tokenhub-us.tencentcloudmaas.com/v1 | Documented international or US access where available |
Model identifiers and capabilities are not universal across accounts. Use the model list and model-detail documentation as the source of truth for availability, supported parameters, pricing and features such as vision, video, tools or structured output.
Make a first request with Chat Completions
Chat Completions is the simplest starting point for a message-based application. It accepts a sequence of messages and returns an assistant response. The following request uses the current OpenAI-compatible mainland China endpoint:
curl -X POST 'https://tokenhub.tencentmaas.com/v1/chat/completions'
-H 'Authorization: Bearer ${TOKENHUB_API_KEY}'
-H 'Content-Type: application/json'
--data-raw '{
"model": "hy3-preview",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain API rate limiting in two sentences."}
]
}'Replace hy3-preview with a model currently enabled for your key. The request contains a system message that sets behavior and a user message containing the task. The server returns a JSON response containing one or more choices, including the assistant message.
Understand the response and select an API style
For a normal Chat Completions response, application code generally reads the assistant content from the first choice. In production, also inspect the response status, error object, finish information and usage fields when provided. Do not assume that every model exposes exactly the same optional fields.
TokenHub also provides an OpenAI-compatible Responses endpoint:
POST /v1/responsesChat Completions is usually the broadest compatibility choice for conversational messages, streaming and tool calls. Responses is more appropriate for response-oriented workflows using instructions, function-call items, structured output and response-style event handling. Confirm the selected model’s support before switching interfaces.
How TokenHub pricing generally works
TokenHub does not have one universal price for every model. Usage charges can vary by model and by the resource being consumed. Depending on the model and account configuration, pricing may include input tokens, output tokens, cached tokens, reasoning tokens, image input, image output, search calls or other model-specific units.
Tencent also offers Token Plan arrangements for enterprise usage, including quota pools, API-key management, model access controls and usage monitoring. The actual price and included quota must be checked in the current TokenHub model catalog or account configuration.
Before estimating costs, identify the exact model, region, request type and expected input and output volume. A token is a unit of text processing; longer prompts and larger generated answers generally affect usage, but the applicable billing dimensions are model-specific. Do not use pricing for one model as a proxy for another.
Important capabilities
Streaming responses
TokenHub supports server-sent event streaming. Streaming lets an application display generated text progressively instead of waiting for the complete answer. Set stream to true in a compatible Chat Completions request and read events until the stream ends.
A stream can contain an error after the initial HTTP response has succeeded. Clients must therefore inspect event payloads for error objects rather than treating the HTTP status alone as proof that generation completed successfully.
Function and tool calling
Function calling lets a model request an application-defined operation, such as retrieving weather data or querying an internal system. The model does not execute the function itself. Your application validates the requested arguments, performs the operation, and sends the result back in the format required by the selected API.
In Chat Completions, define functions in the tools array. In the Responses interface, use the corresponding function-call items and return the function output. Treat tool arguments as untrusted input: validate names, types, permissions and side effects before executing anything.
Structured JSON output
Compatible Chat Completions models can support structured output through response_format, including json_object or json_schema. The Responses-compatible interface uses text.format. A JSON schema is useful when application code needs predictable fields rather than free-form prose.
Support varies by model and endpoint. Validate the returned JSON in your application even when a schema was requested, and provide a clear instruction describing the expected data.
Images, video, files and search
Compatible multimodal models can accept image input through supported image URLs or encoded content. Video support varies by model. File input is available in compatible response workflows through file identifiers, URLs or encoded file data, but file search and other built-in tools depend on the model and endpoint.
Selected models also provide web-search capabilities when the required model capability or resource configuration is enabled. Embeddings are available through an OpenAI-compatible embeddings endpoint for supported text and multimodal embedding models.
Use the OpenAI Python SDK
Because TokenHub is OpenAI-compatible, the current OpenAI Python SDK can be configured with a custom base URL. Install the package with pip install openai and set the API key outside the source code:
import os
from openai import OpenAI
api_key = os.environ["TOKENHUB_API_KEY"]
client = OpenAI(
api_key=api_key,
base_url="https://tokenhub.tencentmaas.com/v1",
)
response = client.chat.completions.create(
model="hy3-preview",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Give three practical retry recommendations."},
],
)
print(response.choices[0].message.content)The package and method style above use the current OpenAI Python client generation. The base URL is the important TokenHub-specific setting; the model must still be enabled for the API key.
Streaming and structured output examples
A streaming request returns an iterator of partial chunks:
stream = client.chat.completions.create(
model="hy3-preview",
messages=[
{"role": "user", "content": "Explain server-sent events briefly."}
],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
print()For structured output, request a schema supported by the selected model:
response = client.chat.completions.create(
model="hy3-preview",
messages=[
{"role": "user", "content": "Return a retry policy."}
],
response_format={
"type": "json_schema",
"json_schema": {
"name": "retry_policy",
"strict": True,
"schema": {
"type": "object",
"properties": {
"max_attempts": {"type": "integer"},
"backoff": {"type": "string"}
},
"required": ["max_attempts", "backoff"],
"additionalProperties": False
}
}
}
)
print(response.choices[0].message.content)Check the model documentation before using this pattern. Structured-output support is not necessarily available for every model.
Limits and production considerations
TokenHub has no single public rate limit that applies to every request. Limits depend on the model, region, account, API key and service plan. HTTP 429 indicates rate limiting. When Tencent supplies a Retry-After header, honor it and use exponential backoff for transient failures.
Older Hunyuan documentation described a default limit of five simultaneous conversations. That legacy figure should not be applied automatically to TokenHub. Confirm current quotas for the selected model and plan instead.
- Use
/v1/modelsand model-detail pages to verify availability before deployment. - Keep API keys in a secret manager and restrict them to required models or services.
- Choose a regional endpoint that matches the application’s deployment and data requirements.
- Implement timeouts, retry handling and logging for HTTP errors and stream-level errors.
- Track usage, latency and failures by model, region, request type and API key.
- Test long prompts, large files, multimodal inputs and tool-call failures separately.
- Review privacy, retention, training-use and regional-processing terms for the particular model and contract.
SDKs, console and developer tools
The TokenHub console provides API-key management, model browsing and usage monitoring. Tencent also provides API documentation, API Explorer facilities and model-specific examples. OpenAI-compatible Python and JavaScript SDKs can be used for inference by changing the base URL. Raw HTTPS requests are another straightforward option.
Tencent Cloud also maintains signed control-plane SDKs for languages including Python, Java, PHP, Go, Node.js, .NET, C++, and Ruby. These SDKs are separate from the OpenAI-compatible inference interface. For ordinary model requests, the compatible client or raw HTTP is generally simpler; use Tencent’s signed SDKs when interacting with Tencent Cloud control-plane services.
When TokenHub is a good or poor choice
TokenHub is a good fit when a team wants access to Tencent Hunyuan and other supported models through familiar OpenAI- or Anthropic-compatible protocols, needs regional Tencent Cloud access, or wants centralized model, key and usage management. Its compatibility can reduce the amount of client code required for an initial integration.
It may be a poor fit when an application requires a feature that the selected model or region does not support, a universal persistent assistant runtime, a guaranteed common retention policy across all models, or a single fixed price and rate limit. Capability, billing and data-handling details must be evaluated at the model, region and account level.
For a new deployment, the practical starting point is to select the region, create a restricted TokenHub key, inspect /v1/models, test a small Chat Completions request, and then add streaming, tools or structured output only after confirming support for the chosen model.
