Hy-World

Hy-World-2.1-scene

by Tencent AI · Current and accessible through Tencent Cloud TokenHub

Tencent Hy-World-2.1-scene creates navigable 3D environments from text prompts or reference images through an asynchronous TokenHub workflow. It can return Gaussian-splat scenes, point clouds, collision meshes, previews, and spatial metadata, with a documented reference cost of about ¥80 per scene.

Image generation Reasoning Coding
Tencent Hy-World-2.1-scene is designed for generating explorable 3D environments rather than ordinary 2D images, videos, or text. A request can start with a text description or one reference image, and the asynchronous service produces scene assets suitable for interactive visualization, prototyping, games, virtual production, and related 3D workflows.
Outputs

What Hy-World-2.1-scene can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
2/10 Speed
3/10 Cost efficiency
Specifications

Technical details

Model family Hy-World
Model type Other
Status Current and accessible through Tencent Cloud TokenHub
Knowledge cutoff notes

Tencent's public service documentation describes the model's API and output behavior but does not specify a knowledge cutoff for this specialized 3D generation service.

Model notes

The canonical API model identifier is hy-world2-scene. The model accepts either a text prompt or one publicly accessible reference image per request, not both simultaneously. Text prompts are limited to 600 characters. Image inputs must be at least 512×512 pixels, no larger than 10 MB, and use JPG, JPEG, PNG, or WEBP. Generation is asynchronous and commonly takes approximately 10–20 minutes. Successful results may include an SPZ 3DGS scene, PLY point cloud, complete collision mesh, simplified collision mesh, preview images, spatial metadata, and token-usage data. The documented reference cost is approximately ¥80 per scene, but actual billing depends on recorded token usage and configuration. Editorial scores are not vendor benchmarks; reasoning and coding are not meaningful primary use cases for this specialized generative model.

Cost

Model pricing

Input ¥10 per 1 million tokens; reference usage is 8,000,000 tokens per generated scene
Output Approximately ¥80 per generated 3D scene at the documented reference usage
Model guide

Hy-World-2.1-scene: Tencent’s Text-to-3D Model for Navigable Worlds

Hy-World-2.1-scene is Tencent’s hosted 3D world-generation model for creating navigable environments from either a text prompt or a reference image. Through Tencent Cloud TokenHub, it can return Gaussian-splat scenes, point clouds, collision meshes, preview images, and spatial metadata.

What is Hy-World-2.1-scene?

Hy-World-2.1-scene is Tencent’s specialized world-generation model for producing interactive 3D environments. It is available through Tencent Cloud TokenHub under the API model identifier hy-world2-scene. Instead of returning a conventional 2D picture, the service generates assets intended to represent a navigable scene that can be explored, rendered, and used in downstream 3D workflows.

The model is part of Tencent’s current TokenHub catalog of 3D-generation services. Its primary output is a scene that can combine 3D Gaussian-splat data, point-cloud information, meshes, preview images, and spatial metadata. Tencent describes the result as supporting free movement and physical collision, which makes the service more relevant to interactive environments than to static image generation.

For a beginner, the practical distinction is simple: a text-to-image model generally produces one flat image, while Hy-World-2.1-scene is intended to provide the underlying representation of a space. That makes it potentially useful when the environment needs to be viewed from different positions or incorporated into an interactive prototype.

How inputs and generation work

Each request uses one supported input route: either a text prompt or a reference image. Text and image inputs cannot be submitted together in the same request. Text prompts can contain up to 600 characters, so the input is best treated as a concise description of the desired environment rather than a long design document.

When using an image, the image must be publicly accessible. Tencent’s documented requirements specify a minimum size of 512 × 512 pixels, a maximum file size of 10 MB, and support for JPG, JPEG, PNG, and WEBP formats. These restrictions matter in production workflows because an otherwise suitable local file may need to be uploaded to an accessible location before the request can be made.

Generation is asynchronous. A successful submission returns a task identifier rather than immediately returning a finished scene. An application must poll the task-status endpoint until the task reaches a terminal state. Tencent documentation indicates that generation commonly takes approximately 10 to 20 minutes. This makes the model more appropriate for queued or batch-style creation workflows than for an interactive application that expects an immediate response after every prompt.

What the model outputs

A completed task can provide several kinds of scene data. The exact result should be treated as a collection of assets rather than a single image file.

  • SPZ 3D Gaussian-splat data: a compact scene representation based on Gaussian splats, which can be rendered to show the generated environment from different viewpoints.
  • PLY point cloud data: a point-based representation that can be used in compatible 3D and visualization tools.
  • Complete collision mesh: geometry intended to support collision-aware interaction with the scene.
  • Simplified collision mesh: a lower-complexity collision representation for workflows where efficient interaction is more important than full geometric detail.
  • Preview images: rendered views that help users inspect the result without immediately loading the full scene assets.
  • Super-resolution preview images: higher-resolution preview output when provided by the completed task.
  • Spatial-position metadata: information describing the scene’s spatial arrangement or positioning.
  • Token-usage information: usage data relevant to billing and accounting.

These outputs give Hy-World-2.1-scene a different profile from a model that only returns a rendered image. However, the supplied documentation does not specify a universal maximum scene size, maximum output token count, or fixed geometry budget. Those limits should therefore not be assumed when planning a pipeline.

Supported modalities and capabilities

AreaDocumented position
Text inputYes; prompts are limited to 600 characters
Image inputYes; one publicly accessible image per request
Audio inputNo documented support
Video inputNo documented support
3D outputYes; scene, Gaussian-splat, point-cloud, and mesh assets
Text outputNot the model’s primary output
Tool or function callingNo documented support
StreamingNo documented support; generation is asynchronous
Structured JSON outputNo documented model capability

The model’s multimodal behavior should be understood as text-to-3D or image-to-3D generation. It accepts more than one input type, but it does not combine text and image inputs in a single request according to the supplied documentation. Its output is also not ordinary conversational text: the useful result is a set of visual and geometric scene assets.

Pricing and cost trade-offs

Tencent’s documented reference pricing is ¥10 per one million tokens. The reference usage for one generated scene is 8,000,000 tokens, giving an approximate cost of ¥80 per scene at that usage level. This is a reference calculation rather than a guaranteed flat price: actual charges depend on recorded token usage and configuration.

The price places the service in a different category from ordinary 2D image generation. A generated scene may include several representations and supporting assets, so the relevant comparison is not simply the cost of one preview image. Organizations should also account for storage, processing, asset conversion, rendering, and any application infrastructure needed to use the returned scene.

Because generation commonly takes 10 to 20 minutes, Hy-World-2.1-scene also trades speed for richer output. It can be a reasonable choice when producing a complete environment is more important than immediate feedback. It is less suitable when a user needs many rapid visual variations during a live interaction.

Main strengths

  • Scene-oriented output: the service is built around navigable 3D environments rather than isolated 2D views.
  • Two practical starting points: users can describe a scene with text or provide a visual reference image.
  • Multiple downstream assets: SPZ Gaussian-splat data, PLY point clouds, collision meshes, previews, and spatial metadata can support different stages of a 3D workflow.
  • Collision-aware potential: complete and simplified collision meshes are useful when a scene must support movement or physical interaction.
  • Useful for rapid environment prototyping: teams can create a 3D starting point without manually building every element before testing an idea.

These strengths are based on the documented output and workflow. They should not be interpreted as a guarantee that every generated scene will be production-ready without cleanup, optimization, or integration work.

Limitations to plan for

The first operational limitation is latency. The asynchronous workflow and typical 10-to-20-minute generation time make the service unsuitable for low-latency inference. A production application needs task tracking, status polling, failure handling, and a way to notify users when assets are ready.

The input rules are also restrictive. Text is limited to 600 characters, and text and image guidance cannot be combined in one request. Image inputs must meet the documented accessibility, dimension, format, and file-size requirements. Users who need detailed art direction may have to encode the most important requirements into a short prompt or prepare a carefully selected reference image.

The service is hosted through Tencent Cloud TokenHub rather than described in the supplied research as a downloadable standalone model. That means deployment depends on the hosted service and its API workflow. The documentation supplied here also does not provide a context window, maximum output-token value, release date, or fixed maximum scene complexity.

Finally, 3D scene generation is not the same as delivering a finished game level or polished digital twin. Even when the output includes meshes and collision assets, a downstream team may still need to inspect geometry, optimize files, configure materials, integrate rendering, and adapt the result to its target engine.

Reasoning, coding, and tool-use expectations

Hy-World-2.1-scene is not a general-purpose language model. It is not designed to answer questions, write application code, perform extended reasoning, or call external tools as part of a conversational workflow. The supplied editorial assessment assigns low practical relevance to reasoning and coding because those are not the model’s primary tasks; these are editorial evaluations, not vendor-published benchmark scores.

Likewise, no tool or function-calling capability is documented. The API workflow itself still requires ordinary application logic for submitting jobs, polling task status, downloading returned assets, and handling errors, but that orchestration is performed by the integrating application rather than by the model through a documented tool-use interface.

Best use cases

  • Interactive environment prototypes: create an explorable starting point for a game, simulation, or virtual-world concept.
  • Virtual production: explore environment ideas before committing to detailed manual modeling or physical set construction.
  • Immersive visualization: generate a spatial concept that stakeholders can inspect from multiple viewpoints.
  • Digital-twin-style mockups: produce an initial scene representation from a visual reference, while recognizing that further accuracy work may be necessary.
  • Gaussian-splat and point-cloud workflows: obtain scene formats suited to compatible viewers, renderers, or 3D processing pipelines.
  • Collision-aware prototypes: use the available collision mesh outputs when movement through the generated space is part of the test.

When to choose Hy-World-2.1-scene

Choose Hy-World-2.1-scene when the deliverable is an explorable 3D environment and the workflow can tolerate asynchronous generation. It is particularly well matched to teams that want to begin from a short textual description or a reference image and need more than a single rendered picture.

Another option may be more appropriate when the priority is rapid 2D ideation, conversational assistance, code generation, audio, video, or real-time interaction. A conventional image model is likely a better fit for quickly producing many flat visual variations, while a general language model is more suitable for reasoning, writing, coding, and tool orchestration. Hy-World-2.1-scene is also not the obvious choice when a team requires a downloadable local model, an explicitly documented streaming interface, or a guaranteed fixed output-size limit.

In short, its central trade-off is specialization versus immediacy. The model is specialized for generating navigable 3D scene assets, but it requires a hosted asynchronous workflow, has meaningful input restrictions, and carries a reference cost of about ¥80 per scene at the documented usage level.


Answers to Frequently Asked Questions

How much does Hy-World-2.1-scene cost?
Tencent’s reference pricing is ¥10 per one million tokens. With the documented reference usage of 8,000,000 tokens per generated scene, the approximate cost is ¥80 per scene. Actual charges depend on recorded token usage and configuration, and additional costs may apply for storage, processing, rendering, and infrastructure.
How long does Hy-World-2.1-scene take to generate a scene?
Generation is asynchronous and commonly takes approximately 10 to 20 minutes. The API returns a task identifier, and the integrating application must poll the task-status endpoint until the task reaches a terminal state.
What output formats does Hy-World-2.1-scene provide?
A completed task can provide SPZ 3D Gaussian-splat data, PLY point clouds, complete and simplified collision meshes, preview images, super-resolution previews when available, spatial-position metadata, and token-usage information.
What is Hy-World-2.1-scene used for?
Hy-World-2.1-scene is Tencent’s text-to-3D and image-to-3D model for generating interactive, navigable environments. It can support game prototypes, virtual production, immersive visualization, digital-twin-style mockups, and other workflows that require more than a single 2D image.
What inputs does Hy-World-2.1-scene support?
Each request supports either a text prompt or one reference image; text and image inputs cannot be combined. Text prompts are limited to 600 characters. Reference images must be publicly accessible, at least 512 × 512 pixels, no larger than 10 MB, and in JPG, JPEG, PNG, or WEBP format.


Sources 4
Provider

About Tencent AI