What is Hy-World-2.1-scene?
Hy-World-2.1-scene is Tencent’s specialized world-generation model for producing interactive 3D environments. It is available through Tencent Cloud TokenHub under the API model identifier hy-world2-scene. Instead of returning a conventional 2D picture, the service generates assets intended to represent a navigable scene that can be explored, rendered, and used in downstream 3D workflows.
The model is part of Tencent’s current TokenHub catalog of 3D-generation services. Its primary output is a scene that can combine 3D Gaussian-splat data, point-cloud information, meshes, preview images, and spatial metadata. Tencent describes the result as supporting free movement and physical collision, which makes the service more relevant to interactive environments than to static image generation.
For a beginner, the practical distinction is simple: a text-to-image model generally produces one flat image, while Hy-World-2.1-scene is intended to provide the underlying representation of a space. That makes it potentially useful when the environment needs to be viewed from different positions or incorporated into an interactive prototype.
How inputs and generation work
Each request uses one supported input route: either a text prompt or a reference image. Text and image inputs cannot be submitted together in the same request. Text prompts can contain up to 600 characters, so the input is best treated as a concise description of the desired environment rather than a long design document.
When using an image, the image must be publicly accessible. Tencent’s documented requirements specify a minimum size of 512 × 512 pixels, a maximum file size of 10 MB, and support for JPG, JPEG, PNG, and WEBP formats. These restrictions matter in production workflows because an otherwise suitable local file may need to be uploaded to an accessible location before the request can be made.
Generation is asynchronous. A successful submission returns a task identifier rather than immediately returning a finished scene. An application must poll the task-status endpoint until the task reaches a terminal state. Tencent documentation indicates that generation commonly takes approximately 10 to 20 minutes. This makes the model more appropriate for queued or batch-style creation workflows than for an interactive application that expects an immediate response after every prompt.
What the model outputs
A completed task can provide several kinds of scene data. The exact result should be treated as a collection of assets rather than a single image file.
- SPZ 3D Gaussian-splat data: a compact scene representation based on Gaussian splats, which can be rendered to show the generated environment from different viewpoints.
- PLY point cloud data: a point-based representation that can be used in compatible 3D and visualization tools.
- Complete collision mesh: geometry intended to support collision-aware interaction with the scene.
- Simplified collision mesh: a lower-complexity collision representation for workflows where efficient interaction is more important than full geometric detail.
- Preview images: rendered views that help users inspect the result without immediately loading the full scene assets.
- Super-resolution preview images: higher-resolution preview output when provided by the completed task.
- Spatial-position metadata: information describing the scene’s spatial arrangement or positioning.
- Token-usage information: usage data relevant to billing and accounting.
These outputs give Hy-World-2.1-scene a different profile from a model that only returns a rendered image. However, the supplied documentation does not specify a universal maximum scene size, maximum output token count, or fixed geometry budget. Those limits should therefore not be assumed when planning a pipeline.
Supported modalities and capabilities
| Area | Documented position |
|---|---|
| Text input | Yes; prompts are limited to 600 characters |
| Image input | Yes; one publicly accessible image per request |
| Audio input | No documented support |
| Video input | No documented support |
| 3D output | Yes; scene, Gaussian-splat, point-cloud, and mesh assets |
| Text output | Not the model’s primary output |
| Tool or function calling | No documented support |
| Streaming | No documented support; generation is asynchronous |
| Structured JSON output | No documented model capability |
The model’s multimodal behavior should be understood as text-to-3D or image-to-3D generation. It accepts more than one input type, but it does not combine text and image inputs in a single request according to the supplied documentation. Its output is also not ordinary conversational text: the useful result is a set of visual and geometric scene assets.
Pricing and cost trade-offs
Tencent’s documented reference pricing is ¥10 per one million tokens. The reference usage for one generated scene is 8,000,000 tokens, giving an approximate cost of ¥80 per scene at that usage level. This is a reference calculation rather than a guaranteed flat price: actual charges depend on recorded token usage and configuration.
The price places the service in a different category from ordinary 2D image generation. A generated scene may include several representations and supporting assets, so the relevant comparison is not simply the cost of one preview image. Organizations should also account for storage, processing, asset conversion, rendering, and any application infrastructure needed to use the returned scene.
Because generation commonly takes 10 to 20 minutes, Hy-World-2.1-scene also trades speed for richer output. It can be a reasonable choice when producing a complete environment is more important than immediate feedback. It is less suitable when a user needs many rapid visual variations during a live interaction.
Main strengths
- Scene-oriented output: the service is built around navigable 3D environments rather than isolated 2D views.
- Two practical starting points: users can describe a scene with text or provide a visual reference image.
- Multiple downstream assets: SPZ Gaussian-splat data, PLY point clouds, collision meshes, previews, and spatial metadata can support different stages of a 3D workflow.
- Collision-aware potential: complete and simplified collision meshes are useful when a scene must support movement or physical interaction.
- Useful for rapid environment prototyping: teams can create a 3D starting point without manually building every element before testing an idea.
These strengths are based on the documented output and workflow. They should not be interpreted as a guarantee that every generated scene will be production-ready without cleanup, optimization, or integration work.
Limitations to plan for
The first operational limitation is latency. The asynchronous workflow and typical 10-to-20-minute generation time make the service unsuitable for low-latency inference. A production application needs task tracking, status polling, failure handling, and a way to notify users when assets are ready.
The input rules are also restrictive. Text is limited to 600 characters, and text and image guidance cannot be combined in one request. Image inputs must meet the documented accessibility, dimension, format, and file-size requirements. Users who need detailed art direction may have to encode the most important requirements into a short prompt or prepare a carefully selected reference image.
The service is hosted through Tencent Cloud TokenHub rather than described in the supplied research as a downloadable standalone model. That means deployment depends on the hosted service and its API workflow. The documentation supplied here also does not provide a context window, maximum output-token value, release date, or fixed maximum scene complexity.
Finally, 3D scene generation is not the same as delivering a finished game level or polished digital twin. Even when the output includes meshes and collision assets, a downstream team may still need to inspect geometry, optimize files, configure materials, integrate rendering, and adapt the result to its target engine.
Reasoning, coding, and tool-use expectations
Hy-World-2.1-scene is not a general-purpose language model. It is not designed to answer questions, write application code, perform extended reasoning, or call external tools as part of a conversational workflow. The supplied editorial assessment assigns low practical relevance to reasoning and coding because those are not the model’s primary tasks; these are editorial evaluations, not vendor-published benchmark scores.
Likewise, no tool or function-calling capability is documented. The API workflow itself still requires ordinary application logic for submitting jobs, polling task status, downloading returned assets, and handling errors, but that orchestration is performed by the integrating application rather than by the model through a documented tool-use interface.
Best use cases
- Interactive environment prototypes: create an explorable starting point for a game, simulation, or virtual-world concept.
- Virtual production: explore environment ideas before committing to detailed manual modeling or physical set construction.
- Immersive visualization: generate a spatial concept that stakeholders can inspect from multiple viewpoints.
- Digital-twin-style mockups: produce an initial scene representation from a visual reference, while recognizing that further accuracy work may be necessary.
- Gaussian-splat and point-cloud workflows: obtain scene formats suited to compatible viewers, renderers, or 3D processing pipelines.
- Collision-aware prototypes: use the available collision mesh outputs when movement through the generated space is part of the test.
When to choose Hy-World-2.1-scene
Choose Hy-World-2.1-scene when the deliverable is an explorable 3D environment and the workflow can tolerate asynchronous generation. It is particularly well matched to teams that want to begin from a short textual description or a reference image and need more than a single rendered picture.
Another option may be more appropriate when the priority is rapid 2D ideation, conversational assistance, code generation, audio, video, or real-time interaction. A conventional image model is likely a better fit for quickly producing many flat visual variations, while a general language model is more suitable for reasoning, writing, coding, and tool orchestration. Hy-World-2.1-scene is also not the obvious choice when a team requires a downloadable local model, an explicitly documented streaming interface, or a guaranteed fixed output-size limit.
In short, its central trade-off is specialization versus immediacy. The model is specialized for generating navigable 3D scene assets, but it requires a hosted asynchronous workflow, has meaningful input restrictions, and carries a reference cost of about ¥80 per scene at the documented usage level.

