What is HunyuanWorld-Mirror?
HunyuanWorld-Mirror is an open-weight, feed-forward 3D reconstruction model from Tencent Hunyuan. Its job is to infer the geometry of a scene from several related views, such as frames extracted from a video or a sequence of photographs. Instead of returning a text description, it produces numerical representations that can be used in 3D processing, rendering, and computer-vision pipelines.
The model is intended to estimate a coherent scene across multiple views. Depending on the inputs and workflow, its outputs can include 3D point maps, depth maps, surface normals, camera parameters, and data for 3D Gaussian Splatting. Gaussian Splatting is a scene-rendering representation that stores many 3D primitives with properties such as position, opacity, scale, rotation, and appearance features.
Tencent released the inference code and pretrained weights on October 22, 2025. The official repository includes inference scripts, a Gradio demo, training code, and evaluation configurations. The checkpoint is also available through Tencent’s Hugging Face organization.
How the model uses visual and geometric information
WorldMirror uses a multimodal prior-prompting approach. In practical terms, it can combine visual observations with additional geometric information when that information is available. The optional priors include camera-to-world pose matrices, camera intrinsic matrices, and depth maps.
Camera poses describe how each camera is positioned and oriented in the scene. Camera intrinsics describe properties such as focal length and the camera’s projection geometry. Depth maps provide an estimated distance from the camera to visible scene points. Supplying these inputs can give the model more reliable geometric evidence, particularly when camera or depth estimation is an important part of the task.
These inputs are optional rather than universally required. That makes the model applicable to workflows ranging from ordinary image-sequence reconstruction to more calibrated pipelines where camera information has already been computed. The model’s unified prediction design supports several related tasks instead of requiring a separate checkpoint for every geometric output.
Inputs and outputs
Supported inputs
- Image sequences from multiple views.
- Frames extracted from video.
- Optional camera-to-world pose matrices.
- Optional camera intrinsic matrices.
- Optional depth maps.
The supplied research does not document a language-model-style context window, token limit, or maximum text output length. Those limits are not applicable to this visual-geometric checkpoint in the same way they are to a conversational model.
Generated representations
- World-coordinate 3D point clouds or point maps.
- Camera-frame depth maps.
- Surface-normal maps.
- Estimated camera poses and intrinsic parameters.
- 3D Gaussian Splatting attributes, including means, opacities, scales, rotations, and spherical-harmonic features.
The outputs are designed for downstream geometric and rendering workflows. The repository supports saving Gaussian Splatting data, exporting reconstruction results to COLMAP-compatible files, rendering novel views, and optionally optimizing the generated Gaussian representation after inference.
Where it fits in Tencent’s current lineup
Tencent’s later HY-World 2.0 repository identifies HunyuanWorld-Mirror as WorldMirror-1 and describes it as a legacy model relative to WorldMirror-2.0. This does not mean that the checkpoint is unavailable: the original repository, weights, and inference workflow remain publicly downloadable. It does mean that users beginning a new project should evaluate whether the newer WorldMirror-2.0 stack is a better fit for their requirements.
WorldMirror-1 is therefore most useful when reproducibility, access to the original implementation, or compatibility with existing research workflows matters. It should not be confused with a general Hunyuan conversational model, a text-to-image system, or a hosted 3D-generation API. Its specialization is visual geometry reconstruction from images and video.
Strengths and practical benefits
- Unified geometric prediction: One model supports point maps, depth, surface normals, camera estimation, and Gaussian Splatting-related outputs.
- Flexible conditioning: Users can provide camera and depth information when available instead of relying only on raw images.
- Video and multi-view support: The model is suited to sequences of related frames rather than only isolated single-image predictions.
- Open-weight access: Researchers can download the checkpoint, inspect the implementation, run local inference, and use the included training configurations.
- Useful downstream formats: COLMAP-compatible exports and Gaussian Splatting data connect the reconstruction process to established 3D and rendering workflows.
These strengths are especially relevant when the goal is to obtain an editable or renderable scene representation rather than merely generate a visual approximation. The optional priors also make it possible to integrate the model into pipelines that already contain camera calibration or depth-estimation stages.
Limitations and technical requirements
The main practical limitation is that HunyuanWorld-Mirror is a self-hosted research model. The original repository recommends Python 3.10, PyTorch 2.4.0, CUDA 12.4, and the gsplat package for Gaussian Splatting rendering. As a result, successful use generally requires a suitable local GPU environment and familiarity with Python, PyTorch, model checkpoints, and 3D data formats.
No official first-party hosted API price, token allowance, batch endpoint, or managed inference service is documented in the supplied research. There is consequently no verified per-image, per-video, or per-request price to report. Operating cost depends on the user’s own hardware or infrastructure rather than on a published Tencent API tariff.
The research also does not provide a fixed context length or maximum output-token limit. This is expected because the model produces geometric tensors and scene data rather than natural-language responses. The exact memory and runtime requirements can depend on factors such as sequence size, image resolution, GPU capacity, and the selected rendering workflow, but the supplied sources do not establish universal numeric limits.
HunyuanWorld-Mirror is not a conversational reasoning model. It has no documented natural-language reasoning score, coding capability, function-calling interface, web-search tool, or text-generation endpoint. Any internal ratings that describe reasoning, coding, speed, or cost should be treated as editorial assessments, not Tencent-published benchmarks or specifications.
Pricing and availability
The checkpoint is publicly downloadable through Tencent’s official GitHub repository and Hugging Face model repository. The supplied research does not identify a recurring subscription, hosted inference price, or usage-based API fee. For that reason, the most accurate pricing description is open-weight download with self-hosted compute costs.
Users should check the repository and model card for the current license terms before incorporating the model into a commercial product. The available research confirms public code and weights but does not provide enough licensing detail here to summarize additional permissions or restrictions safely.
When to choose HunyuanWorld-Mirror
Choose HunyuanWorld-Mirror when you need a downloadable reconstruction model for research or an on-premises 3D pipeline. It is a reasonable option when your inputs consist of image sequences or video frames and your outputs need to include more than a single mesh or rendered image. For example, it can support experiments involving camera recovery, depth estimation, point-map generation, novel-view synthesis, or Gaussian Splatting.
It is also a good fit when you want to use optional calibrated information. A pipeline that already has camera poses, intrinsics, or depth maps can pass those priors to the model and use them as additional geometric evidence.
Another reason to select it is reproducibility. Researchers can work with the released implementation and weights instead of depending on changes to a hosted service. Training configurations also make self-managed fine-tuning possible, although this is a repository-based training workflow rather than a managed fine-tuning API.
When another option may be more appropriate
A hosted 3D service may be more appropriate if you need an easy web API, predictable request pricing, managed GPU infrastructure, or production support. HunyuanWorld-Mirror does not provide a documented first-party hosted endpoint with those guarantees.
A newer model may also be preferable for a new project if current Tencent tooling, features, or maintenance are more important than compatibility with WorldMirror-1. Tencent explicitly positions WorldMirror-2.0 as the newer successor in the HY-World 2.0 catalog, so users should compare the newer stack before committing to the legacy checkpoint.
Finally, this model is the wrong choice for text generation, conversation, text-to-image creation, speech, audio processing, or general-purpose coding assistance. Its outputs are geometric data, not prose or application code. For those tasks, a model designed for the relevant modality will be more suitable.
Bottom line
HunyuanWorld-Mirror is a specialized Tencent model for reconstructing 3D worlds from image sequences and video. Its defining practical feature is the combination of flexible geometric conditioning with multiple reconstruction outputs in one feed-forward system. It remains useful for open research and self-hosted 3D workflows, especially when point maps, depth, camera estimates, surface normals, or Gaussian Splatting data are required.
Its trade-off is operational rather than conceptual: users must manage the local software and GPU environment, and there is no documented hosted API price or language-model-style output limit. Because Tencent now treats WorldMirror-1 as the legacy predecessor to WorldMirror-2.0, it is best selected deliberately for reproducibility, availability, or established research compatibility rather than assumed to be the default choice for every new deployment.

