What is HY-3D-Motion?
HY-3D-Motion is a Tencent Hunyuan model for text-to-human-motion generation. It accepts a natural-language description of an action and creates a corresponding animated 3D motion sequence. The result is delivered as an FBX file, a widely used format for exchanging 3D scenes, characters, and animation data between production tools.
The model is available through Tencent’s TokenHub service. This positioning makes HY-3D-Motion an API-accessible generation component rather than a general conversational assistant or a complete character-creation application. Its role is to supply motion data that can be incorporated into a wider animation, game-development, visualization, or 3D-content pipeline.
How the model works
A request begins with a text prompt describing the desired human movement. For example, a prompt could specify that a character should walk, run, jump, or perform another action. The API then creates an asynchronous job: the submission response provides a task identifier, and a separate query request is used to check progress and obtain the result.
The documented job states include queued, in_progress, completed, and failed. This asynchronous design is important for production integrations because an application should not assume that the FBX file is returned immediately with the submission response. Instead, it should retain the task identifier and poll or otherwise query the job status according to its workflow.
When processing succeeds, the output is an animated FBX file. The service is therefore best understood as a motion-data generator: it does not return a written explanation, an image, a rendered video, or an audio file.
Key controls and documented limits
HY-3D-Motion provides several controls that affect the generated motion and the contents of the returned asset:
- Prompt: The TokenHub endpoint documents a maximum prompt length of 128 characters.
- Duration: Generated motion can be configured from 1 to 12 seconds, with a documented default of 5 seconds.
- Prompt rewriting: An optional rewriting function can be enabled to process the user’s motion description before generation.
- Duration estimation: The API can estimate an appropriate duration when this option is enabled.
- Retargeting: A retargeting file can be supplied for compatible Tencent Hunyuan 3D animation assets. Retargeting adapts motion to a different compatible character or rig.
- Skinning mesh: The caller can choose whether the returned FBX includes a skinned mesh in addition to animation data.
The 128-character prompt limit is a practical constraint. It favors concise action descriptions over long scene directions, dialogue, story context, or detailed multi-character choreography. Similarly, the 1-to-12-second duration range makes the service suitable for short motion clips, but it does not document support for long-form continuous sequences in a single request.
Output and production workflow
The main practical strength of HY-3D-Motion is the format of its output. An FBX result can be moved into a compatible 3D or animation workflow for inspection, editing, retargeting, or combination with other assets. This is more directly useful to a production pipeline than a text-only description of how an animation should look.
A typical integration would collect a concise action prompt, choose a duration, optionally enable rewriting or duration estimation, and submit the job. After receiving the task identifier, the integration would query the job until it reaches a terminal state. A completed job can then be used to retrieve the FBX result, while a failed job should be handled as an unsuccessful generation rather than treated as a valid animation.
The optional mesh setting also allows a workflow to distinguish between a motion-focused result and an asset that includes a skinned character mesh. The supplied documentation confirms that this choice is available, but does not specify the mesh topology, character design, skeleton details, or compatibility guarantees for every downstream application.
Where HY-3D-Motion fits in Tencent’s catalog
HY-3D-Motion belongs to Tencent’s Hunyuan 3D and motion-generation ecosystem. Its specialization is human movement: it is not presented as a general 3D asset generator, a language model, an image model, a video model, a speech system, or an embedding service.
The model should also be distinguished from the open-weight HY-Motion-1.0 project. The supplied documentation identifies them as separate offerings, even though both are associated with Tencent’s broader Hunyuan motion-generation work. HY-3D-Motion is the TokenHub service described here, with an API workflow and per-request points pricing.
Pricing and access
Tencent’s current TokenHub pricing information lists HY-3D-Motion at 10 points per generation request. Tencent states that one point corresponds to CNY 0.12, which makes the listed charge equivalent to CNY 1.20 per request before any account-specific billing conditions or other commercial terms.
This is a request-based price rather than a conventional language-model token price. The supplied documentation does not provide separate input-token and output-token rates, and it does not document a context window or maximum text-output-token limit because the service produces motion files instead of textual model responses.
Actual project cost depends on how many generations are attempted. Iterating on prompts, regenerating unsatisfactory motion, or producing multiple variations can multiply the per-request charge. Teams evaluating the service should therefore estimate cost by generation count rather than by text volume.
Capabilities and trade-offs
HY-3D-Motion’s strongest capability is its narrow alignment with text-to-human-animation work. It accepts text input and returns a non-text 3D animation artifact, making it relevant when the desired result is a short motion clip rather than an answer or a visual concept image.
It is not documented as a reasoning or coding model. It does not provide general tool or function-calling support, streaming responses, fine-tuning, batch processing, or JSON output as distinct capabilities. Those omissions matter when comparing it with general-purpose language models: HY-3D-Motion should not be selected for planning software architecture, writing code, answering research questions, or orchestrating broad tool-based workflows.
Its output also has a different trade-off from image and video generation. An FBX file contains usable animation data for a 3D pipeline, while an image or rendered video is immediately viewable but is less directly editable as skeletal motion. For teams that need a controllable character animation asset, the FBX format can be more useful than a flat visual result. For teams that only need a finished visual clip, an image-to-video or text-to-video system may be more appropriate, although those alternatives solve a different problem.
Editorially, HY-3D-Motion is best rated as a specialized, moderately fast and relatively cost-conscious motion-generation option rather than as a broad AI system. Those are evaluation judgments, not Tencent-published benchmark results. The supplied research does not include latency benchmarks, motion-quality benchmarks, or comparative accuracy tests, so users should validate quality on their own target actions and character workflows.
When to choose HY-3D-Motion
HY-3D-Motion is a reasonable choice when the main requirement is to create short human motion clips from concise natural-language descriptions and receive them as FBX files. Suitable use cases include:
- Rapidly prototyping walking, running, jumping, and other human actions.
- Generating starting points for character-animation blocking.
- Creating motion variations for a 3D project without manually keyframing every initial sequence.
- Testing a Tencent-based pipeline that can use optional retargeting and skinned-mesh controls.
- Producing short animation assets through an asynchronous API rather than operating a local motion-generation system.
The model is less suitable when a project requires long continuous performances, extensive narrative instructions, non-human subjects, finished rendered video, images, speech, text generation, or deep software-development assistance. Its prompt limit and maximum documented duration should be checked before it is adopted for a larger choreography or cinematic workflow.
Limitations to check before use
The most important limitations are the short duration range, the 128-character prompt limit, and the model’s focus on human motion. The documentation does not establish how well it handles complex interactions, precise choreography, multiple characters, unusual body mechanics, or highly specific emotional performance. Those cases should be tested rather than assumed to work from the basic action-generation description.
Retargeting is also conditional: the documentation refers to compatible Tencent Hunyuan 3D animation assets, so a supplied retargeting file should not be presumed to work with every rig or character format. Likewise, enabling a skinned mesh does not by itself guarantee that the result is production-ready for a particular game engine or digital-content-creation tool.
Finally, the asynchronous API requires application-level job management. A robust integration needs to handle queued and in-progress states, retrieve completed results, and respond appropriately to failures. The documentation supplied for this model does not provide a guaranteed completion time, so latency should be measured in the intended deployment environment.
Bottom line
HY-3D-Motion is a focused Tencent TokenHub service for converting concise text descriptions into short animated human-motion FBX files. Its value comes from connecting natural-language action prompts with a format that can enter a 3D production workflow. Configurable duration, prompt rewriting, duration estimation, optional retargeting, and mesh inclusion give developers useful control over the generated asset.
It is not a general AI model and should not be evaluated by language-model features such as text reasoning, coding, or conversational context length. Choose it when the target output is editable 3D human motion; choose another type of system when the project needs text, images, rendered video, audio, long-form choreography, or broader automation capabilities.

