What SAM 3D Body does
SAM 3D Body is Meta’s specialized model for recovering a 3D human body from a single RGB image. Instead of returning a caption, classification, or general image description, it estimates a structured representation of a person’s body, including a full-body mesh, skeletal pose, and body-shape parameters.
A single image does not contain direct information about every hidden surface or the person’s exact depth. SAM 3D Body therefore produces an estimated reconstruction rather than a guaranteed scan or measurement. Its purpose is to provide a useful 3D representation for downstream visualization, analysis, animation, and embodied-AI experiments.
The model belongs to Meta’s SAM 3D release, which also includes SAM 3D Objects for object and scene reconstruction. SAM 3D Body is the human-focused component: its subject is an individual person, and its output is organized around human pose and body shape.
How the model is guided
SAM 3D Body can work from an RGB image and can receive additional guidance through segmentation masks or two-dimensional keypoints. A segmentation mask indicates which pixels belong to the person, while 2D keypoints identify visible body landmarks. These prompts help focus the reconstruction when the image is complicated or when several visual cues are ambiguous.
This prompt-based design is useful in interactive workflows. For example, a user or an upstream vision system can identify the person of interest and provide a mask before asking the model to estimate that individual’s 3D pose and shape. The model is still intended to process people individually rather than solve an entire multi-person scene as one unified human-interaction problem.
Strengths and output representation
The main strength of SAM 3D Body is its focus on full-body reconstruction from difficult real-world images. Meta presents it as being designed for unusual poses, partial occlusion, varied clothing, and other in-the-wild conditions that make monocular 3D estimation difficult.
- Full-body 3D recovery: The output includes a body mesh together with pose and shape information.
- Promptable estimation: Segmentation masks and 2D keypoints can provide additional guidance.
- Human-specific structure: The output is based on Meta’s Momentum Human Rig, or MHR, rather than an unstructured visual prediction.
- Research-oriented release: Meta provides inference code, checkpoints, evaluation resources, and related human-rig assets.
The Momentum Human Rig separates skeletal structure from soft-tissue body shape. That separation can make the result easier to inspect or use in animation and analysis pipelines than a purely image-based output. Meta has also released the related parametric human model under a permissive commercial license, according to the supplied release information.
Technical architecture and technical limits
SAM 3D Body uses a transformer encoder-decoder architecture to predict MHR mesh parameters. Its image encoder uses multiple input pathways intended to preserve high-resolution information from body parts. The mesh decoder supports prompt-based prediction and iterative refinement.
The supplied specifications do not define a conventional context window or maximum output-token limit. Those fields apply mainly to language models and hosted generative APIs, whereas SAM 3D Body produces structured 3D predictions from visual input. The relevant practical limits are instead determined by the image, the quality of the visible evidence, the selected checkpoint, and the local inference environment.
The model accepts a single RGB image as its primary input, with optional mask or 2D-keypoint guidance. It does not have documented audio or video input in the supplied research. Its direct output is a non-text 3D human representation rather than text, image generation, audio, or video generation.
Availability, checkpoints, and pricing
SAM 3D Body was released by Meta as research software and downloadable model checkpoints. The official release date listed in the supplied data is November 19, 2025. Code is available through the Facebook Research repository, and checkpoints are distributed through Hugging Face.
The documented checkpoint variants include facebook/sam-3d-body-dinov3 and facebook/sam-3d-body-vith. Access to the Hugging Face checkpoints requires an access request and authentication. Users should also expect a local machine-learning installation involving PyTorch, Detectron2, and GPU-oriented dependencies.
No official hosted API pricing is published for SAM 3D Body. It is not presented as a metered, general-purpose online model with per-image input and output rates. The financial trade-off is therefore different from an API model: there may be no vendor inference charge for local use, but users must provide compatible hardware, storage, setup time, and maintenance. The actual operating cost depends on the machine and deployment environment, which are not specified in the supplied research.
Reasoning, coding, and tool support
SAM 3D Body is not a language model and should not be evaluated as a chatbot or coding assistant. It does not provide general-purpose reasoning, text generation, code generation, function calling, web search, streaming responses, or JSON-mode text output in the documented release.
It can be integrated into a larger software pipeline, but that is different from having built-in tools. An application may use its mesh and pose predictions as input to an animation system, robotics stack, AR experience, or analysis program. The surrounding application supplies that orchestration; SAM 3D Body itself is the vision component that estimates the human representation.
Editorial capability scores in the supplied data rate it relatively low for reasoning and coding and relatively higher for the speed and cost characteristics of its specialized category. These are comparative editorial assessments, not Meta-published benchmarks. They should not be interpreted as measurements against language models or as guaranteed inference performance.
Limitations to consider
The most important limitation is that SAM 3D Body infers three-dimensional structure from a single two-dimensional view. Hidden limbs, depth, camera properties, body proportions, and fine anatomy can remain ambiguous. A visually plausible mesh is not necessarily an accurate physical measurement or exact scan.
The model processes people individually and does not fully reason about interactions between multiple people, objects, or the surrounding environment. A photograph showing two people interacting may require separate processing and additional application logic. The model is therefore not a complete human-object interaction system or a scene-understanding model.
Hand pose is another stated limitation. Meta reports improvements as part of whole-body estimation, but hand-pose accuracy does not exceed that of specialized hand-only pose-estimation systems. Applications that depend on precise finger articulation may need a dedicated hand-tracking or hand-pose option in addition to, or instead of, SAM 3D Body.
Results can also degrade when the subject is heavily occluded, the body is cropped, visual details are unclear, or the image does not provide enough evidence for a reliable reconstruction. The model should be treated as an estimation tool whose output may need checking and post-processing.
Best use cases
- Computer-vision research: Investigating monocular 3D human pose and shape estimation.
- AR and VR: Creating or prototyping avatars from photographs.
- Interactive media and games: Generating a starting point for character or movement workflows.
- Robotics and embodied AI: Supplying human-pose and body-shape estimates to perception experiments.
- Sports and movement analysis: Studying body configuration in still images or image-based datasets.
- Benchmarking and model development: Testing pipelines that need a structured human mesh rather than a text description.
These uses are strongest when an approximate, structured representation is more valuable than a certified measurement. A researcher can use the mesh to visualize pose, compare predictions, or drive a downstream process, while still validating the result for applications where anatomical accuracy matters.
When to choose SAM 3D Body
Choose SAM 3D Body when the main requirement is a 3D human mesh, pose, and body-shape estimate from one RGB image, especially when local deployment and research-code access are acceptable. Its specialized output makes it more appropriate than a general vision-language model for a pipeline that needs explicit human-body parameters.
It may be a good fit when you need to experiment with difficult poses, occlusion, segmentation guidance, or the Momentum Human Rig representation. It is also a sensible choice when you prefer downloadable checkpoints over a hosted API and can manage the required GPU-oriented software environment.
Another type of option may be more appropriate when you need a simple hosted endpoint, predictable per-request pricing, real-time video tracking, multi-person interaction reasoning, exact body measurement, or highly precise hand pose. A general vision model may be more useful for scene interpretation and natural-language answers, while a dedicated hand or video-tracking system may be better for continuous motion and finger-level detail. Those alternatives solve different problems; SAM 3D Body is specifically optimized around single-image, full-body 3D reconstruction.
SAM 3D Body at a glance
| Attribute | Details |
|---|---|
| Provider | Meta AI |
| Model family | SAM 3D |
| Primary input | Single RGB image, optionally guided by a segmentation mask or 2D keypoints |
| Primary output | 3D human mesh, pose, and body-shape parameters |
| Architecture | Transformer encoder-decoder predicting Momentum Human Rig parameters |
| Deployment | Downloadable research code and gated Hugging Face checkpoints |
| Hosted API pricing | No official hosted API pricing published |
| Built-in tools | None documented |
| Key limitation | Single-person estimation from one image; hidden structure and exact measurements remain uncertain |
SAM 3D Body is best understood as a specialized reconstruction component rather than a general AI assistant. Its value comes from converting limited visual evidence into a structured human representation that other 3D, robotics, animation, or research systems can use.

