What is MedImageParse 3D?
MedImageParse 3D is Microsoft's text-prompted model for segmenting complete three-dimensional medical-imaging studies. In this context, segmentation means assigning voxels—the 3D equivalent of pixels—to a target structure, such as an organ, lesion, or other region of interest. The result is a mask that can be overlaid on the original scan, measured, reviewed, or passed to another imaging system.
The model accepts two main inputs: a natural-language prompt describing what should be identified and a complete medical volume. Microsoft documents NIfTI as the supported volume format. Instead of asking a user to mark a region manually or run a separate model call on every slice, MedImageParse 3D is intended to reason over the volume as a 3D study.
The model is an extension of Microsoft's MedImageParse approach from 2D medical images to volumetric data. Microsoft describes its technical lineage as related to BiomedParse and BoltzFormer. Those names help position the model within Microsoft's biomedical-imaging research work, but MedImageParse 3D should not be treated as a general-purpose conversational model or a medical-report generator.
How text-guided 3D segmentation works
A typical workflow starts with a prepared CT or MRI volume in NIfTI format. The user then supplies a prompt identifying the target, for example an anatomical structure, suspected lesion, or other region represented in the scan. The service returns a segmentation result rather than a prose explanation.
Microsoft's documented response contains a base64-encoded NIfTI segmentation mask. Base64 is an encoding used to transport binary data as text; it does not mean the output is a text description of the scan. After decoding the response, an imaging application or research script can load the NIfTI mask, align it with the source volume, render it in 2D or 3D, and calculate measurements such as approximate volume.
This design is useful when the target changes between requests. A workflow can use different prompts for different structures instead of maintaining a completely separate fixed-label model for every possible segmentation task. However, prompt-based flexibility does not remove the need to validate the resulting masks. Wording, image quality, modality, preprocessing, and the characteristics of the dataset can all affect the output.
Supported inputs and outputs
| Category | Documented behavior |
|---|---|
| Image input | Complete 3D medical-imaging volumes |
| Documented volume format | NIfTI |
| Text input | A prompt describing the anatomy, lesion, or structure to segment |
| Primary output | A three-dimensional segmentation mask |
| Output transport | Base64-encoded NIfTI segmentation result |
| Audio and video | Not supported according to the supplied model information |
| General text generation | Not the model's intended output |
MedImageParse 3D is therefore multimodal in a narrow, task-specific sense: it combines text with 3D medical imagery and produces an image-analysis artifact. It is not a multimodal chat assistant that accepts arbitrary documents, audio, video, or photographs and responds conversationally.
Where it fits in Microsoft's catalog
MedImageParse 3D belongs to Microsoft's healthcare AI and biomedical-imaging model catalog. Microsoft announced the model on March 3, 2025, describing it as an optimization of MedImageParse for 3D CT and MRI data. The model is associated with the MedImageParse family and is also referred to in Microsoft's research catalog as MedImageParse3D.
Its position in the catalog is important. This is a specialized healthcare model, not one of the general-purpose models exposed for broad chat, coding, web search, or content generation. Its value comes from producing a structured segmentation result from volumetric medical data. Users who need a conversational explanation, a clinical narrative, or general medical question answering need a different type of system.
Deployment and access through Foundry classic
Microsoft documents MedImageParse 3D for deployment through the Microsoft Foundry classic healthcare-AI workflow. The documented approach uses a managed compute endpoint with GPU-enabled infrastructure. It is not presented as a simple consumer web application or an unrestricted public chat endpoint.
Deployment requirements described in the documentation include an Azure subscription with a valid payment method, a hub-based project, suitable Azure role-based access, and managed compute capable of running the model. The documentation identifies the workflow as belonging to Foundry classic rather than the newer Foundry portal process. As a result, access can depend on the portal experience, project type, Azure region, subscription permissions, and model-catalog availability.
Operationally, users should plan for more than the model invocation itself. A production or research deployment may involve volume storage, preprocessing, endpoint configuration, GPU runtime, result storage, monitoring, and data-transfer considerations. These surrounding services can be significant for large medical volumes or long-running endpoints.
What MedImageParse 3D does well
- Works on complete volumes: The model is designed for 3D studies rather than only isolated 2D images, which better matches how CT and MRI examinations are commonly organized.
- Uses natural-language prompts: A prompt can identify the desired anatomy or abnormality, making the interface more flexible than a workflow tied to one fixed output label.
- Produces a reusable research artifact: A NIfTI mask can be visualized, measured, compared over time, or used in another imaging pipeline.
- Supports volumetric analysis: Once a mask has been checked, researchers can use it for approximate tumor or organ-volume calculations and other quantitative studies.
- Fits annotation workflows: The output can provide a starting point for human review and dataset labeling, potentially reducing the amount of manual delineation required.
These are practical strengths rather than a claim that the model is accurate for every anatomy, modality, or clinical setting. The supplied documentation does not provide a universal accuracy benchmark or a guarantee of performance across datasets.
Limitations and safety considerations
Microsoft provides its healthcare AI models as-is for research and model-development exploration. The documented model is not intended to be deployed by itself in clinical settings, and its performance for diagnosis or treatment has not been established. A mask from MedImageParse 3D should therefore be treated as an assistive output that requires review, not as ground truth.
Several technical factors can affect results. Teams may need to account for modality-specific preprocessing, volume orientation, voxel spacing, intensity normalization, and prompt wording. False-positive regions, missed structures, incomplete masks, or incorrect boundaries are possible. A visually plausible segmentation can still produce materially wrong measurements if it is not checked against the original study.
The model also does not replace clinical governance. Any use involving patient data should address applicable privacy, security, access-control, retention, validation, and oversight requirements. Research teams should define how outputs are reviewed, how failures are recorded, and when a segmentation is rejected or corrected manually.
Context, output, and feature limits
MedImageParse 3D is not documented with a conventional language-model context window or maximum output-token limit. That omission is expected for a model whose principal output is a segmentation mask rather than generated text. The relevant practical limits are instead connected to the size and characteristics of the supplied volume, the endpoint configuration, available GPU resources, and the resulting mask.
The supplied specifications also do not identify a general tool-calling system, streaming interface, JSON mode, fine-tuning facility, caching feature, or batch API for this model. Its documented output is the segmentation artifact. Developers should not assume that it supports the function-calling and structured-response features commonly associated with general-purpose language models.
Reasoning and coding scores in model databases should be interpreted carefully. MedImageParse 3D can use a text prompt to guide image segmentation, but it is not a reasoning chatbot and is not intended to generate software. The supplied research does not establish general coding ability, chain-of-thought behavior, or broad medical reasoning performance.
Pricing and operational cost
No public per-request or per-volume price specific to MedImageParse 3D was identified in the reviewed authoritative documentation. The documented deployment uses managed Azure compute, so the total cost depends on the selected GPU infrastructure, endpoint configuration, instance count, runtime, storage, and related Azure services.
This means the model may be inexpensive for occasional experiments only if the endpoint is configured and managed carefully, while continuously deployed infrastructure can create ongoing charges even when it is not actively processing requests. Before estimating a project budget, teams should account for volume size, expected request frequency, endpoint uptime, storage, and any preprocessing or postprocessing services. Azure pricing and availability can change, so the absence of a published model-specific price should not be interpreted as free access.
Best use cases
MedImageParse 3D is a strong candidate for teams that need text-guided segmentation of complete CT or MRI volumes and can provide technical and human review. Suitable applications include:
- Research segmentation of organs, lesions, and other structures in CT or MRI studies.
- Annotation assistance for biomedical-imaging datasets.
- Tumor or organ volumetry after clinician or researcher validation.
- Longitudinal studies that compare segmented regions across time points.
- Prototyping custom medical-imaging and image-analysis pipelines.
- Exploring prompt-driven segmentation when a fixed, single-organ model would be too restrictive.
When to choose MedImageParse 3D
Choose MedImageParse 3D when the central task is converting a text description and a complete 3D medical volume into a segmentation mask. It is particularly relevant when the workflow already uses NIfTI, needs volumetric rather than slice-only processing, and can support Azure-managed GPU deployment and validation.
Another type of medical-imaging model may be more appropriate when the target anatomy and labels are fixed, the workflow requires a tightly benchmarked clinical pipeline, or local inference and predictable infrastructure costs are more important than prompt flexibility. A general-purpose vision-language model may be preferable for image descriptions or question answering, but it would not automatically provide a validated 3D NIfTI segmentation mask. Likewise, a conversational AI model is a better fit for report drafting or natural-language interaction, not for this model's primary image-analysis task.
In all cases, the choice should be based on validation results for the intended modality, anatomy, patient population, preprocessing pipeline, and operating environment. The supplied research does not establish that MedImageParse 3D is clinically superior to specialized alternatives.
Bottom line
MedImageParse 3D is a specialized Microsoft healthcare AI model for text-guided segmentation of complete CT and MRI volumes. Its defining output is a three-dimensional NIfTI mask, making it useful for research, annotation, volumetry, and imaging-pipeline prototyping. Its main trade-off is that access and cost are tied to Foundry classic managed GPU deployment, while accuracy and safety require application-specific validation. It is best viewed as a research and development component—not a standalone diagnostic system, general-purpose medical assistant, or replacement for expert review.

