MedImageParse

MedImageParse 3D

by Microsoft Copilot · Available through Microsoft Foundry classic managed compute; research and model-development use only

Microsoft MedImageParse 3D is a text-guided healthcare AI model for segmenting complete CT and MRI volumes. It accepts NIfTI data and a prompt describing the target structure, then returns a three-dimensional NIfTI segmentation mask through a Microsoft Foundry classic managed endpoint. The model is intended for research, annotation assistance, volumetry, and pipeline development, not standalone clinical decision-making.

Image generation Reasoning Coding
MedImageParse 3D brings text-guided image segmentation to complete medical volumes instead of individual 2D slices. A researcher can describe the structure or abnormality of interest, provide a NIfTI-encoded CT or MRI volume, and receive a machine-readable 3D mask for visualization and downstream analysis. Microsoft documents the model for deployment through Foundry classic managed compute and positions it for research and model-development exploration.
Outputs

What MedImageParse 3D can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
5/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family MedImageParse
Model type Other
Release date 2025-03-03
Status Available through Microsoft Foundry classic managed compute; research and model-development use only
Knowledge cutoff notes

No authoritative knowledge-cutoff date was identified for this segmentation model. It is an image-analysis model rather than a general-purpose language model, and its documented behavior is based on learned biomedical imaging data and text-guided segmentation capabilities.

Model notes

The exact model is documented as MedImageParse 3D, while Microsoft's research catalog also refers to it as MedImageParse3D. It accepts a text prompt together with a complete 3D medical volume and returns a base64-encoded NIfTI segmentation mask. Microsoft describes the model as built on the BiomedParse and BoltzFormer research lineage. Deployment is documented through Microsoft Foundry classic using managed compute and GPU infrastructure. No exact public per-volume or per-request model price was identified. Microsoft states that its healthcare AI models are provided as-is for research and model-development exploration and are not intended for standalone clinical use.

Model guide

Microsoft MedImageParse 3D for Text-Guided CT and MRI Segmentation

Microsoft MedImageParse 3D is a research-oriented healthcare AI model that combines a natural-language prompt with a complete 3D medical-imaging volume, such as a CT or MRI study, to produce a three-dimensional segmentation mask. It is designed for organ and lesion delineation, volumetric analysis, annotation assistance, and medical-imaging pipeline development rather than standalone diagnosis or treatment decisions.

What is MedImageParse 3D?

MedImageParse 3D is Microsoft's text-prompted model for segmenting complete three-dimensional medical-imaging studies. In this context, segmentation means assigning voxels—the 3D equivalent of pixels—to a target structure, such as an organ, lesion, or other region of interest. The result is a mask that can be overlaid on the original scan, measured, reviewed, or passed to another imaging system.

The model accepts two main inputs: a natural-language prompt describing what should be identified and a complete medical volume. Microsoft documents NIfTI as the supported volume format. Instead of asking a user to mark a region manually or run a separate model call on every slice, MedImageParse 3D is intended to reason over the volume as a 3D study.

The model is an extension of Microsoft's MedImageParse approach from 2D medical images to volumetric data. Microsoft describes its technical lineage as related to BiomedParse and BoltzFormer. Those names help position the model within Microsoft's biomedical-imaging research work, but MedImageParse 3D should not be treated as a general-purpose conversational model or a medical-report generator.

How text-guided 3D segmentation works

A typical workflow starts with a prepared CT or MRI volume in NIfTI format. The user then supplies a prompt identifying the target, for example an anatomical structure, suspected lesion, or other region represented in the scan. The service returns a segmentation result rather than a prose explanation.

Microsoft's documented response contains a base64-encoded NIfTI segmentation mask. Base64 is an encoding used to transport binary data as text; it does not mean the output is a text description of the scan. After decoding the response, an imaging application or research script can load the NIfTI mask, align it with the source volume, render it in 2D or 3D, and calculate measurements such as approximate volume.

This design is useful when the target changes between requests. A workflow can use different prompts for different structures instead of maintaining a completely separate fixed-label model for every possible segmentation task. However, prompt-based flexibility does not remove the need to validate the resulting masks. Wording, image quality, modality, preprocessing, and the characteristics of the dataset can all affect the output.

Supported inputs and outputs

CategoryDocumented behavior
Image inputComplete 3D medical-imaging volumes
Documented volume formatNIfTI
Text inputA prompt describing the anatomy, lesion, or structure to segment
Primary outputA three-dimensional segmentation mask
Output transportBase64-encoded NIfTI segmentation result
Audio and videoNot supported according to the supplied model information
General text generationNot the model's intended output

MedImageParse 3D is therefore multimodal in a narrow, task-specific sense: it combines text with 3D medical imagery and produces an image-analysis artifact. It is not a multimodal chat assistant that accepts arbitrary documents, audio, video, or photographs and responds conversationally.

Where it fits in Microsoft's catalog

MedImageParse 3D belongs to Microsoft's healthcare AI and biomedical-imaging model catalog. Microsoft announced the model on March 3, 2025, describing it as an optimization of MedImageParse for 3D CT and MRI data. The model is associated with the MedImageParse family and is also referred to in Microsoft's research catalog as MedImageParse3D.

Its position in the catalog is important. This is a specialized healthcare model, not one of the general-purpose models exposed for broad chat, coding, web search, or content generation. Its value comes from producing a structured segmentation result from volumetric medical data. Users who need a conversational explanation, a clinical narrative, or general medical question answering need a different type of system.

Deployment and access through Foundry classic

Microsoft documents MedImageParse 3D for deployment through the Microsoft Foundry classic healthcare-AI workflow. The documented approach uses a managed compute endpoint with GPU-enabled infrastructure. It is not presented as a simple consumer web application or an unrestricted public chat endpoint.

Deployment requirements described in the documentation include an Azure subscription with a valid payment method, a hub-based project, suitable Azure role-based access, and managed compute capable of running the model. The documentation identifies the workflow as belonging to Foundry classic rather than the newer Foundry portal process. As a result, access can depend on the portal experience, project type, Azure region, subscription permissions, and model-catalog availability.

Operationally, users should plan for more than the model invocation itself. A production or research deployment may involve volume storage, preprocessing, endpoint configuration, GPU runtime, result storage, monitoring, and data-transfer considerations. These surrounding services can be significant for large medical volumes or long-running endpoints.

What MedImageParse 3D does well

  • Works on complete volumes: The model is designed for 3D studies rather than only isolated 2D images, which better matches how CT and MRI examinations are commonly organized.
  • Uses natural-language prompts: A prompt can identify the desired anatomy or abnormality, making the interface more flexible than a workflow tied to one fixed output label.
  • Produces a reusable research artifact: A NIfTI mask can be visualized, measured, compared over time, or used in another imaging pipeline.
  • Supports volumetric analysis: Once a mask has been checked, researchers can use it for approximate tumor or organ-volume calculations and other quantitative studies.
  • Fits annotation workflows: The output can provide a starting point for human review and dataset labeling, potentially reducing the amount of manual delineation required.

These are practical strengths rather than a claim that the model is accurate for every anatomy, modality, or clinical setting. The supplied documentation does not provide a universal accuracy benchmark or a guarantee of performance across datasets.

Limitations and safety considerations

Microsoft provides its healthcare AI models as-is for research and model-development exploration. The documented model is not intended to be deployed by itself in clinical settings, and its performance for diagnosis or treatment has not been established. A mask from MedImageParse 3D should therefore be treated as an assistive output that requires review, not as ground truth.

Several technical factors can affect results. Teams may need to account for modality-specific preprocessing, volume orientation, voxel spacing, intensity normalization, and prompt wording. False-positive regions, missed structures, incomplete masks, or incorrect boundaries are possible. A visually plausible segmentation can still produce materially wrong measurements if it is not checked against the original study.

The model also does not replace clinical governance. Any use involving patient data should address applicable privacy, security, access-control, retention, validation, and oversight requirements. Research teams should define how outputs are reviewed, how failures are recorded, and when a segmentation is rejected or corrected manually.

Context, output, and feature limits

MedImageParse 3D is not documented with a conventional language-model context window or maximum output-token limit. That omission is expected for a model whose principal output is a segmentation mask rather than generated text. The relevant practical limits are instead connected to the size and characteristics of the supplied volume, the endpoint configuration, available GPU resources, and the resulting mask.

The supplied specifications also do not identify a general tool-calling system, streaming interface, JSON mode, fine-tuning facility, caching feature, or batch API for this model. Its documented output is the segmentation artifact. Developers should not assume that it supports the function-calling and structured-response features commonly associated with general-purpose language models.

Reasoning and coding scores in model databases should be interpreted carefully. MedImageParse 3D can use a text prompt to guide image segmentation, but it is not a reasoning chatbot and is not intended to generate software. The supplied research does not establish general coding ability, chain-of-thought behavior, or broad medical reasoning performance.

Pricing and operational cost

No public per-request or per-volume price specific to MedImageParse 3D was identified in the reviewed authoritative documentation. The documented deployment uses managed Azure compute, so the total cost depends on the selected GPU infrastructure, endpoint configuration, instance count, runtime, storage, and related Azure services.

This means the model may be inexpensive for occasional experiments only if the endpoint is configured and managed carefully, while continuously deployed infrastructure can create ongoing charges even when it is not actively processing requests. Before estimating a project budget, teams should account for volume size, expected request frequency, endpoint uptime, storage, and any preprocessing or postprocessing services. Azure pricing and availability can change, so the absence of a published model-specific price should not be interpreted as free access.

Best use cases

MedImageParse 3D is a strong candidate for teams that need text-guided segmentation of complete CT or MRI volumes and can provide technical and human review. Suitable applications include:

  • Research segmentation of organs, lesions, and other structures in CT or MRI studies.
  • Annotation assistance for biomedical-imaging datasets.
  • Tumor or organ volumetry after clinician or researcher validation.
  • Longitudinal studies that compare segmented regions across time points.
  • Prototyping custom medical-imaging and image-analysis pipelines.
  • Exploring prompt-driven segmentation when a fixed, single-organ model would be too restrictive.

When to choose MedImageParse 3D

Choose MedImageParse 3D when the central task is converting a text description and a complete 3D medical volume into a segmentation mask. It is particularly relevant when the workflow already uses NIfTI, needs volumetric rather than slice-only processing, and can support Azure-managed GPU deployment and validation.

Another type of medical-imaging model may be more appropriate when the target anatomy and labels are fixed, the workflow requires a tightly benchmarked clinical pipeline, or local inference and predictable infrastructure costs are more important than prompt flexibility. A general-purpose vision-language model may be preferable for image descriptions or question answering, but it would not automatically provide a validated 3D NIfTI segmentation mask. Likewise, a conversational AI model is a better fit for report drafting or natural-language interaction, not for this model's primary image-analysis task.

In all cases, the choice should be based on validation results for the intended modality, anatomy, patient population, preprocessing pipeline, and operating environment. The supplied research does not establish that MedImageParse 3D is clinically superior to specialized alternatives.

Bottom line

MedImageParse 3D is a specialized Microsoft healthcare AI model for text-guided segmentation of complete CT and MRI volumes. Its defining output is a three-dimensional NIfTI mask, making it useful for research, annotation, volumetry, and imaging-pipeline prototyping. Its main trade-off is that access and cost are tied to Foundry classic managed GPU deployment, while accuracy and safety require application-specific validation. It is best viewed as a research and development component—not a standalone diagnostic system, general-purpose medical assistant, or replacement for expert review.


Answers to Frequently Asked Questions

How much does MedImageParse 3D cost?
No public per-request or per-volume price specific to MedImageParse 3D was identified. Costs depend on the Azure GPU infrastructure, endpoint configuration, runtime, instance count, storage, data transfer, and related services, so the model should not be assumed to be free.
How is MedImageParse 3D deployed?
Microsoft documents deployment through the Foundry classic healthcare-AI workflow using managed, GPU-enabled compute. Deployment may require an Azure subscription, a hub-based project, suitable Azure permissions, compatible regions and model availability, as well as supporting services for storage, preprocessing, monitoring, and result management.
Can MedImageParse 3D be used for clinical diagnosis?
No. Microsoft provides the model for research and model-development exploration, and its performance for diagnosis or treatment has not been established. Its segmentations should be treated as assistive outputs that require expert review and application-specific validation.
What input and output formats does MedImageParse 3D support?
The documented input is a complete 3D medical-imaging volume in NIfTI format, together with a text prompt describing the target structure. The model returns a three-dimensional segmentation mask as a base64-encoded NIfTI file.
What is Microsoft MedImageParse 3D used for?
Microsoft MedImageParse 3D is used to segment complete three-dimensional CT and MRI studies based on a natural-language prompt. It can identify organs, lesions, or other regions of interest and produce a reusable 3D segmentation mask for research, annotation, volumetric measurement, and imaging pipeline development.


Sources 5
Provider

About Microsoft Copilot