Florence-2

Florence-2-large

by Microsoft Copilot · Current open-weight model; publicly accessible through Hugging Face for local or self-hosted inference.

Microsoft Florence-2-large is a compact MIT-licensed open-weight vision-language model for image captioning, object detection, visual grounding, region annotation and OCR. It uses image inputs and text task prompts, can run locally through Transformers, and has no identified official hosted Microsoft API price. Its strengths are multi-task coverage, local deployment and relatively small size; its limitations include task-bounded reasoning, text-only output and the need to manage custom code and post-processing.

Text Reasoning Coding
Microsoft Florence-2-large is a compact open-weight model that turns images and task prompts into captions, labels, coordinates and OCR results. It brings several common computer-vision tasks into one sequence-to-sequence system, making it a practical choice for local image analysis, dataset annotation and visual search preprocessing.
Outputs

What Florence-2-large can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

2/10 Reasoning
1/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Florence-2
Model type Multimodal
Context window 1K tokens
Maximum output 1K tokens
Release date 2024-06
Status Current open-weight model; publicly accessible through Hugging Face for local or self-hosted inference.
Knowledge cutoff notes

Microsoft and the model repository do not publish a separate knowledge-cutoff date for Florence-2-large. The model is an image-understanding system rather than a general web-trained conversational model.

Model notes

Florence-2-large is the approximately 0.77B-parameter pretrained model in the Florence-2 family; it is distinct from Florence-2-large-ft, which is fine-tuned on downstream tasks. The model uses text task prompts such as <CAPTION>, <OD>, <OCR> and <OCR_WITH_REGION>. Microsoft reports training with the FLD-5B dataset containing approximately 5.4 billion annotations across 126 million images. The repository is MIT licensed and uses custom Florence-2 model code. Published examples commonly use max_new_tokens=1024, while the main model configuration specifies 1024 maximum text positions. No official Microsoft hosted API price was identified for this exact model.

Model guide

Microsoft Florence-2-large: Compact Open Vision Model for OCR, Detection and Captioning

Florence-2-large is Microsoft's approximately 0.77-billion-parameter open-weight vision-language model for image captioning, object detection, visual grounding, region captioning and OCR. Its unified text-prompt interface makes it useful for local and private computer-vision pipelines, although it is not a general conversational model or a hosted Microsoft API product.

What is Florence-2-large?

Florence-2-large is an open-weight vision-language model from Microsoft. It accepts an image together with a text task prompt and generates a textual or structured-text representation of the requested result. Depending on the prompt, that result may be a caption, object labels, bounding boxes, region descriptions or recognized text with coordinates.

The model contains approximately 0.77 billion parameters. That is relatively compact for a foundation model covering multiple vision tasks, which is important for developers who want to run inference locally or process large image collections without relying on a managed multimodal API.

Florence-2-large belongs to Microsoft's Florence-2 model family and is distributed through the Microsoft Hugging Face repository. It is distinct from the Florence-2-large-ft variant mentioned in the repository ecosystem; the large model is the pretrained open-weight checkpoint described here, while the ft name refers to a downstream fine-tuned variant.

What can Florence-2-large do?

Florence-2-large uses task tokens and prompts to select the kind of image understanding required. Common examples include <CAPTION> for a basic caption, <OD> for object detection and <OCR> for optical character recognition. The processor converts the generated sequence into more useful task-specific structures where appropriate.

  • Image captioning: Generates captions, detailed captions and more detailed descriptions.
  • Object detection: Identifies objects and returns labels with bounding boxes.
  • Dense-region captioning: Describes multiple regions of an image rather than producing only one global caption.
  • Visual grounding: Connects phrases or descriptions to image regions, supporting caption-to-phrase grounding and related annotation workflows.
  • Region proposals: Produces candidate regions for downstream computer-vision processing.
  • OCR: Extracts visible text from images.
  • OCR with regions: Returns recognized text alongside quadrilateral coordinates describing where the text appears.

These functions share one model and a common prompt-based interface, but they are not identical to a general-purpose chat experience. The developer normally selects the task, supplies the image, decodes the output and applies the processor's post-processing for coordinates, labels or other structured results.

Model design and training background

Florence-2 uses a unified sequence-to-sequence architecture rather than a separate specialist model for every supported task. In practical terms, the model is trained to represent an image and generate different textual encodings of visual information depending on the task prompt.

Microsoft reports that the Florence-2 family was trained with the FLD-5B dataset, containing approximately 5.4 billion visual annotations across 126 million images. This training approach is intended to give the model broad coverage across captioning, detection, grounding and OCR-related tasks. Those figures are provider-reported research details, not a guarantee that every domain or image style will receive equally reliable results.

The model repository is MIT licensed, according to the supplied research. That permissive license can make Florence-2-large attractive for experimentation, internal tools and self-hosted applications, subject to the developer's responsibility to review the repository, dependencies, data rights and deployment requirements.

Verified specifications at a glance

SpecificationFlorence-2-large
ProviderMicrosoft
ReleaseJune 2024
Model sizeApproximately 0.77 billion parameters
Model typeOpen-weight multimodal vision-language model
InputsImages and text task prompts
OutputsText, captions, labels, coordinates and other task-specific textual structures
Maximum text positions1,024 in the main model configuration
Common published generation settingmax_new_tokens=1024
LicenseMIT
Hosted Microsoft API priceNo official price identified for this exact model

The 1,024-position configuration limit and the commonly published 1,024-token generation setting should not be confused with a hosted API context-window guarantee. Florence-2-large is generally loaded as a local model, and actual memory use, throughput and usable generation length also depend on the Transformers implementation, precision, hardware and decoding configuration.

How it is deployed

Florence-2-large is available from Microsoft's Hugging Face repository for local or self-hosted inference. The published implementation uses the Transformers library, AutoProcessor and a Florence-compatible causal language-model class or AutoModelForCausalLM. The repository requires trust_remote_code=True for its custom implementation.

That custom-code requirement is useful because it provides the model-specific processing needed for Florence-2 tasks, but it also creates an operational responsibility. Teams should pin a model revision, inspect the repository code and dependencies, test upgrades in isolation and avoid treating downloaded remote code as automatically trustworthy.

There is no identified official Microsoft hosted API price for Florence-2-large, and the model should not be described as a standard Copilot or Azure model endpoint with published per-token billing. The direct cost is instead associated with the hardware, hosting service or third-party platform used to run it. This can be economical for repeated batch processing, but the operator must manage installation, compute capacity, monitoring and scaling.

Main strengths and trade-offs

The model's main strength is breadth relative to its size. A single approximately 0.77-billion-parameter checkpoint can support captioning, detection, grounding and OCR instead of requiring a different model for each basic task. Its open-weight distribution also supports private or offline workflows where sending images to a hosted service is undesirable.

Its compact scale may offer a useful speed and cost trade-off for local inference and batch jobs. The supplied research rates its speed and cost favorably as editorial evaluations, not Microsoft-published scores. Actual performance depends heavily on hardware, quantization, image resolution, batch size and implementation choices, so those assessments should be treated as practical guidance rather than fixed benchmarks.

The trade-off is that Florence-2-large is specialized for image understanding rather than broad reasoning. It does not natively generate images, audio or video. It is not designed as a general chat assistant, coding model or autonomous agent, and the research identifies no tool or function-calling interface. It also does not provide a normal structured-output mode in the sense of a managed API; its outputs are generated text that the processor may parse into task-specific structures such as boxes and OCR coordinates.

Reasoning, coding and supported modalities

Florence-2-large has image input and text input, with text as its direct output modality. It can perform visual interpretation tasks, but that should not be mistaken for advanced general reasoning. Its useful reasoning is task-bounded: it can associate image content with captions, labels, regions or text, while complex multi-step analysis outside those tasks is not its intended role.

Coding is not a primary capability. A developer can write software around the model, but Florence-2-large is not a coding assistant and should not be selected for code generation. It also has no supplied evidence of native web search, external tools, streaming, batch API or function calling. Fine-tuning or downstream adaptation is possible within the model ecosystem, but the existence of the related large-ft checkpoint does not mean that every deployment includes a managed fine-tuning service.

Published performance and how to interpret it

Microsoft reports several zero-shot research results for Florence-2-large, including a COCO captioning CIDEr score of 135.6, a NoCaps validation CIDEr score of 120.8, a TextCaps validation CIDEr score of 72.8 and a COCO detection validation mAP of 37.5. The large model also outperformed the Florence-2-base variant on several reported visual-grounding and referring-expression benchmarks.

These are provider-reported research benchmarks on named datasets. They are useful for understanding the model's intended capabilities, but they are not production guarantees. Results can change with image quality, domain, prompt selection, decoding settings and post-processing. A document OCR workflow, retail inventory system or specialized scientific dataset should be tested with representative images before deployment.

Limitations to consider

Florence-2-large may produce incorrect captions, missed detections, inaccurate coordinates or OCR errors. Such failures are especially important when images are low resolution, text is stylized, objects are partially hidden or the deployment domain differs from the model's evaluation data.

The generated representations also require task-specific interpretation. Bounding boxes and quadrilateral coordinates are not automatically a complete application-level detection or document-processing system. Developers may need validation, confidence handling, coordinate conversion and rules for rejecting malformed or implausible outputs.

Because the repository uses custom model code, compatibility and supply-chain review matter. Pinning revisions and testing the exact processor and model versions can reduce surprises. Human review or additional validation is appropriate for safety-critical, identity-sensitive, legal or high-accuracy OCR and detection workflows.

When to choose Florence-2-large

Choose Florence-2-large when you need a compact, open-weight image-understanding model and are prepared to operate it locally or through a compatible third-party service. It is particularly suitable for:

  • Generating captions and metadata for image collections.
  • Annotating datasets with object boxes, labels or regions.
  • Preprocessing document images with OCR and text locations.
  • Building visual search or inventory pipelines.
  • Running private or offline computer-vision inference.
  • Using one model for several related vision tasks rather than deploying multiple specialist checkpoints.

Another option may be more appropriate if the project requires image, audio or video generation, a conversational assistant, advanced general reasoning, code generation, built-in tools, managed scaling or a provider-backed service-level agreement. A larger or more specialized vision system may also be preferable when accuracy on a narrow production domain matters more than model size and local operating cost. Florence-2-large is best understood as a practical multi-task image-analysis component, not as a replacement for a general multimodal assistant.


Answers to Frequently Asked Questions

What are the main limitations of Florence-2-large?
Florence-2-large is designed for task-specific image understanding rather than general chat, advanced reasoning, coding or autonomous agents. It can produce incorrect captions, missed detections, inaccurate coordinates and OCR errors, especially with low-quality or domain-specific images. Production workflows should include validation, confidence handling and human review where accuracy or safety is critical.
Is Florence-2-large available through a paid Microsoft API?
No official Microsoft hosted API price was identified for this exact model. Florence-2-large is generally run locally, on self-managed infrastructure or through a compatible third-party platform, so costs depend on hardware, hosting, compute capacity and operational requirements.
How is Florence-2-large deployed and accessed?
The model is available through Microsoft's Hugging Face repository for local or self-hosted inference. It is typically used with Transformers, AutoProcessor and a Florence-compatible causal language-model class, and its implementation requires trust_remote_code=True. Teams should review the custom code, pin model revisions and test dependencies before deployment.
What is Microsoft Florence-2-large?
Florence-2-large is an open-weight vision-language model from Microsoft with approximately 0.77 billion parameters. It accepts images and text task prompts, then generates captions, object labels, bounding boxes, region descriptions, OCR text or other task-specific outputs.
What tasks can Florence-2-large perform?
Florence-2-large supports image captioning, detailed descriptions, object detection, dense-region captioning, visual grounding, region proposals, OCR and OCR with quadrilateral text coordinates. Developers select these tasks with prompts or task tokens such as , and .


Sources 4
Provider

About Microsoft Copilot