What is Florence-2-large?
Florence-2-large is an open-weight vision-language model from Microsoft. It accepts an image together with a text task prompt and generates a textual or structured-text representation of the requested result. Depending on the prompt, that result may be a caption, object labels, bounding boxes, region descriptions or recognized text with coordinates.
The model contains approximately 0.77 billion parameters. That is relatively compact for a foundation model covering multiple vision tasks, which is important for developers who want to run inference locally or process large image collections without relying on a managed multimodal API.
Florence-2-large belongs to Microsoft's Florence-2 model family and is distributed through the Microsoft Hugging Face repository. It is distinct from the Florence-2-large-ft variant mentioned in the repository ecosystem; the large model is the pretrained open-weight checkpoint described here, while the ft name refers to a downstream fine-tuned variant.
What can Florence-2-large do?
Florence-2-large uses task tokens and prompts to select the kind of image understanding required. Common examples include <CAPTION> for a basic caption, <OD> for object detection and <OCR> for optical character recognition. The processor converts the generated sequence into more useful task-specific structures where appropriate.
- Image captioning: Generates captions, detailed captions and more detailed descriptions.
- Object detection: Identifies objects and returns labels with bounding boxes.
- Dense-region captioning: Describes multiple regions of an image rather than producing only one global caption.
- Visual grounding: Connects phrases or descriptions to image regions, supporting caption-to-phrase grounding and related annotation workflows.
- Region proposals: Produces candidate regions for downstream computer-vision processing.
- OCR: Extracts visible text from images.
- OCR with regions: Returns recognized text alongside quadrilateral coordinates describing where the text appears.
These functions share one model and a common prompt-based interface, but they are not identical to a general-purpose chat experience. The developer normally selects the task, supplies the image, decodes the output and applies the processor's post-processing for coordinates, labels or other structured results.
Model design and training background
Florence-2 uses a unified sequence-to-sequence architecture rather than a separate specialist model for every supported task. In practical terms, the model is trained to represent an image and generate different textual encodings of visual information depending on the task prompt.
Microsoft reports that the Florence-2 family was trained with the FLD-5B dataset, containing approximately 5.4 billion visual annotations across 126 million images. This training approach is intended to give the model broad coverage across captioning, detection, grounding and OCR-related tasks. Those figures are provider-reported research details, not a guarantee that every domain or image style will receive equally reliable results.
The model repository is MIT licensed, according to the supplied research. That permissive license can make Florence-2-large attractive for experimentation, internal tools and self-hosted applications, subject to the developer's responsibility to review the repository, dependencies, data rights and deployment requirements.
Verified specifications at a glance
| Specification | Florence-2-large |
|---|---|
| Provider | Microsoft |
| Release | June 2024 |
| Model size | Approximately 0.77 billion parameters |
| Model type | Open-weight multimodal vision-language model |
| Inputs | Images and text task prompts |
| Outputs | Text, captions, labels, coordinates and other task-specific textual structures |
| Maximum text positions | 1,024 in the main model configuration |
| Common published generation setting | max_new_tokens=1024 |
| License | MIT |
| Hosted Microsoft API price | No official price identified for this exact model |
The 1,024-position configuration limit and the commonly published 1,024-token generation setting should not be confused with a hosted API context-window guarantee. Florence-2-large is generally loaded as a local model, and actual memory use, throughput and usable generation length also depend on the Transformers implementation, precision, hardware and decoding configuration.
How it is deployed
Florence-2-large is available from Microsoft's Hugging Face repository for local or self-hosted inference. The published implementation uses the Transformers library, AutoProcessor and a Florence-compatible causal language-model class or AutoModelForCausalLM. The repository requires trust_remote_code=True for its custom implementation.
That custom-code requirement is useful because it provides the model-specific processing needed for Florence-2 tasks, but it also creates an operational responsibility. Teams should pin a model revision, inspect the repository code and dependencies, test upgrades in isolation and avoid treating downloaded remote code as automatically trustworthy.
There is no identified official Microsoft hosted API price for Florence-2-large, and the model should not be described as a standard Copilot or Azure model endpoint with published per-token billing. The direct cost is instead associated with the hardware, hosting service or third-party platform used to run it. This can be economical for repeated batch processing, but the operator must manage installation, compute capacity, monitoring and scaling.
Main strengths and trade-offs
The model's main strength is breadth relative to its size. A single approximately 0.77-billion-parameter checkpoint can support captioning, detection, grounding and OCR instead of requiring a different model for each basic task. Its open-weight distribution also supports private or offline workflows where sending images to a hosted service is undesirable.
Its compact scale may offer a useful speed and cost trade-off for local inference and batch jobs. The supplied research rates its speed and cost favorably as editorial evaluations, not Microsoft-published scores. Actual performance depends heavily on hardware, quantization, image resolution, batch size and implementation choices, so those assessments should be treated as practical guidance rather than fixed benchmarks.
The trade-off is that Florence-2-large is specialized for image understanding rather than broad reasoning. It does not natively generate images, audio or video. It is not designed as a general chat assistant, coding model or autonomous agent, and the research identifies no tool or function-calling interface. It also does not provide a normal structured-output mode in the sense of a managed API; its outputs are generated text that the processor may parse into task-specific structures such as boxes and OCR coordinates.
Reasoning, coding and supported modalities
Florence-2-large has image input and text input, with text as its direct output modality. It can perform visual interpretation tasks, but that should not be mistaken for advanced general reasoning. Its useful reasoning is task-bounded: it can associate image content with captions, labels, regions or text, while complex multi-step analysis outside those tasks is not its intended role.
Coding is not a primary capability. A developer can write software around the model, but Florence-2-large is not a coding assistant and should not be selected for code generation. It also has no supplied evidence of native web search, external tools, streaming, batch API or function calling. Fine-tuning or downstream adaptation is possible within the model ecosystem, but the existence of the related large-ft checkpoint does not mean that every deployment includes a managed fine-tuning service.
Published performance and how to interpret it
Microsoft reports several zero-shot research results for Florence-2-large, including a COCO captioning CIDEr score of 135.6, a NoCaps validation CIDEr score of 120.8, a TextCaps validation CIDEr score of 72.8 and a COCO detection validation mAP of 37.5. The large model also outperformed the Florence-2-base variant on several reported visual-grounding and referring-expression benchmarks.
These are provider-reported research benchmarks on named datasets. They are useful for understanding the model's intended capabilities, but they are not production guarantees. Results can change with image quality, domain, prompt selection, decoding settings and post-processing. A document OCR workflow, retail inventory system or specialized scientific dataset should be tested with representative images before deployment.
Limitations to consider
Florence-2-large may produce incorrect captions, missed detections, inaccurate coordinates or OCR errors. Such failures are especially important when images are low resolution, text is stylized, objects are partially hidden or the deployment domain differs from the model's evaluation data.
The generated representations also require task-specific interpretation. Bounding boxes and quadrilateral coordinates are not automatically a complete application-level detection or document-processing system. Developers may need validation, confidence handling, coordinate conversion and rules for rejecting malformed or implausible outputs.
Because the repository uses custom model code, compatibility and supply-chain review matter. Pinning revisions and testing the exact processor and model versions can reduce surprises. Human review or additional validation is appropriate for safety-critical, identity-sensitive, legal or high-accuracy OCR and detection workflows.
When to choose Florence-2-large
Choose Florence-2-large when you need a compact, open-weight image-understanding model and are prepared to operate it locally or through a compatible third-party service. It is particularly suitable for:
- Generating captions and metadata for image collections.
- Annotating datasets with object boxes, labels or regions.
- Preprocessing document images with OCR and text locations.
- Building visual search or inventory pipelines.
- Running private or offline computer-vision inference.
- Using one model for several related vision tasks rather than deploying multiple specialist checkpoints.
Another option may be more appropriate if the project requires image, audio or video generation, a conversational assistant, advanced general reasoning, code generation, built-in tools, managed scaling or a provider-backed service-level agreement. A larger or more specialized vision system may also be preferable when accuracy on a narrow production domain matters more than model size and local operating cost. Florence-2-large is best understood as a practical multi-task image-analysis component, not as a replacement for a general multimodal assistant.

