What is North Micro Vision Instruct?
North Micro Vision Instruct is Cohere’s 2.4-billion-parameter open-weight vision-language model. A vision-language model processes visual information and language in the same prompt: for example, a user can provide an image of a form followed by a question asking for the customer’s address. The model then produces a text answer based on the image and the surrounding instructions.
The model is officially presented as North Micro Vision by Cohere and is released on Hugging Face under the repository name North-Micro-Vision-Instruct. It is licensed under Apache 2.0, which makes the checkpoint suitable for research, experimentation, customization, and private deployment subject to the license terms.
North Micro Vision Instruct is not a consumer chatbot subscription and is not documented as a metered Cohere hosted generation model with public per-token pricing. Its primary role is an open-weight model that developers and organizations can run, evaluate, or fine-tune using their own infrastructure and inference stack.
Primary purpose and position in Cohere’s lineup
Within Cohere’s model ecosystem, North Micro Vision Instruct addresses multimodal understanding rather than general-purpose text generation alone. Its intended tasks include reading and interpreting visual content, extracting information from documents, answering questions about images, and combining several images with textual instructions.
The model occupies a smaller, more customizable position than a large hosted multimodal service. Its 2.4B parameter count is useful for teams looking to experiment with a vision-language checkpoint without starting with a much larger model. That compactness can help with local testing and specialized fine-tuning, although actual hardware and latency depend on factors such as numerical precision, image resolution, batching, quantization, and inference software.
North Micro Vision Instruct should also be distinguished from Cohere Parse. Parse is a separate document-parsing service that uses a proprietary North Micro Vision architecture. The open-weight North Micro Vision Instruct checkpoint and the Parse endpoint are related in architecture, but they are different offerings with different deployment and access models.
Inputs, outputs, and supported modalities
The model accepts interleaved text and images. Interleaving means that a prompt can place written instructions and visual inputs together rather than treating the image as an isolated attachment. It also supports multiple images, which enables tasks such as comparing two pages, examining several product photographs, or relating a chart to a screenshot and a written question.
Its native output is text. North Micro Vision Instruct does not natively generate images, audio, video, music, or other visual media. It can describe or analyze an image, but it cannot return a newly generated image as the answer.
| Capability | Supported or documented behavior |
|---|---|
| Text input | Yes |
| Image input | Yes |
| Multiple images | Yes |
| Text output | Yes |
| Image, audio, or video output | No native output documented |
| License | Apache 2.0 |
| Parameter count | 2.4B |
The supplied documentation does not identify a separate native tool-use, function-calling, web-search, or code-execution interface for this checkpoint. Those capabilities should not be assumed merely because an application can place the model inside a larger workflow.
What it can understand
North Micro Vision Instruct is designed for visual question answering, image captioning, OCR, chart and table understanding, document analysis, and visual grounding. Visual grounding refers to connecting an answer to a particular region or element in an image, such as identifying which line in a diagram supports an answer.
Its native-resolution processing is intended to preserve image proportions and fine visual detail. This matters for documents, forms, screenshots, tables, and charts, where resizing an image too aggressively can remove small text or alter spatial relationships. Developers should still test the model on their own materials because real-world performance can vary with scan quality, handwriting, layout complexity, typography, and language.
The model is described as supporting multilingual visual understanding across languages including English, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Chinese, and Arabic. This makes it relevant to multilingual document workflows and image-based questions where the text in the image and the user’s question may not use the same language.
Architecture and context limits
The model combines a custom-trained 400-million-parameter native-resolution vision encoder, a projector that connects visual representations to the language model, and an approximately 2-billion-parameter language backbone called North Micro LLM. The language backbone follows the Command A+ architectural approach and uses a hybrid attention design.
The language backbone has a stated 128K-token context window. A context window is the amount of text and other represented prompt information the model can process in one interaction. However, the model documentation identifies up to 8K tokens as the validated operating range for multimodal prompts. This distinction is important: the 128K figure describes the language backbone, while 8K is the practical documented range for prompts that combine images and text.
For document workloads, users should therefore avoid assuming that a very long collection of images and extracted text will work reliably simply because the language backbone has a large theoretical context. Long-document applications may need page selection, chunking, staged extraction, or external retrieval. The research does not provide a maximum output-token limit for this model.
Deployment, customization, and pricing
The open-weight release is intended for local and private deployment rather than use through a documented Cohere hosted endpoint with a published token price. The checkpoint is available through Cohere Labs on Hugging Face and is documented for use with the Transformers ecosystem.
There is no verified official Cohere input or output price for this exact open-weight model in the supplied research. That does not mean running it is cost-free: organizations still need to account for hardware, storage, inference software, engineering work, and operational maintenance. The main economic advantage is control over deployment and the possibility of adapting the model to a specialized task, not a guaranteed zero operating cost.
Fine-tuning is identified as a supported use case. A team could use it as a starting point for a domain-specific visual system, such as extracting fields from a particular class of forms or answering questions about specialized diagrams. The feasibility and cost of that work depend on the available training data, hardware, precision, and fine-tuning method; the supplied research does not specify a required hardware configuration.
Main strengths and trade-offs
- Compact open-weight design: The 2.4B parameter scale is more approachable for experimentation than much larger vision-language checkpoints, although it still requires suitable hardware and software.
- Native-resolution vision: Preserving image dimensions can help with small text, document structure, charts, and screenshots.
- Multilingual support: The documented language coverage extends beyond English and includes several European, Asian, and Middle Eastern languages.
- Multi-image prompts: Several images can be supplied when comparison or combined visual context is needed.
- Customization and deployment control: Apache 2.0 licensing and open weights support private deployment and task-specific experimentation.
- Important context limitation: The 128K language-backbone context should not be treated as a validated 128K multimodal context. The documented validated multimodal range is up to 8K tokens.
- No media generation: It returns text and is not a model for creating images, audio, or video.
Best use cases
North Micro Vision Instruct is a good fit when the central problem is understanding visual information and the team values an open checkpoint or private deployment. Suitable examples include:
- Extracting fields from forms, invoices, reports, and other documents.
- OCR and visual information extraction from scans or screenshots.
- Answering questions about charts, tables, diagrams, and images.
- Comparing multiple images or document pages in one prompt.
- Generating text captions and descriptions for images.
- Building multilingual visual question-answering workflows.
- Fine-tuning a compact multimodal model for a specialized domain.
- Research or private deployments where sending images to a hosted service is undesirable or not permitted.
When to choose this model
Choose North Micro Vision Instruct when you need a compact, customizable vision-language model for image understanding rather than a hosted consumer assistant or a media-generation system. It is especially attractive when open weights, Apache 2.0 licensing, local execution, or private infrastructure are important. Its documented multilingual and multi-image behavior also suits applications that must inspect varied visual inputs.
A hosted multimodal service may be more appropriate when the priority is a managed API, predictable provider-side scaling, a documented production SLA, or minimal infrastructure work. Cohere Parse may be more suitable when the specific requirement is a managed document-parsing service rather than operating and customizing an open-weight checkpoint. A larger vision-language model may be preferable when evaluations show that the compact model cannot handle the required visual complexity, long multimodal prompts, or domain-specific accuracy target.
Conversely, North Micro Vision Instruct may be a better choice than a larger model when the task is narrow, the deployment environment is private or resource-constrained, and the team can validate or fine-tune the checkpoint. The research does not provide standardized benchmark results, so these trade-offs should be confirmed with representative documents, languages, image resolutions, and failure cases.
Limitations and evaluation guidance
The most important implementation constraint is the difference between the theoretical language context and the validated multimodal context. Applications should design around the 8K-token multimodal range unless Cohere documents a broader validated limit. They should also test how image resolution, batching, quantization, and checkpoint precision affect speed, memory use, and answer quality.
The supplied information does not establish a maximum output length, hosted per-token price, tool-calling interface, structured-output guarantee, or benchmark score. These should be treated as unknown rather than inferred from the model’s architecture or its availability through general machine-learning tooling.
Overall, North Micro Vision Instruct is best understood as a compact open-weight model for multilingual visual understanding, document analysis, and customization. Its value comes from combining useful image capabilities with deployment control, while its practical limits include the 8K validated multimodal context, text-only output, and the operational work required to run and evaluate an open checkpoint.

