Yi-VL

Yi-VL-34B

by 01.AI · Available as open-weight downloadable model; no current first-party hosted API availability verified

Yi-VL-34B is a 34-billion-parameter open-weight vision-language model from 01.AI for English and Chinese image understanding. It accepts text and images, produces text responses, supports visual question answering and image text recognition, and is intended for self-hosted or research use. Its key trade-offs are a 4,096-token context limit, text-only output, substantial recommended GPU hardware, and no verified current hosted API pricing.

Text Reasoning Coding
Yi-VL-34B extends 01.AI’s Yi language-model family with visual understanding. It is designed for organizations and researchers that want to run a downloadable vision-language model on their own hardware rather than use a clearly documented hosted API. Its main strengths are bilingual image understanding, image-based question answering, and open-weight deployment flexibility. Its main trade-offs are a substantial multi-GPU hardware requirement, a 4,096-token context limit, text-only output, and the absence of verified current hosted pricing or a first-party API offering for this exact model.
Outputs

What Yi-VL-34B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

6/10 Reasoning
5/10 Coding
3/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Yi-VL
Model type Multimodal
Context window 4K tokens
Release date 2024-01-23
Status Available as open-weight downloadable model; no current first-party hosted API availability verified
Knowledge cutoff notes

No explicit model knowledge-cutoff date was found. The model documentation says that reported MMMU and CMMMU benchmark results used data available up to January 2024, but that statement does not establish the model's training-data cutoff.

Model notes

Yi-VL-34B is an open-weight multimodal model based on a LLaVA-style architecture. The official model card describes text and image inputs, text output, English and Chinese support, multi-round conversations with one image, and 448x448 image understanding. Its configuration lists max_position_embeddings of 4096. The model uses a CLIP ViT-H/14 vision encoder and a Yi-34B-Chat language backbone. 01.AI reported historical first-place open-source results on MMMU and CMMMU using benchmark data available up to January 2024; this is not a verified knowledge-cutoff date. The model card recommends four RTX 4090 GPUs or an A800 80 GB for inference. The model is governed by the Yi Series Models Community License Agreement 2.1, and no exact current hosted API price was verified.

Model guide

Yi-VL-34B: Open-Weight Bilingual Vision Model for Self-Hosted Image Understanding

Yi-VL-34B is an open-weight 34-billion-parameter vision-language model from 01.AI. It accepts text and images and produces text responses, with a focus on bilingual English and Chinese visual question answering, image text recognition, image comprehension, and multi-round conversations about a single image.

What is Yi-VL-34B?

Yi-VL-34B is an open-weight multimodal vision-language model released by 01.AI on January 23, 2024. The model accepts both text and images as input and generates text responses. In practical terms, a user can provide an image with a question such as “What information appears in this document?” or “Describe the objects in this scene,” and the model will respond in English or Chinese.

The model is part of 01.AI’s Yi family, but it is not simply a text-only language model with an image upload feature added at the application layer. Yi-VL-34B combines a language model with a vision encoder so that visual information can be converted into representations the language model can interpret. The result is intended for visual question answering, image description, image text recognition, information extraction, and multi-round conversations about one image.

Its published weights make it primarily a self-hosted or research-oriented model. The supplied documentation does not verify a current first-party hosted API product or per-token pricing for Yi-VL-34B.

Where Yi-VL-34B fits in 01.AI’s lineup

01.AI develops the Yi family of foundation models and also offers broader enterprise AI, agent, model deployment, fine-tuning, and private-deployment services. Yi-VL-34B occupies the vision-language part of that ecosystem. Its role is narrower than a general consumer assistant: it is a downloadable model for interpreting images and answering questions with text.

The model should therefore be evaluated as an open-weight vision-language model rather than as a complete chatbot platform. The available materials do not establish integrated web search, persistent memory, a consumer subscription plan, native image generation, audio or video generation, or a current hosted service dedicated to this model.

Supported inputs and outputs

Yi-VL-34B supports two input types:

  • Text: Questions, instructions, descriptions, and follow-up messages.
  • Images: Visual content supplied for analysis, question answering, recognition, or information extraction.

Its output is text. The model does not generate images, audio, video, speech, music, embeddings, or other separately documented media outputs. This distinction matters because “multimodal” describes its ability to process more than one input type; it does not mean that the model produces every type of media.

The model documentation describes English and Chinese interaction, image text recognition, visual comprehension, and multi-round text-image conversations involving one image. The published image-understanding resolution is 448 by 448 pixels. The supplied research does not verify support for multiple images in a single conversation, video input, audio input, or document-specific parsing beyond the model’s general image and text capabilities.

Architecture and context limit

Yi-VL-34B uses a LLaVA-style architecture. A CLIP ViT-H/14 vision transformer encodes the image, a projection module aligns the visual features with the language representation, and a Yi-34B-Chat language model generates the response. This design lets the language component reason over information extracted from an image while continuing to handle ordinary text prompts.

The published configuration lists 60 transformer layers, a hidden size of 7,168, and a maximum position length of 4,096 tokens. The 4,096-token figure is the documented context limit, not a guaranteed number of words or characters. In a vision-language interaction, the usable room for conversation can also be affected by how the image is represented internally and by the length of the user’s prompts and prior messages.

No maximum output-token value is verified in the supplied model information. Users should not assume that the model has a separately documented long-form generation limit beyond the stated context configuration.

What Yi-VL-34B is good at

Visual question answering

The model’s central use case is answering questions about an image. For example, it can be used to ask about visible objects, scene content, basic relationships between elements, or information presented in an image. This makes it relevant to image inspection tools, visual search prototypes, educational applications, and research workflows that need a text explanation of visual content.

Image text recognition and extraction

Yi-VL-34B is also documented for image text recognition. A user can submit an image containing text and ask the model to read, summarize, translate, or extract relevant information. Results should still be checked when the source image is blurry, densely formatted, unusually styled, or important for compliance or business decisions. The available materials establish the capability category but do not provide a guaranteed accuracy rate for every document type.

English and Chinese conversations

English and Chinese support is one of the model’s clearest differentiators. It can be used for bilingual visual question answering and text-image conversations, which may be useful when an image contains information that needs to be discussed in either language. The research does not establish equal performance across all dialects, domains, image types, or translation tasks, so bilingual support should not be interpreted as a guarantee of professional translation quality.

Multi-round discussion of one image

The model card describes multi-round conversations with one image. This allows a user to ask an initial question and then follow up for clarification, additional extraction, or a different explanation without submitting a completely unrelated request each time. The documented limitation is important: the supplied research does not confirm a general multi-image conversation workflow.

Performance, hardware, and speed trade-offs

01.AI reported that Yi-VL-34B ranked first among existing open-source models on MMMU and CMMMU using benchmark data available up to January 2024. These are historical provider-reported evaluation claims. They describe the model’s position against the comparison set used at that time and should not be treated as a current leaderboard ranking or a guarantee of performance on a particular application.

The model card recommends substantial inference hardware: four RTX 4090 GPUs or an A800 with 80 GB of memory. That requirement is a major practical consideration. Yi-VL-34B is not positioned as a lightweight model for an ordinary laptop or a small single-GPU deployment. Its 34-billion-parameter size and multimodal architecture can provide a useful foundation for organizations with suitable infrastructure, but the hardware footprint can increase setup cost and reduce deployment flexibility.

In the supplied editorial assessment, the model receives a speed score of 3 out of 10, a cost score of 8 out of 10, a reasoning score of 6 out of 10, and a coding score of 5 out of 10. These are editorial evaluations, not scores published by 01.AI. The cost score reflects the attractiveness of downloadable weights when an organization already has hardware, not a verified hosted price. The model is not primarily a coding model, and its coding score should not be used as evidence of a provider-tested coding benchmark.

API access, tools, and pricing

No current first-party hosted API availability is verified for Yi-VL-34B. The model is available as downloadable open weights, and the supplied sources do not provide a current per-input-token price, per-output-token price, subscription price, or guaranteed managed endpoint for this exact model.

Tool calling and function calling are not verified. Structured-output enforcement, streaming, caching, batch processing, and fine-tuning support are also not established by the supplied research. These features may be important in production applications, so teams should not infer their availability from the model’s ability to produce text or participate in multi-round conversations.

The Yi Series Models Community License Agreement 2.1 governs the model according to the supplied model information. Organizations should review that license and any applicable deployment obligations before using the weights commercially or redistributing a system built with them.

Strengths and limitations

Strengths include:

  • Open-weight distribution that supports self-hosted experimentation and deployment.
  • Text and image understanding in a single model.
  • English and Chinese visual interaction.
  • Support for visual question answering, image text recognition, and information extraction.
  • Multi-round discussion about one image.
  • A documented architecture and configuration that can help technical teams evaluate deployment requirements.

Limitations include:

  • The recommended hardware is substantial, with four RTX 4090 GPUs or an A800 80 GB cited by the model card.
  • The maximum position length is 4,096 tokens.
  • Output is text only; the model is not an image, audio, video, or speech generator.
  • No current hosted API, pricing, or maximum output-token limit is verified for this exact model.
  • Tool use, structured outputs, streaming, batch processing, caching, and fine-tuning are not confirmed.
  • Historical benchmark claims may not represent performance against newer vision-language models.

When to choose Yi-VL-34B

Yi-VL-34B is a reasonable choice when the main requirement is self-hosted bilingual image understanding and the organization has enough GPU capacity to operate a large open-weight model. It may fit research projects, private visual-document workflows, image question-answering systems, and prototypes where keeping inference under the organization’s control is more important than using a managed API.

It is less suitable when the priority is low-cost or low-latency inference, easy access through a hosted endpoint, long context, guaranteed structured output, native tool calling, or a complete consumer assistant experience. A smaller vision-language model may be more appropriate when hardware and response speed are the main constraints. A newer managed multimodal model may be preferable when a team needs documented API pricing, production support, current benchmark performance, or integrated tools. A dedicated OCR system may be a better fit for high-volume document extraction when reliable text recognition and layout handling matter more than open-ended visual conversation.

Overall, Yi-VL-34B’s appeal comes from its combination of downloadable weights, bilingual visual understanding, and a clearly described vision-language architecture. Its practical value depends heavily on whether the user can justify the hardware and operational effort required to run it. It should be selected for controlled image-understanding workloads rather than treated as a general replacement for hosted multimodal assistants or specialized media-processing systems.


Answers to Frequently Asked Questions

What are the main limitations of Yi-VL-34B?
Yi-VL-34B has a documented maximum context length of 4,096 tokens, produces text only, and requires substantial hardware. Support for multiple images, video, audio, tool calling, structured outputs, streaming, batch processing, caching, fine-tuning, and a maximum output-token limit is not confirmed.
Does Yi-VL-34B have a hosted API or published pricing?
No current first-party hosted API or pricing for Yi-VL-34B is verified in the available information. The model is distributed as downloadable open weights for self-hosted experimentation and deployment.
What hardware is required to run Yi-VL-34B?
The model card recommends substantial inference hardware, such as four RTX 4090 GPUs or an A800 GPU with 80 GB of memory. Its 34-billion-parameter size makes it unsuitable for most ordinary laptops and small single-GPU deployments.
What is Yi-VL-34B used for?
Yi-VL-34B is an open-weight vision-language model for analyzing images and generating text responses. It supports visual question answering, image description, image text recognition, information extraction, and multi-round conversations about one image.
Does Yi-VL-34B support English and Chinese?
Yes. Yi-VL-34B is designed for English and Chinese text-image interactions, including bilingual visual question answering and discussions about image content. However, equal performance across all domains, dialects, and translation tasks is not guaranteed.


Sources 6
Provider

About 01.AI