Yi-1.5

Yi-1.5-34B-Chat

by 01.AI · Open-weight model remains available for self-hosted deployment; 01.AI hosted model-platform API service ended on 2026-09-03

Yi-1.5-34B-Chat is a 34-billion-parameter Apache 2.0 language model from 01.AI for English-Chinese conversation, text generation, coding, mathematics, and reasoning. Its 4,096-token standard checkpoint remains available as open weights for local deployment and fine-tuning, but it requires substantial infrastructure and does not provide native multimodal input, web search, or a verified current first-party API price.

Text Reasoning Coding
Yi-1.5-34B-Chat is the conversational version of 01.AI's Yi-1.5 34B model. Released on May 13, 2024, it combines a large 34-billion-parameter architecture with instruction tuning for dialogue and general text tasks. Its standard checkpoint uses a 4,096-token context window, BF16 safetensors, and an Apache 2.0 license. Today, its main practical appeal is control: developers can download the weights, run the model locally or on private infrastructure, and adapt it with compatible open-source tools instead of depending on a first-party hosted API.
Outputs

What Yi-1.5-34B-Chat can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
4/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Yi-1.5
Model type General Purpose
Context window 4K tokens
Release date 2024-05-13
Status Open-weight model remains available for self-hosted deployment; 01.AI hosted model-platform API service ended on 2026-09-03
Shutdown date 2026-09-03
Knowledge cutoff notes

01.AI's official Yi-1.5 model documentation does not provide a directly verified knowledge-cutoff date for this exact chat checkpoint. Some third-party model directories report March 31, 2024, but that date was not confirmed in a first-party source and is therefore not used as the structured value.

Model notes

Yi-1.5-34B-Chat is the standard 4K chat checkpoint, distinct from Yi-1.5-34B-Chat-16K. The official model card reports approximately 34B parameters and BF16 safetensors. 01.AI describes Yi-1.5 as continuously pretrained with about 500B additional tokens and fine-tuned on about 3M samples. The exact model is open-weight and can be served with Transformers, vLLM, SGLang, Ollama, or similar tooling. Fine-tuning is supported through external frameworks rather than a documented first-party managed fine-tuning endpoint. Editorial scores are comparative estimates, not vendor-published ratings. No exact current first-party hosted token price or maximum output limit was verified. The 01.AI hosted model platform announced that model experience and API services stopped on September 3, 2026, but the model weights remain available for self-hosting.

Model guide

Yi-1.5-34B-Chat: Open-Weight Bilingual Model for Self-Hosted AI

Yi-1.5-34B-Chat is a 34-billion-parameter, Apache 2.0-licensed instruction-tuned language model from 01.AI. It is designed for bilingual English-Chinese conversation, text generation, coding, mathematics, reasoning, reading comprehension, and local deployment. The model remains available as downloadable open weights, while 01.AI's hosted model-platform API service ended on September 3, 2026.

What is Yi-1.5-34B-Chat?

Yi-1.5-34B-Chat is an instruction-tuned causal language model from 01.AI. In practical terms, it is a text model trained to respond to prompts and conversations rather than merely continue raw text. The chat tuning makes it suitable for assistants, question answering, coding help, educational software, document processing, and other applications that need dialogue-style responses.

The model belongs to the Yi-1.5 family, which also includes smaller 6B and 9B variants. The 34B version is the largest standard Yi-1.5 checkpoint described in the supplied research. That size gives it more capacity for difficult language and reasoning tasks than smaller local models, but it also makes deployment substantially more demanding.

01.AI released Yi-1.5-34B-Chat on May 13, 2024. The provider describes the Yi-1.5 series as having been continuously pretrained with approximately 500 billion tokens and fine-tuned with approximately 3 million diverse instruction samples. Those figures are provider-reported training details rather than independent performance measurements.

Technical specifications and limits

The standard Yi-1.5-34B-Chat checkpoint is a text-in, text-out model. It accepts written prompts and produces written responses. It does not natively process images, audio, or video, and it does not generate non-text media. This distinguishes it from multimodal models elsewhere in 01.AI's broader platform lineup.

SpecificationVerified detail
Provider01.AI
Release dateMay 13, 2024
Model familyYi-1.5
Parameter countApproximately 34 billion
Model typeInstruction-tuned causal language model
Standard context length4,096 tokens
WeightsBF16 safetensors
InputText
OutputText
LicenseApache 2.0
Maximum output tokensNot verified for this exact checkpoint

The 4,096-token context is the total working window available to the prompt and response, subject to the serving configuration. It is adequate for ordinary conversations, moderate code snippets, and shorter documents, but it is not a long-context model by current standards. The separate Yi-1.5-34B-Chat-16K checkpoint should not be confused with this standard 4K version.

No provider-verified maximum output limit was supplied for this checkpoint. The practical limit can also depend on the inference framework, memory available, and how much of the context window is occupied by the input.

What can the model do?

Yi-1.5-34B-Chat is intended for general-purpose language work, with particular relevance to English-Chinese use cases. Its documented target capabilities include instruction following, language understanding, commonsense reasoning, reading comprehension, coding, mathematics, and logical reasoning.

  • Bilingual conversation: It is designed for English and Chinese dialogue, translation-adjacent workflows, and applications serving users in both languages.
  • Coding assistance: It can generate and explain code, help interpret programming questions, and support software-development workflows. The supplied research rates its coding capability as an editorial 7 out of 10, not as a provider-published benchmark score.
  • Mathematics and reasoning: The model is positioned for mathematical computation and multi-step reasoning. It may be useful for working through explanations and intermediate steps, but the research does not establish a guaranteed accuracy level.
  • Text processing: It can support summarization, question answering, rewriting, classification through prompting, and other tasks that remain within a 4K-token text context.
  • Fine-tuning: Developers can adapt the open weights with external tools such as LLaMA-Factory, Swift, XTuner, and Firefly. This is a self-managed workflow rather than evidence of a current first-party managed fine-tuning service for this exact model.

The model does not have a documented native web-search tool, function-calling system, or provider-managed tool-use interface in the supplied specifications. An application can still place the model inside a larger software system, but retrieval, tool orchestration, validation, and execution would need to be implemented by the developer or supported by the chosen inference stack.

Deployment and current availability

Yi-1.5-34B-Chat is primarily a self-hosted model. Its official weights are available through the 01.AI Hugging Face repository, and the model can be served with Transformers, vLLM, SGLang, Ollama, or comparable compatible tools. These options give developers control over the serving environment and make it possible to keep prompts and generated responses within privately managed infrastructure.

The 34B parameter count is the main deployment trade-off. Running the official BF16 checkpoint requires high-memory GPU infrastructure compared with 6B or 9B models. Quantized community versions can reduce memory requirements, but their behavior, quality, and compatibility may differ from the official BF16 release. The supplied research does not specify a single minimum hardware configuration, so a precise hardware requirement should not be assumed.

Availability of the original hosted service should be separated from availability of the weights. The supplied service-status research records that 01.AI's hosted model experience and API services ended on September 3, 2026. As a result, new projects should treat this checkpoint as an open-weight deployment target or use a third-party host, rather than assuming that a current first-party API endpoint and token billing are available.

Pricing and API status

There is no verified current first-party token price for Yi-1.5-34B-Chat. The model weights are distributed under the Apache 2.0 license, but that does not make inference free: users still pay for GPUs, servers, electricity, storage, operations, or any third-party hosting they choose.

The model database contains no verified input price, output price, recurring subscription price, or maximum output-token allowance for this exact checkpoint. Because the 01.AI hosted platform service ended according to the supplied research, comparing it with current hosted models on a direct per-token basis is not possible without selecting a separate provider.

For a private deployment, the relevant cost comparison is infrastructure cost versus model size. A 34B model may offer stronger quality than smaller local alternatives for some language, coding, and reasoning tasks, but it will generally be slower and more expensive to operate than a compact 6B or 9B model, especially without quantization or efficient batching.

Strengths and limitations

Where it is strong

  • Open deployment: Downloadable weights allow local, private, or specialized deployments rather than requiring a permanent first-party API dependency.
  • Useful bilingual focus: English-Chinese language work is a central positioning advantage.
  • Broad text capability: The same checkpoint can support conversation, coding, mathematics, reasoning, and document-oriented tasks.
  • Adaptation flexibility: The Apache 2.0 license and compatibility with external fine-tuning tools make it suitable for research and customization, subject to the user's legal and operational review.
  • Common serving support: Transformers, vLLM, SGLang, Ollama, and similar tools provide several routes to local inference.

Where it is limited

  • Short standard context: The 4,096-token window limits long documents, extended conversations, and large codebases in a single request.
  • Text only: The checkpoint does not natively accept images, audio, or video and cannot produce those media types.
  • No verified native tools: Web search, function calling, structured-output mode, caching, and batch API support are not documented for this exact checkpoint in the supplied research.
  • Heavy infrastructure needs: A 34B BF16 model is not an easy fit for low-memory hardware. Quantization can help but introduces deployment and quality considerations.
  • No current first-party hosted path: The recorded end of 01.AI's hosted API service means users must manage deployment themselves or rely on another host.
  • Unverified knowledge cutoff: No first-party knowledge-cutoff date was confirmed for this exact checkpoint.

When to choose Yi-1.5-34B-Chat

Choose Yi-1.5-34B-Chat when you need an open-weight bilingual model and can operate the required infrastructure. It is especially appropriate for private assistants, English-Chinese support tools, coding applications, research projects, internal document workflows, and fine-tuning experiments where data-control requirements make a self-hosted model preferable.

It is also a reasonable choice when a 34B model's quality and reasoning capacity are more important than minimal hardware cost, and when a 4K context is sufficient. The Apache 2.0 license and broad compatibility with open-source serving tools make it easier to integrate into a controlled deployment than a hosted-only model.

A smaller local model may be more appropriate when response speed, low memory use, or inexpensive operation is the priority. The Yi-1.5 6B and 9B variants are relevant sibling options when the task does not justify the 34B model's infrastructure requirements. A longer-context model is a better fit for large documents or long-running conversations, while a current managed API model is more suitable when the project needs provider-operated scaling, current web grounding, native multimodal input, structured outputs, or a published token-pricing model.

Bottom line

Yi-1.5-34B-Chat is best understood as a capable, open-weight bilingual text model rather than a current hosted AI service. Its core value is the combination of a 34B parameter scale, Apache 2.0 licensing, coding and reasoning support, and the ability to run or fine-tune it with third-party infrastructure. Its main costs are operational: a modest 4K context, text-only behavior, significant hardware demands, and the loss of the original 01.AI hosted API route. For teams that value deployment control and can manage infrastructure, it remains a practical self-hosted option; for users seeking convenience, multimodality, web access, or predictable managed pricing, another model type will usually be a better fit.


Answers to Frequently Asked Questions

Is Yi-1.5-34B-Chat available through a current official API?
The supplied service-status research states that 01.AI's hosted model experience and API services ended on September 3, 2026. New users should therefore treat Yi-1.5-34B-Chat primarily as a self-hosted open-weight model or use a third-party hosting provider.
What is the context length of Yi-1.5-34B-Chat?
The standard Yi-1.5-34B-Chat checkpoint has a 4,096-token context window shared between the prompt and response. It is suitable for ordinary conversations, moderate code snippets, and shorter documents, but not for extensive long-context workloads.
What languages does Yi-1.5-34B-Chat support?
Yi-1.5-34B-Chat is designed primarily for English and Chinese use cases, including bilingual conversation and translation-adjacent workflows.
Can Yi-1.5-34B-Chat run locally?
Yes. Yi-1.5-34B-Chat is an open-weight model that can be self-hosted using tools such as Transformers, vLLM, SGLang, Ollama, or similar inference frameworks. The official BF16 checkpoint requires substantial GPU memory, while quantized versions may reduce hardware requirements.
What is Yi-1.5-34B-Chat?
Yi-1.5-34B-Chat is a 34-billion-parameter instruction-tuned causal language model from 01.AI. It is designed for text-based conversations, question answering, coding assistance, reasoning, mathematics, document processing, and other general-purpose language tasks.


Sources 4
Provider

About 01.AI