Yi Large

Yi Large Turbo

by 01.AI · Documented in the 01.AI API catalog; current operational availability should be verified through the provider

Yi Large Turbo is 01.AI’s text-only model for fast, cost-sensitive chat, summarization, classification, coding assistance, and general text generation. It has a documented 4,096-token context window, streaming, and OpenAI-compatible API access priced at $0.19 per million input and output tokens.

Text Reasoning Coding
Yi Large Turbo is 01.AI’s speed- and cost-oriented model for applications that need dependable general-purpose text generation without paying for a larger or more specialized model. It accepts text and returns text through an OpenAI-compatible chat-completions API. The provider documents a 4,096-token context window and pricing of $0.19 per million input tokens plus $0.19 per million output tokens. Its main trade-off is scope: it is best suited to short- and medium-length language tasks, not long-context analysis, native multimodal work, or advanced function-calling agents.
Outputs

What Yi Large Turbo can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

6/10 Reasoning
5/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Yi Large
Model type General Purpose
Context window 4K tokens
Knowledge cutoff 2024-03-31
Release date 2024-08-02
Status Documented in the 01.AI API catalog; current operational availability should be verified through the provider
Knowledge cutoff notes

The March 31, 2024 cutoff is reported in OpenRouter's model metadata. A directly published 01.AI knowledge-cutoff statement for Yi Large Turbo was not located.

Model notes

01.AI's API documentation describes Yi Large Turbo as a high-performance, cost-effective model for complex inference and high-quality text generation. The same documentation lists a 4K context window and pricing of $0.19 per million input and output tokens. OpenRouter reports an August 2, 2024 release date and a March 31, 2024 knowledge cutoff. 01.AI documents function calling specifically for Yi Large FC rather than Yi Large Turbo. Some third-party catalogs report different context lengths or prices, so the provider's current API documentation should be treated as the pricing authority.

Cost

Model pricing

Input $0.19 per 1 million tokens
Output $0.19 per 1 million tokens
Model guide

Yi Large Turbo: A Low-Cost, Fast Model for Short-Context Text Generation

Yi Large Turbo is a text-only language model from 01.AI aimed at fast, cost-sensitive chat, summarization, classification, and general text generation. It offers a documented 4,096-token context window, streaming, and OpenAI-compatible chat-completions access at $0.19 per million input tokens and $0.19 per million output tokens according to 01.AI documentation.

What is Yi Large Turbo?

Yi Large Turbo is a proprietary text-generation model developed by 01.AI. It belongs to the Yi model family and is positioned for fast, cost-efficient inference rather than for the widest possible context window or the most specialized agent features.

In practical terms, the model takes text messages as input and produces generated text. Suitable workloads include chat completion, drafting, rewriting, summarization, classification, question answering, and other general natural-language transformations. The model can also be used behind a retrieval system, where an external application supplies relevant documents before asking it to answer. That arrangement should not be confused with built-in web search: the available research does not document native browsing or web grounding for Yi Large Turbo.

01.AI describes the model as suitable for complex inference and high-quality text generation. Those are provider-positioning claims rather than independent benchmark results, so performance should be evaluated against the particular prompts, languages, and response-quality requirements of a planned application.

Where Yi Large Turbo fits in the Yi lineup

Yi Large Turbo is one of the general-purpose text models in 01.AI’s catalog. Its role is different from models that add specialized capabilities. For example, the supplied 01.AI documentation identifies Yi Vision for visual understanding and Yi Large FC for function calling. Those comparisons help define Yi Large Turbo’s boundaries: it is intended primarily for text in and text out, rather than image understanding or native tool orchestration.

This makes Yi Large Turbo a relatively straightforward choice when an application needs a language model for ordinary text processing and does not require a model-specific vision or function-calling interface. Developers should not assume that capabilities documented for another Yi model automatically apply to Yi Large Turbo.

Capabilities and supported modalities

The documented input and output format is text. Yi Large Turbo does not have documented image, audio, or video input, and it does not produce images, audio, or video. It is therefore not an appropriate primary model for visual question answering, image generation, speech applications, or video analysis.

Within text workflows, it can support common language operations such as:

  • Conversational assistants and customer-service chat
  • Drafting, rewriting, and tone transformation
  • Summarization and information extraction
  • Question answering over text supplied by an application
  • Classification and routing
  • Short-form coding assistance and code-related text generation
  • Knowledge-search interfaces paired with an external retrieval system

The provider’s wider platform materials describe coding and logical reasoning among Yi-series use cases. For this exact model, the supplied catalog characterizes it as a general-purpose language model rather than documenting a separate reasoning mode or a coding-specialist configuration. Coding and reasoning should therefore be treated as ordinary text-generation workloads, with quality validated against the application’s own examples.

Context window and technical profile

01.AI’s API documentation lists a 4,096-token context window for Yi Large Turbo. A token is a unit of text used by the model; the context window covers the material the model can consider in a request, including the conversation history and the requested response as applicable. A 4K limit is adequate for short conversations, compact instructions, individual documents, and focused transformations, but it can become restrictive when an application sends long chat histories, large source files, or multiple retrieved documents at once.

The available research does not specify a maximum output-token limit for this model. Applications should not invent one in their own capacity planning. Instead, developers should consult the current 01.AI API documentation and design requests so that the combined prompt and response remain within the provider’s current limits.

The API supports streaming responses. Streaming sends generated text incrementally rather than waiting for the complete answer, which can make an interactive interface feel faster even though it does not necessarily reduce total generation time or token cost.

OpenRouter lists an August 2, 2024 release date and a March 31, 2024 knowledge cutoff. These are hosting-platform metadata, not a directly published 01.AI knowledge-cutoff statement, so they should be treated as reference information rather than a guarantee about every deployment.

Pricing and API access

01.AI’s API documentation lists Yi Large Turbo at $0.19 per 1 million input tokens and $0.19 per 1 million output tokens. Input tokens are the text sent to the model, while output tokens are the text it generates. Because the two rates are the same in the cited documentation, an application’s cost depends on both prompt size and response length.

At that rate, Yi Large Turbo is designed for workloads where a low per-token price matters: high-volume classification, short summaries, routine drafting, lightweight chat, and other requests that do not require a large context window. Actual availability, billing, and current pricing can change, so the provider’s platform should be checked before deployment.

The model identifier documented for API use is yi-large-turbo. 01.AI provides an OpenAI-compatible chat-completions interface, which can reduce integration work for applications already structured around that request format. Compatibility with a familiar API shape does not mean that every OpenAI feature is supported. In particular, the supplied research does not clearly document structured outputs, caching, batch processing, or fine-tuning for this exact model.

Reasoning, coding, and tool use

Yi Large Turbo is suitable for general reasoning expressed through text, including classification decisions, comparisons, extraction tasks, and multi-step instructions of modest length. The model is also a plausible option for basic coding assistance, such as generating snippets, explaining code, rewriting text, or producing structured descriptions of programming tasks. These are practical use cases, not evidence of a separately trained reasoning or coding mode.

Native function calling or tool use is not clearly documented for Yi Large Turbo. 01.AI specifically documents function calling for Yi Large FC, so developers building an agent that must call tools, APIs, or business systems should not assume that Yi Large Turbo provides the same interface. A surrounding application could still parse model text and implement its own workflow, but that is different from provider-supported native function calling and may require additional validation and error handling.

Strengths and limitations

Strengths

  • Low documented price: $0.19 per million input tokens and $0.19 per million output tokens makes it suitable for cost-sensitive workloads.
  • Speed-oriented positioning: 01.AI presents Turbo as a faster alternative within the Yi family, making it a candidate for interactive applications and high-volume generation.
  • Broad text-task coverage: It can handle ordinary chat, rewriting, summarization, classification, extraction, and general drafting without requiring a specialized interface.
  • Simple integration path: The OpenAI-compatible chat-completions API and streaming support are useful for applications using common chat-generation patterns.
  • Multilingual potential: The Yi family is associated with bilingual and multilingual language workflows, although application-specific language quality should be tested rather than assumed.

Limitations

  • Small context window by current standards: The documented 4,096-token limit is not well suited to very long documents, extensive conversation memory, or large retrieval batches.
  • Text-only operation: It does not provide documented image, audio, or video input or output.
  • No clearly documented native tool calling: Applications requiring reliable structured tool invocation may need another model, such as the Yi Large FC model referenced in 01.AI’s documentation.
  • Unspecified output ceiling: The supplied research does not establish a maximum output-token limit.
  • Unclear advanced API features: Structured outputs, fine-tuning, caching, and batch support are not clearly documented for this exact model.
  • Limited grounding: It has no documented built-in web search, so current-information answers require an external retrieval or search layer.

When to choose Yi Large Turbo

Choose Yi Large Turbo when the main requirement is economical, relatively fast text generation and the requests can fit within a 4,096-token context. It is a sensible candidate for customer-support chat, short knowledge-base answers, document summaries, email drafting, text classification, content transformation, and other high-volume workflows where every request is relatively focused.

It is especially attractive when an existing application benefits from an OpenAI-compatible chat-completions format and streaming responses, but does not need advanced agent features. Its low documented token price can also make it useful for first-pass processing, such as classifying incoming text before sending only selected cases to a more capable or specialized model.

Consider another option when the application must process long documents or lengthy conversations, understand images, generate media, browse the web, or call external tools through a provider-supported function-calling mechanism. Yi Vision is the more relevant named sibling for visual understanding, while Yi Large FC is the model specifically associated with function calling in the supplied 01.AI documentation. A larger-context model may be more appropriate for long documents, and a model with native search or retrieval integration may be preferable for current factual information.

Bottom line

Yi Large Turbo is a focused general-purpose text model rather than an all-in-one multimodal or agent platform. Its strongest practical case is short-context text generation where low cost, straightforward API integration, and responsive output matter more than the largest context window or the broadest set of built-in capabilities. The key decision is whether the application can stay within 4,096 tokens and handle retrieval, tools, and other extensions outside the model itself. If so, Yi Large Turbo offers a clear cost-and-speed trade-off; if not, a longer-context, multimodal, or function-calling model will likely be a better fit.


Answers to Frequently Asked Questions

How can developers access Yi Large Turbo?
Developers can access the model through 01.AI’s OpenAI-compatible chat-completions API using the model identifier "yi-large-turbo". The API also supports streaming responses, although supported features and current availability should be confirmed in the latest provider documentation.
Does Yi Large Turbo support function calling, web search, or multimodal input?
Native function calling and built-in web search are not clearly documented for Yi Large Turbo. It is a text-only model with no documented image, audio, or video input or output. Applications can add retrieval or custom tool workflows externally, but those are not the same as provider-supported native capabilities.
How much does Yi Large Turbo cost?
01.AI documents Yi Large Turbo at $0.19 per 1 million input tokens and $0.19 per 1 million output tokens. Pricing and availability may change, so developers should verify the current rates before deployment.
What is Yi Large Turbo best used for?
Yi Large Turbo is best suited to low-cost, fast text-generation tasks such as chat, summarization, rewriting, classification, information extraction, question answering over supplied text, drafting, and short-form coding assistance.
What is the context window of Yi Large Turbo?
Yi Large Turbo has a documented 4,096-token context window. This is suitable for short conversations, compact instructions, individual documents, and focused transformations, but it may be restrictive for long documents, extensive chat histories, or large retrieval results.


Sources 4
Provider

About 01.AI