What is Yi-Large?
Yi-Large is a proprietary large language model developed by 01.AI and released in May 2024. A large language model processes text prompts and generates text responses, making it suitable for conversational software, document analysis, drafting, summarization, classification, and other language-focused tasks.
01.AI describes Yi-Large as a 100-billion-parameter foundation model for chat, text generation, complex inference, prediction, and natural-language understanding. In practical terms, it is positioned as a high-end general-purpose model in the company’s API catalog rather than as a specialized image, audio, or video system.
The canonical API model identifier is yi-large. The model is available through 01.AI’s platform and can be accessed using an API interface compatible with OpenAI’s chat-completions style. This compatibility can reduce integration work for developers who already use that request pattern, although an application still needs to be configured with 01.AI’s endpoint and credentials.
Where Yi-Large fits in 01.AI’s lineup
Yi-Large belongs to 01.AI’s Yi family of foundation models. It is the standard text-generation model described in the supplied documentation, and it should not be treated as a bundle of every capability available elsewhere in the Yi catalog.
For example, 01.AI lists Yi-Large-FC separately for function or tool-use scenarios. Tool calling should therefore not be assumed for standard Yi-Large merely because a related model supports it. Similarly, image understanding is associated with the separate Yi-Vision model, not with the standard yi-large endpoint.
This distinction matters when selecting a model. Yi-Large is best understood as a general text model with a large context window. Applications that need verified built-in function calling or visual input should evaluate the specifically documented model for that requirement instead.
Verified specifications
| Specification | Yi-Large |
|---|---|
| Provider | 01.AI |
| Release | May 2024 |
| Parameters | 100 billion |
| Model identifier | yi-large |
| Context window | 32K tokens |
| Input | Text |
| Output | Text |
| API format | OpenAI-compatible chat-completions interface |
| Input price | $3 per 1 million tokens |
| Output price | $3 per 1 million tokens |
The 32K context window is the amount of text the model can consider within a request and its associated conversation context, subject to the provider’s implementation. It is useful for longer documents, extended instructions, and multi-turn conversations, but it does not guarantee that every long document will be handled equally well. Developers should still test how their application divides, retrieves, and presents information.
Capabilities and practical uses
01.AI positions Yi-Large for chat and text generation involving complex inference, prediction, and natural-language understanding. These provider-described capabilities support several practical use cases:
- Long-context conversation: building assistants that need to retain and process more text within a request than smaller-context models.
- Document and knowledge work: summarizing, comparing, extracting information from, and asking questions about lengthy text.
- Multilingual language processing: generating and analyzing text across languages, where the application’s language requirements match the model’s tested performance.
- Drafting and transformation: producing first drafts, rewriting material, classifying text, and turning unstructured language into more consistent formats.
- General reasoning tasks: handling analysis, prediction, and complex written instructions that do not require native visual input or a separately verified tool-calling model.
The model’s 100-billion-parameter scale and 32K context window suggest a design aimed at demanding language workloads rather than the lowest possible latency or cost. That is an editorial assessment of its positioning, not a provider-published benchmark result. Actual quality will depend on the prompt, language, subject matter, and evaluation criteria used by an application team.
Reasoning, coding, and tool support
Yi-Large is documented for complex inference and natural-language understanding, so it can be evaluated for reasoning-heavy text tasks such as comparing alternatives, following multi-step instructions, and organizing evidence. The supplied research does not identify a separate reasoning mode, dedicated reasoning-token budget, or verified reasoning benchmark for this model.
The model can be considered for code-related text generation because it is a general-purpose language model and 01.AI’s platform describes coding capabilities in its broader Yi documentation. However, the supplied model-specific material does not provide a coding benchmark or a guarantee of language-specific programming accuracy. Developers should test code generation, debugging, and explanation separately before relying on it in production.
Standard Yi-Large should not automatically be treated as a function-calling or agent model. The supplied documentation associates tool-use scenarios with the separately listed Yi-Large-FC model. If an application needs the model to emit structured tool calls that an external system will execute, the tool-specific model and its current documentation should be checked instead.
Input, output, and modality limits
The standard Yi-Large model is a text-in, text-out system. The supplied research verifies no native image, audio, or video output for this model. Image understanding is documented for Yi-Vision rather than Yi-Large, so visual prompts should not be assumed to work with the standard endpoint.
Streaming is supported through the API interface, which allows an application to receive generated text progressively instead of waiting for the complete response. The research does not directly verify a maximum output-token limit for Yi-Large. It also does not establish fine-tuning, prompt caching, batch processing, or a separate legacy JSON-mode capability for this exact model. These should be treated as unverified rather than assumed to be unavailable or supported.
Pricing and API access
01.AI’s API documentation lists Yi-Large at $3 per 1 million input tokens and $3 per 1 million output tokens. These are token-based API prices, not a monthly subscription price. Input tokens are the text sent to the model, while output tokens are the text it generates.
At the listed rates, a request using 100,000 input tokens would have an input charge of approximately $0.30, before any other provider-specific considerations. A response containing 10,000 output tokens would add approximately $0.03. Actual billing should be confirmed against the provider’s current pricing and tokenization rules.
The API uses the yi-large model identifier and an OpenAI-compatible chat-completions pattern. Compatibility can make migration from another API simpler, but it does not mean that every feature of another provider’s SDK or endpoint is available. Teams should verify authentication, current endpoint details, error handling, rate limits, and production availability in 01.AI’s documentation.
Strengths and trade-offs
Yi-Large’s clearest strengths are its large parameter scale, 32K context window, general-purpose language focus, and relatively straightforward API integration. It is a plausible choice when an application needs to process substantial text and values broad language capability over access to media-generation features.
Those strengths come with trade-offs. A 100-billion-parameter model may be a less economical choice than a smaller model for simple classification, short replies, or very high-volume workloads, although the supplied research does not provide a direct speed or benchmark comparison. The listed $3 per million input and output tokens also makes cost dependent on how much context an application sends and how long its responses are.
Yi-Large’s standard endpoint is not the appropriate default when the main requirement is image understanding, native media generation, verified function calling, integrated web search, or a documented maximum output limit. Related Yi models or another provider may be more suitable for those needs, but the exact alternative should be selected from current capability documentation rather than inferred from the shared model family.
When to choose Yi-Large
Choose Yi-Large when the application needs a general-purpose text model with a documented 32K context window, API-based access, and support for substantial language-processing tasks. It is especially relevant for:
- long-form document analysis and summarization;
- enterprise assistants centered on text and knowledge work;
- multilingual generation and transformation;
- complex written instructions and analysis;
- applications that can use an OpenAI-compatible chat-completions interface; and
- workloads where a larger general model is preferable to a minimal, low-cost text model.
Consider another option when low latency and minimum cost are more important than model scale, when the application requires image or video input, or when built-in tool calling is central to the workflow. For those cases, Yi-Large-FC, Yi-Vision, a smaller language model, or a specialized multimodal system may be more appropriate depending on the verified requirement.
Known limitations and unverified details
The supplied first-party material does not directly establish Yi-Large’s knowledge-cutoff date, maximum output-token limit, fine-tuning availability, prompt-caching support, batch API support, or a distinct JSON-mode specification. It also does not provide model-specific benchmark results for reasoning, coding, speed, or factual reliability.
These gaps do not invalidate the documented specifications, but they are important for production planning. A team evaluating Yi-Large should test representative prompts, measure latency and token usage, validate generated code or factual claims, and confirm the current API contract with 01.AI before making the model a production dependency.

