Yi

Yi-Lightning

by 01.AI · Proprietary hosted API model; current public availability and lifecycle status are not clearly documented

Yi-Lightning is 01.AI’s proprietary hosted mixture-of-experts language model for Chinese-English generation, reasoning, mathematics and coding. Released in October 2024, it offers an approximately 16K-token context and 4,096-token maximum output, with historical API pricing of about $0.14 per million input and output tokens. Current availability, pricing and advanced API features require verification.

Text Reasoning Coding
Yi-Lightning is a proprietary large language model from 01.AI designed to deliver capable text generation with relatively low serving cost and latency. Released in October 2024, it was positioned as a flagship API model for Chinese and English chat, reasoning, mathematics, coding and other general-purpose language tasks. 01.AI describes Yi-Lightning as a 100-billion-parameter mixture-of-experts model, although the company has not publicly established every architectural detail, including the number of active parameters used for each request. The model was historically available through an OpenAI-compatible developer API, with reported pricing of approximately $0.14 per million input tokens and $0.14 per million output tokens. Because current public model availability and pricing are not clearly documented, those figures should be treated as historical rather than guaranteed current rates.
Outputs

What Yi-Lightning can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Yi
Model type General Purpose
Context window 16K tokens
Maximum output 4K tokens
Release date October 2024
Status Proprietary hosted API model; current public availability and lifecycle status are not clearly documented
Knowledge cutoff notes

No authoritative provider-published knowledge-cutoff date was found for Yi-Lightning.

Model notes

01.AI describes Yi-Lightning as a 100-billion-parameter mixture-of-experts model. The technical report reports sixth place overall on Chatbot Arena at launch and strong results in Chinese, mathematics, coding and hard prompts. The model is proprietary and was offered through 01.AI’s developer API, not as a downloadable open-weight checkpoint. Historical third-party pricing references report approximately $0.14 per million input and output tokens and an approximately 16K context window. Current public model availability, pricing, structured-output support, tool calling, caching and batch processing are not sufficiently documented in current first-party material. Yi-Lightning should not be confused with Yi-Vision or the open-weight Yi model releases.

Cost

Model pricing

Input $0.14 per 1 million input tokens historically reported; verify current pricing with 01.AI
Output $0.14 per 1 million output tokens historically reported; verify current pricing with 01.AI
Model guide

Yi-Lightning: 01.AI’s Fast, Cost-Efficient Mixture-of-Experts Model

Yi-Lightning is 01.AI’s proprietary hosted language model, released in October 2024 as a fast and relatively inexpensive option for Chinese-English generation, reasoning, mathematics, coding and other text-based workloads. It uses a provider-described 100-billion-parameter mixture-of-experts architecture, supports an approximately 16K-token context window and was historically offered through an OpenAI-compatible API. Its main trade-offs are limited public documentation, uncertain current availability, a shorter context window than many newer models, no verified native multimodal output and no downloadable open-weight release.

What is Yi-Lightning?

Yi-Lightning is a proprietary language model developed by 01.AI. It generates text from text prompts and was built for general-purpose tasks such as conversation, reasoning, mathematics, coding, summarization, classification and multilingual text generation. The model was released in October 2024 and became available through 01.AI’s developer platform rather than as a downloadable model checkpoint.

01.AI describes Yi-Lightning as a 100-billion-parameter mixture-of-experts, or MoE, model. In a mixture-of-experts system, a request is routed to selected specialist components instead of activating the entire model for every token. This can provide a large overall model capacity without requiring the full parameter set to run on every request. In practical terms, the architecture was intended to improve the balance between response quality, speed and serving cost.

The 100-billion-parameter figure is a provider description of total model size. The supplied technical information does not establish the exact number of active parameters per request or provide a complete public architecture specification. Those details matter when comparing Yi-Lightning with other models, so the total parameter count should not be interpreted as a direct measure of its runtime cost or capability.

Where Yi-Lightning fits in 01.AI’s lineup

Yi-Lightning belongs to 01.AI’s Yi family of language models and was presented as the company’s flagship model for hosted API use at launch. It should not be confused with Yi-Vision, which is a separate model associated with visual understanding, or with 01.AI’s open-weight Yi releases. Yi-Lightning itself was offered as a proprietary hosted service rather than a model that developers could download and operate independently.

That distinction affects deployment decisions. A hosted API can avoid the hardware, model-serving and operational work required for self-hosting, but it also makes an application dependent on the provider’s endpoint, account requirements, pricing and model lifecycle. The current public information does not clearly confirm whether Yi-Lightning remains generally available, so teams evaluating it should verify access and supported model identifiers with 01.AI before building a production dependency.

Architecture and performance positioning

Yi-Lightning’s main technical distinction is its mixture-of-experts design. The model’s routing system selects experts for different parts of a request, while the provider’s technical report describes additional work on expert segmentation, routing and key-value caching. Key-value caching is a serving optimization that reuses information from earlier tokens during generation, helping reduce repeated computation in a conversation.

At launch, 01.AI reported strong results in Chinese, mathematics, coding and difficult prompts. The technical report also states that Yi-Lightning reached sixth place overall on Chatbot Arena at that time. These are historical launch results and provider-associated research claims, not a guarantee of current ranking or performance. Model rankings can change as evaluation sets, competing models and serving implementations change.

Editorially, Yi-Lightning is best understood as a cost- and speed-oriented general language model rather than a specialist model for very long documents or multimodal generation. Its historical positioning makes it most interesting when an application needs a hosted text model that can handle both Chinese and English while keeping token costs low.

Context window and maximum output

Available model references identify an approximately 16K-token context window, commonly represented as 16,384 tokens. The context window is the combined space available for the prompt, conversation history and generated response. In practical use, long instructions, large pasted documents and extensive chat history consume part of that space, leaving less room for the answer.

The supplied model data lists a maximum output of 4,096 tokens. This is a generation limit rather than a promise that every request will produce that many tokens. Applications should still handle shorter responses, truncation and provider-side request restrictions.

A 16K context can be sufficient for ordinary chat, extraction, coding assistance and moderate-sized documents. It is less suitable for workflows that routinely pass large books, lengthy repositories or extensive conversation histories in one request. Applications with those requirements may need chunking, retrieval, summarization or a different model with a longer documented context window.

Supported modalities and outputs

Yi-Lightning is documented as a text-generation model. Its supported input is text, and its direct output is text. There is no verified evidence in the supplied research that Yi-Lightning accepts image, audio or video input, and it is not documented as a model for generating images, audio or video.

This limitation is important because 01.AI’s broader platform includes visual-understanding capabilities through Yi-Vision. Those capabilities should not automatically be attributed to Yi-Lightning. If an application needs image interpretation or another non-text modality, it should use a model explicitly documented for that purpose rather than assuming that the Yi-Lightning endpoint supports it.

Reasoning, mathematics and coding

Yi-Lightning was designed for general reasoning and was specifically positioned around mathematics, coding and difficult prompts. It can therefore be considered for tasks such as explaining a calculation, transforming structured text, generating code, reviewing a short function, classifying incoming records or producing a first draft from instructions.

These capabilities should be interpreted as task positioning rather than a guarantee of correctness. The supplied research does not provide a current standardized benchmark table that would establish how Yi-Lightning compares with every newer reasoning or coding model. For high-stakes calculations, production code and business decisions, outputs still require validation.

Yi-Lightning’s coding role is strongest where the task fits within its context and does not require specialized development-environment access. It can generate or explain text-based code through the language API, but the supplied information does not verify built-in code execution, repository indexing or a provider-managed development environment.

API access, tools and structured output

Yi-Lightning was offered through an OpenAI-compatible chat-completions-style API. This compatibility can reduce migration work for applications already organized around chat messages, although compatibility with a common request format does not mean that every provider-specific feature behaves identically.

Streaming was documented as supported, allowing an application to receive generated text incrementally instead of waiting for the complete response. The supplied research does not verify current support for function calling, tool use, structured JSON output, response caching or batch processing. Those features should therefore be treated as unconfirmed rather than assumed from the API’s general compatibility.

For a basic integration, developers should confirm the current endpoint, model identifier, authentication process, token accounting, error behavior, rate limits and lifecycle status directly with 01.AI. Historical platform documentation may not reflect the provider’s current enterprise-focused product structure.

Historical pricing and cost trade-offs

Historical pricing references list approximately $0.14 per million input tokens and $0.14 per million output tokens. These figures are not presented as a currently guaranteed price: the supplied research specifically notes that current availability and pricing should be confirmed with 01.AI.

If still applicable, equal input and output rates would make Yi-Lightning relatively straightforward to budget for workloads with substantial generation, such as chat, extraction and high-volume text processing. Its mixture-of-experts architecture was also intended to support efficient serving. However, low historical token pricing should not be evaluated separately from availability, latency, rate limits and operational reliability.

The practical trade-off is therefore clear: Yi-Lightning may be attractive when a hosted model’s cost and response speed matter more than a very long context window, broad multimodal support or a large set of documented provider features. A more expensive or newer model may be preferable when the application needs stronger current guarantees, advanced tool integration or extensive long-context processing.

Main strengths and limitations

Strengths

  • Designed for efficient hosted inference through a mixture-of-experts architecture.
  • Strong historical positioning for Chinese-English use, mathematics, coding and general reasoning.
  • Historically low reported token pricing compared with many frontier API models.
  • OpenAI-compatible chat-completions-style access can simplify basic application integration.
  • Streaming support was documented for incremental text generation.
  • Suitable for general text workloads that do not require image, audio or video generation.

Limitations

  • Current availability, pricing and lifecycle status are not clearly documented in the supplied first-party material.
  • The approximately 16K-token context is limited for very long documents and large repository-level prompts.
  • The model is proprietary and hosted, with no verified downloadable open-weight checkpoint.
  • There is no verified native image, audio or video output.
  • Current support for tool calling, structured output, caching and batch processing is unconfirmed.
  • Historical launch rankings should not be treated as current performance guarantees.

When to choose Yi-Lightning

Yi-Lightning is a reasonable candidate when an application needs a relatively inexpensive hosted text model for Chinese-English conversation, summarization, classification, extraction, coding assistance, mathematical prompts or high-volume generation. It is particularly relevant for teams that prefer API access over operating model infrastructure and whose prompts fit comfortably within an approximately 16K-token context.

It may be less appropriate when the application requires a verified current service commitment, a long context window, downloadable weights, image understanding, native media generation or documented function-calling and structured-output features. In those cases, a different model category may be a better fit: a long-context model for large documents, a multimodal model for visual input, an open-weight model for private deployment or a provider with clearly documented tool and schema support for agent workflows.

Before selecting Yi-Lightning for production, verify that the model is still exposed through 01.AI’s current platform, confirm the live price and model name, test Chinese and English workloads representative of the application, and measure latency and output quality under realistic prompts. Its historical combination of speed, cost and general text capability is useful, but those benefits depend on current access and documentation.


Answers to Frequently Asked Questions

Who should choose Yi-Lightning?
Yi-Lightning may suit teams seeking a relatively inexpensive hosted text model for Chinese-English applications, summarization, classification, extraction, coding assistance, mathematics and high-volume generation. It is less suitable for workloads requiring very long context, downloadable weights, multimodal processing or verified support for tools and structured output.
What types of input and output does Yi-Lightning support?
Yi-Lightning is documented as a text-generation model that accepts text input and produces text output. There is no verified evidence that it supports image, audio or video input or generation; those capabilities should not be assumed from 01.AI’s separate Yi-Vision model.
What is Yi-Lightning’s context window and maximum output length?
Yi-Lightning is commonly documented with an approximately 16,384-token context window and a maximum output of 4,096 tokens. The context window includes the prompt, conversation history and generated response, so long documents and extensive chat histories may require chunking or summarization.
What is Yi-Lightning?
Yi-Lightning is a proprietary general-purpose language model developed by 01.AI and released in October 2024. It uses a 100-billion-parameter mixture-of-experts architecture for tasks such as conversation, reasoning, mathematics, coding, summarization, classification and multilingual text generation.
Is Yi-Lightning an open-source or downloadable model?
No. Yi-Lightning was offered as a proprietary hosted API service rather than a downloadable open-weight model. Developers should access it through 01.AI’s platform if the service and model identifier are currently available.


Sources 4
Provider

About 01.AI