What is K-EXAONE-236B-A23B?
K-EXAONE-236B-A23B is an open-weight multilingual language model developed by LG AI Research. The model is designed to generate and analyze text, solve multi-step problems, write code, process long documents, and participate in tool-enabled workflows. Its canonical public identifier is LGAI-EXAONE/K-EXAONE-236B-A23B.
The model was released through LG AI Research’s official K-EXAONE repository and Hugging Face model page on December 31, 2025, according to the supplied official research. Within LG AI Research’s current model work, it represents a large-scale open-weight option for research, enterprise experimentation, and developers who need to run or control inference themselves. It is not presented as a conventional subscription chatbot with a standard consumer plan.
The name describes the model’s scale and sparse architecture: it contains 236 billion total parameters, while approximately 23 billion are active during each inference step. Parameters are the learned values used by a neural network to process input and produce output. Because a Mixture-of-Experts model routes each token through selected parts of the network, total model capacity and per-token computation are not the same thing.
Sparse architecture and 256K context
K-EXAONE-236B-A23B uses a fine-grained Mixture-of-Experts architecture with a hybrid attention design. In practical terms, the model has a very large overall parameter capacity but does not evaluate every parameter for every token. This can reduce per-token computation compared with a dense model containing the same total number of parameters, although the complete model still has substantial memory and infrastructure requirements.
Its documented native context window is up to 256,000 tokens. The context window is the amount of input and generated conversation or document content that the model can consider within one request. A 256K-token window is useful for long reports, large codebases, extended transcripts, and multi-document analysis, but it does not make the model a reliable database. Long inputs can still contain irrelevant or ambiguous information, and the model’s answers should be checked against source material.
LG AI Research’s technical materials describe a staged extension from 8K to 32K and then to 256K tokens. The model’s hybrid attention combines global and local attention, including a sliding-window component, to help manage the memory and computation involved in long-context processing.
Reasoning, languages, and coding
Reasoning and non-reasoning modes
K-EXAONE-236B-A23B supports both reasoning and non-reasoning modes. Reasoning mode is enabled by default in the recommended chat-template configuration. It is intended for tasks that benefit from more deliberate, multi-step processing, such as mathematics, science, coding, complex analysis, and structured problem solving.
Developers can disable extended reasoning with enable_thinking=False. This provides a practical latency trade-off: non-reasoning mode may be more appropriate for straightforward classification, extraction, rewriting, or routine responses where extended deliberation is unnecessary. The supplied research does not establish a single guaranteed speed improvement or a fixed reasoning-token budget for every deployment.
Six supported languages
The model supports Korean, English, Spanish, German, Japanese, and Vietnamese. LG AI Research describes expanded multilingual training and cross-lingual knowledge-transfer methods, with particular emphasis on Korean and broader multilingual capability. This makes the model relevant to applications that need to work across these languages rather than only translating between them.
Language support should not be interpreted as identical performance in every language or task. The supplied research identifies the supported languages but does not provide a universal quality guarantee for each one. Teams should evaluate the model with their own terminology, documents, and safety requirements.
Coding and document analysis
Coding is one of the model’s intended uses. It can assist with code generation, explanation, transformation, debugging, and software-related reasoning. Its long context can also help when a task requires examining several files or a lengthy technical specification, provided the deployment can accommodate the requested context length.
The model is also suited to long-document question answering and analysis. Example applications include extracting information from large reports, comparing sections across documents, summarizing technical material, and asking follow-up questions over an extended source. These uses still require verification because the model can produce inaccurate or unsupported statements.
Tool calling and deployment options
K-EXAONE-236B-A23B supports tool calling through OpenAI-compatible and Hugging Face tool specifications. A developer can provide a tool schema describing available functions, their arguments, and expected structure, allowing the model to request actions such as retrieval, calculations, or calls to an external business system. The model produces the tool request; the surrounding application remains responsible for executing the function, validating arguments, handling errors, and returning the result.
The official materials demonstrate agentic workflows and document deployment with Transformers, vLLM, SGLang, TensorRT-LLM, and llama.cpp-compatible GGUF files. This breadth gives infrastructure teams several runtime choices, but it does not mean every environment offers identical performance or feature support. Runtime behavior depends on the serving engine, quantization format, batch size, context length, and hardware.
LG AI Research states that vLLM and SGLang deployments can provide the full 256K context length using tensor parallelism across four H200 GPUs, subject to the relevant runtime and implementation requirements. This is deployment guidance, not a minimum hardware guarantee for every configuration. The model’s size makes it unsuitable for most ordinary laptops, low-memory servers, and low-resource single-GPU setups without significant quantization and additional compromises.
Performance, speed, and cost trade-offs
The model’s main practical trade-off is between capability and infrastructure complexity. Its 236B total parameters, 23B active parameters, long context, reasoning modes, and tool support target demanding workloads. However, even with sparse activation, hosting the full model requires substantial memory, multi-GPU infrastructure, and operational expertise.
Quantized and GGUF variants can reduce memory requirements and make lower-resource deployment more feasible. Quantization can also affect output quality, throughput, supported context length, or runtime compatibility, so the most suitable format depends on the application. A team seeking low latency or inexpensive high-volume inference may prefer a smaller model, a non-reasoning configuration, shorter contexts, or a managed service operated by a third party.
No official first-party hosted token pricing was identified in the supplied research. There is therefore no verified input-price or output-price schedule to report for K-EXAONE-236B-A23B. The financial cost of using it depends on hardware, cloud rental, electricity, storage, engineering time, runtime configuration, and any separate commercial service or licensing arrangement. Open weights do not mean that production deployment is cost-free.
Modalities and output limits
The documented public model is text-only. It accepts text input and produces text output; it is not natively documented as an image, audio, or video input or output model. It should not be selected when the primary requirement is image understanding, image generation, speech recognition, audio generation, or video processing.
The supplied official examples use max_new_tokens values of 16,384 for reasoning and tool-use examples and 1,024 for non-reasoning examples. These are usage examples rather than a clearly documented hard maximum output limit. The maximum usable output can also be constrained by the serving engine, available memory, context window, and deployment configuration, so no verified universal maximum-output figure is available.
Main strengths and limitations
Where the model is strong
- Large open-weight model with 236B total and approximately 23B active parameters.
- Very long 256K-token context for document-heavy and code-heavy workloads.
- Reasoning and non-reasoning modes for balancing deliberation and responsiveness.
- Support for Korean, English, Spanish, German, Japanese, and Vietnamese.
- Coding, mathematical, scientific, and complex-analysis use cases.
- Tool calling and agent workflows using OpenAI-compatible and Hugging Face specifications.
- Multiple deployment paths, including vLLM, SGLang, TensorRT-LLM, Transformers, and llama.cpp/GGUF.
Important limitations
- Large hardware requirements make full deployment impractical for many individuals and small teams.
- There is no verified official first-party hosted API price for this model.
- API access, where available through another operator, may have separate terms, pricing, availability, and performance.
- The model is text-only and does not natively provide image, audio, or video capabilities.
- It can generate inaccurate, biased, harmful, or syntactically incorrect content.
- LG AI Research warns that the model does not reflect the latest information, so it is not a real-time knowledge source without retrieval and verification.
- The supplied materials do not state an exact knowledge-cutoff date or a universal hard maximum for generated output.
When to choose K-EXAONE-236B-A23B
Choose K-EXAONE-236B-A23B when you need an open-weight model for self-hosted or controlled infrastructure and your workload benefits from long context, multilingual capability, reasoning, coding, or tool-enabled agents. It is particularly relevant for Korean-language applications, research environments, enterprise experimentation, long-document analysis, and teams that need to inspect or adapt the deployment rather than rely entirely on a consumer application.
It may also be a good fit when the model’s sparse architecture offers a useful capacity-versus-computation compromise and the organization already operates suitable multi-GPU infrastructure. The ability to switch between reasoning and non-reasoning modes can help separate complex tasks from high-throughput routine processing.
Another option may be more appropriate when low latency, low infrastructure cost, simple installation, or predictable hosted pricing is the priority. A smaller language model is likely to be easier to run for lightweight applications. A managed API may be preferable when a team does not want to operate multi-GPU serving infrastructure. A multimodal model is a better choice when image, audio, or video input is central to the workflow. For current facts, the model should be paired with an external retrieval or search system rather than used alone.
Bottom line
K-EXAONE-236B-A23B is a specialized large open-weight model rather than a ready-made consumer chatbot. Its defining combination is a 236B-parameter sparse MoE architecture, approximately 23B active parameters, a 256K context window, six-language support, reasoning controls, coding capability, and tool calling. Those features make it attractive for capable self-hosted multilingual and long-context systems, but its hardware demands, uncertain hosted costs, text-only design, and lack of real-time knowledge make careful deployment planning essential.

