EXAONE 4.0

EXAONE-4.0-32B

by LG AI Research · Available open-weight model; no official deprecation or shutdown date found

EXAONE-4.0-32B is LG AI Research’s 32B open-weight text model for Korean, English, and Spanish. It supports reasoning and non-reasoning modes, a 131,072-token context window, coding, and application-controlled tool calling. The model is intended mainly for self-hosted deployment through Transformers, llama.cpp, TensorRT-LLM, or vLLM. No official hosted API pricing was verified, and its non-commercial license, hardware requirements, text-only modality, and unknown maximum output limit are important constraints.

Text Reasoning Coding
EXAONE-4.0-32B is a 32B-class open-weight causal language model from LG AI Research. It is designed for users who want to run a capable multilingual model themselves rather than depend on a conventional consumer chatbot or a broadly documented first-party API. The model supports Korean, English, and Spanish, combines reasoning and non-reasoning modes, accepts long text contexts up to 131,072 tokens, and provides documented function and tool-calling support. It is a strong candidate for research, Korean-language applications, coding, document processing, and controlled deployments, but its hardware demands, licensing conditions, and lack of official hosted token pricing make it less convenient than a managed API model.
Outputs

What EXAONE-4.0-32B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family EXAONE 4.0
Model type Reasoning
Context window 131K tokens
Release date 2025-07-15
Status Available open-weight model; no official deprecation or shutdown date found
Knowledge cutoff notes

No authoritative exact knowledge-cutoff date was identified in the official model documentation or model card.

Model notes

EXAONE-4.0-32B is a 32B-class dense causal language model in the EXAONE 4.0 family. LG AI Research describes it as a hybrid model with selectable non-reasoning and reasoning modes. The official model configuration lists 30.95B parameters excluding embeddings, 64 layers, grouped-query attention with 40 attention heads and 8 key-value heads, and a 131,072-token context length. It supports English, Korean, and Spanish. Official documentation describes agentic function/tool calling through supplied tool schemas. The model is distributed under the EXAONE AI Model License Agreement 1.2 - NC, which includes non-commercial restrictions and restrictions on developing competing models. Official deployment guidance covers Transformers, llama.cpp, TensorRT-LLM, and vLLM. Editorial scores are comparative estimates rather than vendor-provided ratings. No exact knowledge cutoff, maximum output-token limit, official hosted API pricing, fine-tuning support, prompt-caching support, batch API, legacy JSON mode, or first-party web-search capability was verified.

Cost

Model pricing

Input No official hosted API pricing found; intended primarily for self-hosted deployment
Output No official hosted API pricing found; intended primarily for self-hosted deployment
Model guide

EXAONE-4.0-32B: Open-Weight Korean Reasoning Model for Local Deployment

EXAONE-4.0-32B is LG AI Research’s 32-billion-parameter open-weight language model for Korean, English, and Spanish text generation, reasoning, coding, long-context processing, and tool use. It offers selectable reasoning and non-reasoning modes, a 131,072-token context window, and deployment support through Transformers, llama.cpp, TensorRT-LLM, and vLLM. Its main trade-offs are the lack of official hosted pricing, substantial hardware requirements, non-commercial licensing restrictions, and no native image, audio, or video input.

What is EXAONE-4.0-32B?

EXAONE-4.0-32B is an open-weight language model provided by LG AI Research as part of the EXAONE 4.0 family. It is a dense causal language model: given a sequence of text, it predicts and generates the next tokens. In practical terms, it can answer questions, write and transform text, reason through problems, generate code, process long documents, and call external tools when an application supplies the required tool definitions.

The model is aimed primarily at self-hosted and research-oriented use. Its weights and deployment instructions are publicly available through LG AI Research’s official GitHub and Hugging Face releases. This makes EXAONE-4.0-32B different from a typical consumer chatbot, where the provider hosts the model and exposes capabilities through a web application or subscription. With EXAONE-4.0-32B, the operator is generally responsible for the hardware, software stack, security, monitoring, and licensing compliance.

The model was released on July 15, 2025, according to the supplied release information. LG AI Research describes the EXAONE 4.0 generation as supporting both reasoning and non-reasoning operation. The 32B model is the larger professional-scale variant in the documented EXAONE 4.0 release, alongside a separate 1.2B on-device model that should not be confused with this item.

Position in the EXAONE family

EXAONE-4.0-32B sits at the intersection of LG AI Research’s general language-model and enterprise AI work. It is not the same product as the EXAONE Showroom, a consumer-facing demonstration environment, and it is not a hosted subscription tier. It is the model artifact intended for developers, researchers, and organizations that need control over deployment.

Its most important positioning distinction is the combination of a relatively large 32B parameter scale, multilingual support focused on Korean, English, and Spanish, selectable reasoning behavior, and open-weight availability. LG AI Research also publishes other EXAONE-related systems, including vision-language and specialized industrial solutions, but those are separate models or products. EXAONE-4.0-32B itself is text-only for both input and output.

Verified specifications

SpecificationEXAONE-4.0-32B
ProviderLG AI Research
Model familyEXAONE 4.0
Model typeDense causal language model with reasoning and non-reasoning modes
Parameter scale32B class; official configuration lists 30.95B parameters excluding embeddings
Context length131,072 tokens
LanguagesEnglish, Korean, and Spanish
InputText
OutputText
Tool useFunction and tool calling through supplied tool schemas
Official hosted priceNot found; intended primarily for self-hosted deployment
Maximum output tokensNot verified

The 131,072-token context window is one of the model’s most useful technical characteristics. A context window is the amount of text and other conversational information the model can consider in one request, including the prompt and the generated response. A long window can help with large documents, extended coding sessions, and multi-step tasks, although real-world capacity depends on the serving implementation, available memory, prompt structure, and generation settings.

No authoritative maximum output-token limit was verified in the supplied research. The context window should therefore not be interpreted as a guaranteed amount of generated text. Operators should check the selected inference framework and configuration for practical generation limits.

Reasoning, coding, and tool use

EXAONE-4.0-32B offers selectable reasoning and non-reasoning modes. Reasoning mode is intended for tasks where the model benefits from spending additional computation on intermediate problem solving, such as multi-step analysis, difficult instructions, or structured decision support. Non-reasoning mode is better suited to straightforward generation and tasks where lower latency is more important than extended deliberation. The supplied research verifies the availability of these modes but does not provide a single universal quality level for every task.

The model is also documented for agentic function and tool calling. In this setup, an application supplies tool schemas that describe available functions, their names, and their expected arguments. The model can then produce a structured request for the application to execute. The external function, such as a database lookup or business-system action, is not executed by the model itself. Developers must implement execution, validation, permissions, error handling, and any confirmation step.

Coding is a supported use case rather than a separate programming-only model mode. EXAONE-4.0-32B can be considered for code generation, explanation, transformation, and debugging workflows, particularly when the surrounding system can provide a long repository or specification context. The research does not establish a universal benchmark ranking, so coding quality should be evaluated against the languages, frameworks, and repository patterns relevant to a particular project.

Modalities and deployment options

This model accepts text and produces text. It does not natively accept images, audio, or video, and it does not directly output those media types. That limitation matters when comparing it with a vision-language model or a multimodal assistant. An image or document-understanding workflow would need a separate preprocessing component, or a different EXAONE model specifically designed for vision-language input.

Official deployment guidance covers Transformers, llama.cpp, TensorRT-LLM, and vLLM. These options span common research and production-oriented inference environments. Transformers is widely used for experimentation and model integration; llama.cpp can support optimized local inference workflows; TensorRT-LLM is relevant to NVIDIA-oriented optimized serving; and vLLM is commonly used for high-throughput model serving. The precise hardware requirements, quantization options, throughput, and latency depend on the chosen framework and configuration.

A 32B-class model generally requires substantially more memory and operational planning than a small on-device model. Quantization can reduce memory requirements, but the supplied research does not establish a single minimum hardware specification or guarantee a particular speed after quantization. Users with limited hardware should test the intended model format and workload rather than assume that a local installation will be practical.

Pricing and licensing

No official hosted API input or output price was found for EXAONE-4.0-32B. The model is primarily presented for self-hosted deployment, so the financial calculation is different from a per-token API comparison. Costs may include GPU or cloud-instance time, storage, electricity, engineering, monitoring, and maintenance. A self-hosted deployment can be attractive when usage is predictable or data-control requirements are important, but it is not automatically cheaper for occasional or low-volume workloads.

The model is distributed under the EXAONE AI Model License Agreement 1.2 - NC. The supplied research identifies non-commercial restrictions and restrictions related to developing competing models. These conditions are important: public availability of model weights does not mean unrestricted commercial use. Before deploying the model in a business, customer-facing service, or revenue-generating workflow, the operator should read the current license and obtain any required permission or commercial arrangement.

Because there is no verified maximum output-token limit or official hosted price, a precise cost-per-request comparison cannot be calculated from the available information. Any cost or performance score should be treated as an editorial estimate, not a provider-published specification.

Main strengths and limitations

Strengths

  • Strong Korean-language positioning: Korean is a central supported language, alongside English and Spanish, making the model relevant to multilingual applications that are not designed only around English.
  • Open-weight access: Organizations can study and deploy the model through supported local or self-managed infrastructure rather than relying exclusively on a consumer application.
  • Reasoning flexibility: Selectable reasoning and non-reasoning modes allow applications to trade response depth against latency for different tasks.
  • Long context: The verified 131,072-token context length is suitable for large prompts, extended coding contexts, and long-document workflows, subject to serving limits.
  • Agent integration: Documented function and tool calling can connect model responses to application-defined operations.
  • Deployment choice: Support for Transformers, llama.cpp, TensorRT-LLM, and vLLM gives technical teams several implementation paths.

Limitations

  • Text-only operation: Native image, audio, and video input are not supported by this model.
  • Hardware and operations: A 32B model is more demanding to serve than a small local model, especially without quantization or optimized inference.
  • No verified hosted pricing: Users seeking a simple pay-as-you-go API cannot rely on an official LG-hosted token price from the supplied information.
  • License restrictions: The NC license includes non-commercial and competing-model restrictions that may rule out some commercial deployments.
  • Incomplete service-style limits: The exact knowledge cutoff, maximum output tokens, fine-tuning support, caching, batch API, and legacy JSON-mode support were not verified.
  • No first-party web search: The supplied model data does not identify built-in web-search capability. Current information retrieval would require an external system if permitted by the deployment design.

When to choose this model

Choose EXAONE-4.0-32B when you need a self-managed multilingual text model and can support the operational requirements of a 32B deployment. It is especially suitable for Korean-language research, internal document and knowledge workflows, coding assistance, long-context analysis, and agent experiments where the application needs to control tool definitions and execution. It may also fit organizations that prefer keeping inference on their own infrastructure, subject to the license and applicable security requirements.

The model is a reasonable choice when the ability to inspect or locally operate open weights matters more than the convenience of a managed API. Its reasoning switch can also be useful in systems that need a faster ordinary response for simple prompts and a more deliberate mode for complex tasks.

Another option may be more appropriate when the workload requires native image understanding, audio or video processing, guaranteed hosted API availability, transparent token pricing, or a mature consumer application. A smaller model may be preferable when low latency, low memory use, or edge deployment is the priority. A managed commercial model may be easier for production teams that need service-level commitments, standardized billing, built-in web access, or a clearly documented JSON and batch-processing interface. A vision-language member of the broader EXAONE ecosystem would be more suitable for image and document inputs than EXAONE-4.0-32B itself.

Overall assessment

EXAONE-4.0-32B is best understood as a capable open-weight research and deployment model rather than a ready-made chatbot subscription. Its distinguishing combination is a 131,072-token context window, Korean-English-Spanish text support, selectable reasoning behavior, coding capability, and application-controlled tool use. Those features make it relevant to organizations building their own systems, particularly where Korean-language performance and deployment control are important.

Its trade-offs are equally practical. The model requires meaningful infrastructure, has no verified official hosted price, does not provide native multimodal input or output, and carries license conditions that must be reviewed before commercial use. Editorial assessments in the supplied data rate its reasoning and coding relatively highly, with moderate speed and strong estimated cost efficiency for an open-weight model, but these are comparative editorial judgments rather than LG AI Research specifications. Prospective users should validate performance, memory use, latency, and license fit on their own workloads before selecting it for production.


Answers to Frequently Asked Questions

What license does EXAONE-4.0-32B use, and can it be used commercially?
EXAONE-4.0-32B is distributed under the EXAONE AI Model License Agreement 1.2 - NC. The license includes non-commercial and competing-model restrictions, so users must review the current terms and obtain any required permission before commercial or revenue-generating deployment.
Can EXAONE-4.0-32B be deployed locally?
Yes. EXAONE-4.0-32B is intended primarily for self-hosted and research-oriented deployment. Official guidance covers Transformers, llama.cpp, TensorRT-LLM, and vLLM, although a 32B-class model requires substantial memory and operational planning.
What is the context length of EXAONE-4.0-32B?
EXAONE-4.0-32B has a verified context length of 131,072 tokens. Its practical capacity depends on the inference framework, available memory, prompt structure, and generation settings.
What is EXAONE-4.0-32B?
EXAONE-4.0-32B is a 32B-class open-weight dense causal language model from LG AI Research. It supports text generation, reasoning, coding, long-context analysis, and application-controlled tool calling, with a focus on Korean, English, and Spanish.
What languages and input modalities does EXAONE-4.0-32B support?
EXAONE-4.0-32B supports Korean, English, and Spanish. It is text-only: it accepts text input and produces text output, without native image, audio, or video processing.


Sources 3
Provider

About LG AI Research