EXAONE Deep

EXAONE-Deep-32B

by LG AI Research · Current open-weight research model; downloadable and locally deployable

EXAONE-Deep-32B is LG AI Research’s 32-billion-parameter open-weight reasoning model for mathematics, science, and coding. It offers a 32,768-token context window, local deployment through several inference tools, and provider-reported reasoning benchmarks. The model is text-only, has no documented hosted API pricing or built-in web search and tool use, and is licensed primarily for research unless a separate commercial agreement is obtained.

Text Reasoning Coding
EXAONE-Deep-32B is the largest model in LG AI Research’s EXAONE Deep series. Released on March 18, 2025, it was developed from EXAONE 3.5 32B Instruct and further trained to improve multi-step reasoning in mathematics, science, and programming. The model is available as downloadable weights, making it suitable for local and on-premise experimentation, but its research-focused license and lack of a published first-party hosted API price are important considerations for anyone evaluating it for production use.
Outputs

What EXAONE-Deep-32B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
5/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family EXAONE Deep
Model type Reasoning
Context window 33K tokens
Release date 2025-03-18
Status Current open-weight research model; downloadable and locally deployable
Knowledge cutoff notes

LG AI Research states that the model does not reflect the latest information, but it does not publish a specific knowledge-cutoff date for EXAONE-Deep-32B.

Model notes

The canonical Hugging Face identifier is LGAI-EXAONE/EXAONE-Deep-32B. The model has approximately 30.95B parameters excluding embeddings, 64 layers, grouped-query attention with 40 query heads and 8 key-value heads, and a 102,400-token vocabulary. It is fine-tuned from EXAONE 3.5 32B Instruct using supervised fine-tuning, direct preference optimization, and online reinforcement learning. Official documentation supports Transformers and describes deployment with vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio. LG AI Research recommends EXAONE 3.5 Instruct models for broader practical use cases. The EXAONE AI Model License Agreement 1.1 - NC limits use to research purposes unless a separate commercial agreement is obtained. No provider-hosted API, batch API, prompt-caching service, or model-specific JSON/structured-output mode is documented. Knowledge cutoff is not published in the official model documentation.

Cost

Model pricing

Input No official hosted API price; downloadable weights for local or third-party deployment
Output No official hosted API price; downloadable weights for local or third-party deployment
Model guide

EXAONE-Deep-32B: An Open-Weight Model for Mathematics, Science and Coding

EXAONE-Deep-32B is LG AI Research’s 32-billion-parameter open-weight reasoning model for mathematical problem solving, scientific reasoning, and software coding. It offers a 32,768-token context window and local deployment options, but is distributed primarily for research use rather than as a conventional hosted chatbot or broadly documented API.

What is EXAONE-Deep-32B?

EXAONE-Deep-32B is a 32-billion-parameter reasoning language model from LG AI Research. Its canonical model identifier is LGAI-EXAONE/EXAONE-Deep-32B. The model generates text from text input and is designed primarily for difficult mathematical, scientific, and programming tasks rather than for image understanding, media generation, or general consumer chat.

It is the largest model in the EXAONE Deep family, alongside 2.4B and 7.8B variants. Within LG AI Research’s broader catalog, EXAONE-Deep-32B is positioned as a reasoning-focused open-weight release. That positioning matters: the model is intended for researchers and developers who want to download and operate the weights themselves, not for users looking for a finished subscription chatbot with integrated search, mobile applications, or a simple pay-as-you-go interface.

LG AI Research released the model on March 18, 2025. It was based on EXAONE 3.5 32B Instruct and subsequently optimized with supervised fine-tuning, direct preference optimization, and online reinforcement learning. In practical terms, these stages were intended to improve the model’s ability to work through problems rather than only produce a quick surface-level answer.

Technical specifications and context window

The model has approximately 30.95 billion parameters excluding embeddings, 64 transformer layers, grouped-query attention with 40 query heads and 8 key-value heads, and a 102,400-token vocabulary. Its documented context length is 32,768 tokens. A context window includes the user’s prompt, conversation history, intermediate reasoning content, and generated response, so long reasoning traces can reduce the space available for later turns.

SpecificationVerified detail
Model familyEXAONE Deep
ProviderLG AI Research
ParametersApproximately 30.95B excluding embeddings
Context length32,768 tokens
WeightsBF16
Architecture details64 layers; grouped-query attention; 40 query heads and 8 key-value heads
Input and outputText input and text generation output
Knowledge cutoffNot published

There is no documented maximum output-token value separate from the 32,768-token context limit. The usable response length depends on the deployment configuration and on how much of the context is already occupied by the prompt and conversation. LG AI Research recommends managing earlier reasoning traces in multi-turn use because they can consume a substantial portion of the available context.

Reasoning, mathematics, science and coding

EXAONE-Deep-32B’s main distinction is its emphasis on reasoning. The model is intended to work through multi-step problems in areas such as mathematics, scientific problem solving, and programming. This makes it a more specialized choice than a general instruction model optimized mainly for everyday conversation, summarization, or broad assistant workloads.

In provider-published evaluations, LG AI Research reports scores of 95.7 on MATH-500, 72.1 on AIME 2024 pass@1, 65.8 on AIME 2025 pass@1, 94.5 on the CSAT Math 2025 evaluation, 66.1 on GPQA Diamond, and 59.5 on LiveCodeBench. These are provider-reported results under particular prompts, sampling settings, and benchmark procedures; they should not be treated as guaranteed performance for every task or deployment.

The coding capability is best understood as code generation and programming problem solving through text. The supplied specifications do not document a built-in code interpreter, execution sandbox, or first-party tool environment. As a result, the model can propose or explain code, but an application would need a separate execution system if it must compile, run, test, or inspect that code automatically.

Supported inputs, outputs and tools

EXAONE-Deep-32B is text-only. It does not natively accept images, audio, or video, and it does not generate images, audio, video, or other non-text media. This rules it out for multimodal document inspection, image question answering, speech workflows, and visual coding tasks unless another model is added to the system.

The supplied model information does not document native web search, function calling, or tool-use support. It also does not document a model-specific JSON mode or structured-output guarantee. Developers can still build applications around the model, but retrieval, function execution, schema validation, and other orchestration would need to be implemented outside the model itself.

Streaming generation is supported in the documented Transformers example through a text iterator streamer. Streaming can make a locally hosted application feel more responsive, but it does not change the model’s total generation speed or context capacity.

Deployment and availability

The primary access method is downloading the model weights from Hugging Face and running them locally or on managed infrastructure. Official documentation provides a Transformers example and describes deployment options including vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio. The range of supported tools gives experienced users flexibility across research workstations, servers, and local experimentation, although the practical hardware requirements depend on quantization, batching, sequence length, and the chosen inference engine.

LG AI Research recommends using the model’s chat template and beginning reasoning generation with the appropriate thought marker. Implementations should follow the official repository and model-card instructions rather than assuming that a generic prompt format will produce the intended behavior. For multi-turn conversations, removing or managing previous reasoning traces may help preserve context for later requests.

There is no official provider-hosted API price published for EXAONE-Deep-32B in the supplied research. It is not presented as a conventional recurring subscription product or as a broadly documented first-party token-priced API. The cost model is therefore primarily infrastructure-based: users operating the weights pay for their own hardware or hosting, while third-party platforms may set separate prices and terms.

License and important limitations

The model is distributed under the EXAONE AI Model License Agreement 1.1 - NC. According to the supplied research, access, downloading, modification, and distribution are permitted for research purposes under stated conditions. Commercial use of the model, derivatives, or output requires a separate commercial agreement with LG AI Research. Organizations should review the current license directly before using the model in a commercial product, internal business workflow, or customer-facing service.

The research-focused license is one of the most important limitations for production teams. Even if the model runs successfully on local infrastructure, technical deployability does not by itself establish permission for commercial deployment. Licensing, redistribution, derivative models, and generated output should be assessed separately.

The model may generate inaccurate, biased, harmful, or outdated text. Its knowledge-cutoff date is not published, and LG AI Research states that the model does not reflect the latest information. It has no documented first-party web-search capability, so it should not be relied upon for current events, changing regulations, live prices, or other time-sensitive facts without an external retrieval system and verification process.

Strengths and trade-offs

  • Reasoning specialization: Its training and published evaluations focus on mathematics, science, and coding rather than only conversational fluency.
  • Open-weight access: Downloadable weights allow local, private, and on-premise experimentation when the license and infrastructure requirements are acceptable.
  • Flexible deployment: Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio are identified as deployment options.
  • Large context for a local model: The 32,768-token context window supports substantial problem statements, code, and research material, although reasoning traces can use that capacity quickly.
  • Operational responsibility: Local deployment avoids dependence on a documented provider-hosted endpoint, but shifts hardware, scaling, monitoring, security, and maintenance responsibilities to the operator.
  • Limited product features: There is no documented native web search, tool execution, multimodal input, structured-output mode, or official model-specific hosted API.

In editorial terms, EXAONE-Deep-32B offers a compelling capability-versus-control trade-off for research users: it targets demanding reasoning while permitting local operation. The trade-off is speed and convenience. A 32B model generally requires substantially more compute and operational work than a smaller model, and the supplied information does not provide a fixed tokens-per-second figure. Users should benchmark it on their own hardware rather than assuming a particular response speed.

When to choose EXAONE-Deep-32B

Choose EXAONE-Deep-32B when the central requirement is a downloadable reasoning model for mathematical analysis, scientific problem solving, coding evaluation, or controlled local research. It is particularly relevant when an organization wants to inspect or manage the model locally, needs an open-weight workflow, or wants to experiment with inference engines such as vLLM, SGLang, llama.cpp, or Ollama.

It may also suit benchmarking and model research where reproducible local access is more important than a polished end-user interface. The 32B size provides a larger reasoning-oriented option within the EXAONE Deep family, while the smaller 2.4B and 7.8B siblings may be more appropriate when lower memory use or faster inference is the priority, subject to their own capabilities and licenses.

Another option may be more appropriate for ordinary customer support, broad conversational assistance, or applications that require current web information. LG AI Research itself recommends EXAONE 3.5 Instruct models for a wider range of practical use cases. A multimodal model should be selected for image, audio, or video input; a tool-enabled hosted model may be preferable when reliable function calling, web retrieval, managed scaling, or structured JSON output is central to the application.

Overall, EXAONE-Deep-32B is best viewed as a specialized open-weight reasoning model rather than a complete AI application. Its value comes from its focus on difficult text-based reasoning and the ability to deploy the weights through local or third-party infrastructure. Its research-only licensing position, text-only interface, unavailable hosted pricing, and lack of documented built-in tools make careful evaluation necessary before adopting it for production.


Answers to Frequently Asked Questions

Can EXAONE-Deep-32B be used commercially?
EXAONE-Deep-32B is distributed under the EXAONE AI Model License Agreement 1.1 - NC. Research use is permitted under stated conditions, while commercial use of the model, derivatives, or output requires a separate commercial agreement with LG AI Research. Organizations should review the current license before deployment.
Does EXAONE-Deep-32B support images, web search, function calling, or code execution?
No native support for images, audio, video, web search, function calling, or code execution is documented. The model accepts text and generates text, so retrieval, tool execution, code testing, and multimodal processing must be provided by external systems.
How can EXAONE-Deep-32B be deployed?
Users can download the model weights from Hugging Face and run EXAONE-Deep-32B locally or on managed infrastructure. Documented deployment options include Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio.
What is EXAONE-Deep-32B?
EXAONE-Deep-32B is a 32-billion-parameter, text-only reasoning language model from LG AI Research. Its canonical model identifier is LGAI-EXAONE/EXAONE-Deep-32B, and it is designed primarily for mathematics, science, and programming tasks.
What are the context length and technical specifications of EXAONE-Deep-32B?
EXAONE-Deep-32B has approximately 30.95 billion parameters excluding embeddings, 64 transformer layers, grouped-query attention with 40 query heads and 8 key-value heads, BF16 weights, and a 32,768-token context length. Its knowledge cutoff has not been published.


Sources 5
Provider

About LG AI Research