EXAONE Deep

EXAONE-Deep-2.4B

by LG AI Research · Current and publicly downloadable open-weight model

EXAONE-Deep-2.4B is a compact open-weight reasoning model from LG AI Research, fine-tuned for mathematics, science, and coding. It offers a 32,768-token context window, downloadable weights, and support for multiple local inference frameworks, but it is text-only and has no documented first-party token-priced API or native web and tool features.

Text Reasoning Coding
EXAONE-Deep-2.4B is LG AI Research's smallest EXAONE Deep reasoning model. Fine-tuned from EXAONE 3.5-2.4B-Instruct, it focuses on mathematical and coding problems rather than acting as a broad consumer chatbot. The downloadable weights, 32,768-token context window, and support for several local inference frameworks make it especially relevant to developers and researchers who want a reasoning model without relying on a hosted API.
Outputs

What EXAONE-Deep-2.4B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

7/10 Reasoning
6/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family EXAONE Deep
Model type Reasoning
Context window 33K tokens
Maximum output 33K tokens
Release date 2025-03-18
Status Current and publicly downloadable open-weight model
Knowledge cutoff notes

No exact knowledge-cutoff date was published in the reviewed official model card, repository, configuration, or paper. The model documentation explicitly warns that the model does not reflect the latest information.

Model notes

Canonical Hugging Face identifier: LGAI-EXAONE/EXAONE-Deep-2.4B. The model has approximately 2.14B parameters excluding embeddings, 30 layers, grouped-query attention with 32 query heads and 8 key-value heads, and tied word embeddings. It is a reasoning-focused fine-tune of EXAONE 3.5-2.4B-Instruct using supervised fine-tuning, direct preference optimization, and online reinforcement learning. Official documentation recommends starting reasoning prompts with the <thought> tag, avoiding system prompts, and using temperature 0.6 with top_p 0.95. The model is downloadable rather than offered with a documented first-party token-priced API. It is licensed under the EXAONE AI Model License Agreement 1.1 - NC. Editorial scores are comparative estimates, not vendor specifications.

Model guide

EXAONE-Deep-2.4B: Compact Open-Weight Reasoning for Local Math and Coding

EXAONE-Deep-2.4B is a compact open-weight reasoning language model from LG AI Research. It is designed for mathematics, science, coding, and deliberate problem solving, while its 2.4-billion-parameter class makes local deployment more practical than using a much larger reasoning model.

What is EXAONE-Deep-2.4B?

EXAONE-Deep-2.4B is an open-weight text-generation model developed by LG AI Research. The canonical model identifier is LGAI-EXAONE/EXAONE-Deep-2.4B, and the weights are distributed through the official Hugging Face repository. It belongs to the EXAONE Deep family, alongside larger 7.8B and 32B variants, but this 2.4B model is intended for comparatively resource-conscious use.

The model is a reasoning-focused fine-tune of EXAONE 3.5-2.4B-Instruct. LG AI Research describes the training process as involving supervised fine-tuning, direct preference optimization, and online reinforcement learning. In practical terms, the model is intended to spend more effort working through difficult problems before producing an answer, particularly in mathematics, science, and programming.

This is not a consumer subscription service or a general-purpose provider chatbot. It is a downloadable model that users can run themselves, subject to its license and their available hardware.

Where it fits in the EXAONE lineup

Within the EXAONE Deep series, EXAONE-Deep-2.4B is the smallest published variant covered by the supplied research. Its smaller size gives it a practical speed and deployment advantage over larger reasoning models, although larger models may offer more capacity for difficult or broad tasks. LG AI Research recommends EXAONE 3.5 Instruct models for a wider range of everyday instruction-following use cases, which is an important distinction: EXAONE-Deep-2.4B is optimized for reasoning, not for being the most flexible assistant in the EXAONE catalog.

The model should also be distinguished from LG AI Research's multimodal EXAONE releases. EXAONE-Deep-2.4B is text-only and does not provide the image understanding or document-vision features associated with the provider's vision-language work.

Technical specifications and limits

SpecificationVerified detail
ProviderLG AI Research
Model familyEXAONE Deep
Model sizeApproximately 2.4B parameters; about 2.14B excluding embeddings
Architecture details30 transformer layers, grouped-query attention, 32 query heads, and 8 key-value heads
Vocabulary102,400 tokens
Context length32,768 tokens
Maximum generation settingUp to 32,768 new tokens in the documented usage example
InputText
OutputText, including reasoning traces when prompted according to the model guidance
WeightsDownloadable from Hugging Face
LicenseEXAONE AI Model License Agreement 1.1 - NC

The 32,768-token context window is the maximum documented context length, not a promise that every local deployment will have enough memory to use it efficiently. Actual speed and usable context depend on the inference engine, numerical precision, hardware, and settings chosen by the operator. The documented maximum of up to 32,768 new tokens is a generation setting, so it should not be confused with the amount of input that can be supplied at the same time.

Reasoning and benchmark results

Reasoning is the model's central purpose. It is designed to work through multi-step problems rather than only predict a short answer from a simple prompt. LG AI Research's recommended prompting approach uses a <thought> tag to signal reasoning-oriented generation. The supplied model notes also recommend avoiding system prompts and starting with a temperature of 0.6 and a top-p value of 0.95.

LG AI Research reports the following results for EXAONE-Deep-2.4B: 92.3 pass@1 on MATH-500, 52.5 pass@1 on AIME 2024, 47.9 pass@1 on AIME 2025, 54.3 on GPQA Diamond, and 46.6 on LiveCodeBench. These are provider-reported benchmark results, not independent guarantees of performance on a user's own problems. The model was reported to outperform DeepSeek-R1-Distill-Qwen-1.5B on the listed mathematics, science, and coding comparisons.

For a beginner, the practical interpretation is that this model is a better fit for tasks requiring intermediate steps, calculations, proofs, algorithm design, or code reasoning than for simple conversational responses. Benchmark scores should still be treated as directional because prompting, sampling settings, evaluation versions, and hardware can affect results.

Mathematics, coding, and language use

Mathematics and coding are the clearest target workloads. Suitable examples include asking the model to solve a structured algebra problem, explain a mathematical approach, identify an error in a program, generate a small algorithm, or reason through a programming challenge. Its training and evaluation emphasis makes it more suitable for these tasks than a small model optimized only for conversational fluency.

The model can be useful for English-Korean experimentation and compact bilingual applications, as reflected in the supplied evaluation notes. However, the research does not establish a complete language-coverage list or guarantee equal quality across languages. Users should test the particular Korean, English, or mixed-language workload they care about.

Reasoning traces can make the output more informative for debugging or study, but they can also make responses longer and slower. A generated explanation should not be treated as proof that every intermediate step is correct. The model card warns that outputs may contain factual errors, bias, inappropriate content, or outdated information.

Modalities, tools, and structured output

EXAONE-Deep-2.4B accepts text and produces text. It does not natively process images, audio, or video, and it does not generate non-text media. It therefore cannot directly inspect a photograph, hear a recording, or analyze a video without an external system converting that material into text first.

The supplied research does not verify native tool calling, function calling, web search, JSON mode, or a built-in code-execution environment for this model. A developer could potentially place the model inside a larger application that handles tools externally, but that would be an application-layer integration rather than a verified model capability. The model also has no documented knowledge-cutoff date and should not be used as a current-information system without external retrieval and verification.

Deployment, speed, and cost

LG AI Research documents local and self-hosted deployment rather than a provider-operated, token-priced API for this model. The official usage path supports Transformers with remote code enabled, and the model can also be served with vLLM and SGLang. Additional documented deployment options include TensorRT-LLM, llama.cpp, Ollama, and LM Studio. Streaming generation is available through standard generation utilities such as TextIteratorStreamer.

The official example uses bfloat16 weights, automatic device mapping, Transformers 4.43.1 or later, and a maximum generation setting of up to 32,768 new tokens. These details are useful starting points, but they do not establish a single hardware requirement or a guaranteed response speed. Quantization and different serving engines may reduce memory demands, while long reasoning traces and large contexts can increase latency.

There is no verified recurring subscription price or first-party per-token API price in the supplied research. The main financial advantage is therefore the availability of downloadable weights for users who already have suitable hardware or infrastructure. Self-hosting still carries costs for compute, storage, operations, and engineering. It may be less economical than a hosted API for occasional use, but more attractive for repeated local experimentation, privacy-sensitive workflows, or deployments where an external endpoint is undesirable.

Main strengths and limitations

Strengths

  • Reasoning focus: The model is purpose-built for deliberate mathematical, scientific, and coding work.
  • Compact scale: Its 2.4B-class size is easier to experiment with locally than larger reasoning models, although the required resources depend on precision and context length.
  • Open-weight access: Developers can download the model and select from several local serving frameworks instead of depending on a documented hosted API.
  • Long context for its size: The 32,768-token context window is useful for longer problem statements, code, and explanations.
  • Transparent deployment options: Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio are identified in the documentation.

Limitations

  • Text-only operation: It cannot directly handle image, audio, or video input.
  • Reasoning is not general instruction following: For broad everyday assistant behavior, LG AI Research points users toward EXAONE 3.5 Instruct models instead.
  • No verified built-in current information: There is no documented web search capability or exact knowledge-cutoff date.
  • No verified native tool or JSON mode: The supplied documentation does not establish function calling, structured output, or code execution as model features.
  • License restrictions: The EXAONE AI Model License Agreement 1.1 - NC must be reviewed before commercial use or redistribution.
  • Uncertain factual reliability: As with other reasoning models, longer explanations can still contain incorrect steps or conclusions.

When to choose EXAONE-Deep-2.4B

Choose this model when you want a downloadable, relatively compact reasoning model for local mathematics, coding experiments, education, research, or bilingual English-Korean testing. It is especially compelling when avoiding a hosted API matters, when you need to inspect or control the deployment, or when a larger reasoning model would be unnecessarily expensive or slow.

A larger reasoning model may be more appropriate for unusually difficult problems, complex software engineering, or workloads where maximum answer quality matters more than local efficiency. A general instruction-tuned model is a better choice for broad assistant tasks, natural conversational interaction, and routine instruction following. A multimodal model is required for images, audio, or video. A hosted service with retrieval and tool integration is more suitable when the application needs current web information, managed scaling, or verified function calling.

Availability and licensing

EXAONE-Deep-2.4B is described as current and publicly downloadable through the official LGAI-EXAONE Hugging Face repository and documented in the official EXAONE Deep GitHub repository. Before deploying it, users should read the EXAONE AI Model License Agreement 1.1 - NC, confirm whether their intended use is permitted, and account for the model's warnings about factual errors, bias, inappropriate content, and outdated information.


Answers to Frequently Asked Questions

What license applies to EXAONE-Deep-2.4B?
EXAONE-Deep-2.4B is distributed under the EXAONE AI Model License Agreement 1.1 - NC. Users should review the license carefully to determine whether commercial use, redistribution, and their intended deployment are permitted.
How can EXAONE-Deep-2.4B be deployed locally?
The model can be downloaded from Hugging Face and deployed with Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, or LM Studio. Actual memory requirements and speed depend on hardware, numerical precision, quantization, inference engine, context length, and generation settings.
What can EXAONE-Deep-2.4B be used for?
EXAONE-Deep-2.4B is suitable for structured mathematics, coding challenges, algorithm design, program debugging, proofs, education, research, and compact bilingual English-Korean experiments. It is optimized for reasoning rather than broad conversational assistance.
What are the main technical specifications of EXAONE-Deep-2.4B?
The model has approximately 2.4 billion parameters, 30 transformer layers, grouped-query attention, 32 query heads, 8 key-value heads, and a 102,400-token vocabulary. It supports a documented context length of up to 32,768 tokens and accepts and generates text.
What is EXAONE-Deep-2.4B?
EXAONE-Deep-2.4B is an open-weight, text-only reasoning model developed by LG AI Research. Its canonical identifier is LGAI-EXAONE/EXAONE-Deep-2.4B, and it is designed primarily for mathematics, science, programming, and other multi-step reasoning tasks.


Sources 5
Provider

About LG AI Research