EXAONE Deep

EXAONE-Deep-7.8B

by LG AI Research · Current and accessible open-weight research release

A practical profile of LG AI Research's EXAONE-Deep-7.8B, covering its reasoning focus, reported mathematics and coding benchmarks, 32,768-token context, local deployment options, text-only limits, missing hosted API pricing, and research-only licensing.

Text Reasoning Coding
EXAONE-Deep-7.8B is LG AI Research's compact reasoning model for users who want to run an advanced language model locally rather than depend on a consumer chatbot or hosted API. The model focuses on extended problem solving in mathematics, science, and programming, and supports both English and Korean text. It is available as an open-weight research release, but its non-commercial license and lack of identified official hosted API pricing are important limitations for production users.
Outputs

What EXAONE-Deep-7.8B can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

8/10 Reasoning
8/10 Coding
7/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family EXAONE Deep
Model type Reasoning
Context window 33K tokens
Maximum output 33K tokens
Release date 2025-03-18
Status Current and accessible open-weight research release
Knowledge cutoff notes

LG AI Research's official model documentation does not specify a precise training-data knowledge cutoff for EXAONE-Deep-7.8B. The model card warns that responses may not reflect the latest information.

Model notes

EXAONE-Deep-7.8B is an open-weight causal language model with approximately 7.8 billion parameters, including 6.98 billion parameters excluding embeddings. The official configuration specifies 32 layers, grouped-query attention with 32 query heads and 8 key-value heads, a 102,400-token vocabulary, BF16 weights, and a 32,768-token context length. It is a reasoning-specialized fine-tune of EXAONE-3.5-7.8B-Instruct and was trained using supervised fine-tuning, direct preference optimization, and online reinforcement learning. LG AI Research reports MATH-500 94.8, AIME 2024 pass@1 70.0, AIME 2024 consensus@64 83.3, AIME 2025 pass@1 59.6, AIME 2025 consensus@64 76.7, CSAT Math 2025 89.9, GPQA Diamond 62.6, and LiveCodeBench 55.2. The model is available through Hugging Face and can be deployed with Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio. Official guidance recommends using the chat template, beginning reasoning with the <thought> marker, avoiding system prompts, and using temperature 0.6 with top_p 0.95 for evaluation. The EXAONE AI Model License Agreement 1.1 - NC grants research-use rights and prohibits commercial use unless separately authorized. No official hosted API token pricing was identified. Editorial scores are comparative estimates, not vendor-provided ratings.

Model guide

EXAONE-Deep-7.8B: Open-Weight Reasoning for Mathematics, Science and Code

EXAONE-Deep-7.8B is a 7.8-billion-parameter open-weight reasoning language model from LG AI Research. It is designed for mathematics, science, coding, and Korean- and English-language text generation, with a 32,768-token context window and local deployment through established inference tools. Its main trade-off is that it is a research-focused, non-commercial release rather than a hosted general-purpose assistant or broadly documented commercial API.

EXAONE-Deep-7.8B is an open-weight causal language model from LG AI Research, built specifically for reasoning-heavy text tasks. Its main areas are mathematics, science problem solving, coding, and bilingual English-Korean generation. The model has approximately 7.8 billion parameters, or about 6.98 billion excluding embeddings, making it substantially smaller than many frontier hosted models and more practical for local experimentation.

Unlike a consumer chatbot, EXAONE-Deep-7.8B is primarily a model checkpoint for researchers and developers. It can be downloaded from Hugging Face and run with local inference software such as Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio. That gives users control over deployment, but also means they must provide compatible hardware, configure the runtime, and handle serving and safety considerations themselves.

What EXAONE-Deep-7.8B is

EXAONE-Deep-7.8B is a reasoning-specialized fine-tune of EXAONE-3.5-7.8B-Instruct. LG AI Research trained it with supervised fine-tuning, direct preference optimization, and online reinforcement learning. In practical terms, the model was adapted to spend more effort on multi-step problems instead of only producing short instruction-following responses.

The model belongs to LG AI Research's EXAONE family, but it has a narrower role than the organization's broader multimodal and enterprise-oriented releases. EXAONE-Deep-7.8B is text-only: it accepts text and produces text. It is not the appropriate EXAONE model for image understanding, document images, visual reasoning, or other multimodal workloads.

Its public release status is current and accessible as an open-weight research model. The model card and repository provide the primary usage information, including the model files, chat-template guidance, evaluation settings, and supported deployment approaches.

Core specifications and limits

SpecificationVerified detail
ProviderLG AI Research
Model typeReasoning-focused causal language model
ParametersApproximately 7.8 billion; approximately 6.98 billion excluding embeddings
LanguagesEnglish and Korean
Context length32,768 tokens
Maximum output32,768 tokens in the supplied model specification
InputText
OutputText
WeightsBF16 configuration
LicenseEXAONE AI Model License Agreement 1.1 - NC

A 32,768-token context window is large enough for lengthy prompts, source code, mathematical derivations, and substantial documents, although the usable amount depends on the inference configuration and available memory. The stated maximum output is also 32,768 tokens, but generating such a long response can be slow and memory-intensive. In ordinary use, shorter output limits are likely to be more practical.

The official configuration specifies 32 layers, 32 query-attention heads, eight key-value heads, and a 102,400-token vocabulary. Grouped-query attention uses fewer key-value heads than query heads, a design that can reduce inference memory requirements compared with some conventional attention configurations. These are model-architecture details rather than guarantees of a particular speed on a user's hardware.

Reasoning, mathematics and coding performance

Reasoning is the model's defining capability. LG AI Research reports results including 94.8 on MATH-500, 70.0 pass@1 on AIME 2024, 83.3 consensus@64 on AIME 2024, 59.6 pass@1 on AIME 2025, 76.7 consensus@64 on AIME 2025, 89.9 on CSAT Math 2025, 62.6 on GPQA Diamond, and 55.2 on LiveCodeBench.

These results indicate a model aimed at multi-step mathematical and scientific problems rather than simple conversational completion. Pass@1 measures whether a first sampled answer is correct, while consensus@64 reflects the result obtained when many samples are generated and the most consistent answer is selected. The two measures should not be treated as equivalent, and benchmark results do not guarantee the same performance on a particular prompt.

For coding, the reported LiveCodeBench result supports using the model for code-generation experiments, algorithmic problems, debugging assistance, and evaluation of reasoning over program logic. However, the supplied specifications mark tool use as unsupported. The model should therefore be treated as a text-based coding assistant, not as an agent that can independently execute code, inspect a repository, call external functions, or browse the web.

LG AI Research's recommended prompting guidance is especially relevant to reasoning workflows. The official instructions recommend using the supplied chat template, beginning reasoning with the <thought> marker, avoiding system prompts, and using temperature 0.6 with top_p 0.95 for evaluation. These recommendations should be followed when attempting to reproduce reported behavior. They are provider guidance, not universal requirements for every deployment.

Deployment and developer use

EXAONE-Deep-7.8B is intended for local or self-hosted inference. Supported deployment paths identified in the supplied research include Hugging Face Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio. The choice depends on whether the user prioritizes a research workflow, a high-throughput server, broad hardware compatibility, or a simpler desktop interface.

Transformers is a conventional starting point for researchers who want direct access to model configuration and generation settings. vLLM and SGLang are more relevant when serving the model to multiple requests or integrating it into an application. llama.cpp, Ollama, and LM Studio can make local experimentation more accessible, subject to the availability and quality of a compatible model conversion.

The model supports streaming according to the supplied specification, and fine-tuning is also listed as supported. Those flags describe the available model and deployment ecosystem rather than a hosted service commitment. Users still need to verify the exact runtime, quantization, hardware, and fine-tuning method they plan to use.

There is no identified official hosted API token price for EXAONE-Deep-7.8B. The model should not be presented as having a standard monthly plan or a first-party pay-as-you-go endpoint. A team that needs managed hosting must arrange its own infrastructure or use a third-party service if one supports the model and its license.

Modalities and feature boundaries

The model has text input and text output only. It does not provide image, audio, or video input; it does not generate images, audio, or video; and it does not produce embeddings or other non-text output according to the supplied model record.

It also lacks verified built-in web search, real-time data access, structured-output mode, batch API support, and function or tool calling. This makes it unsuitable as a ready-made web-grounded assistant or autonomous workflow agent. Developers can build surrounding application logic, but the model itself should not be assumed to call tools or return schema-constrained JSON reliably.

Its training-data knowledge cutoff has not been specified by LG AI Research. The model documentation warns that responses may not reflect the latest information. Even when the model gives a confident answer, current facts should be checked against authoritative sources rather than inferred from the model's fluency.

License, cost and commercial use

The EXAONE AI Model License Agreement 1.1 - NC provides research-use rights and restricts commercial use unless LG AI Research grants separate permission. This is the most important practical distinction between downloading the weights and using them in a commercial product. A business should review the complete license and obtain the necessary authorization before deploying the model for paid services, internal commercial operations, or other uses that may fall within the restriction.

The weights are available for research use through Hugging Face, but free access to the files does not mean unrestricted commercial use. There is no verified provider-advertised input or output price for this model, so cost is determined mainly by hardware, electricity, engineering time, storage, and any third-party hosting arrangement.

Main strengths and limitations

Strengths

  • Reasoning specialization: The model is explicitly optimized for extended mathematical, scientific, and coding problems.
  • Local control: Open weights allow research, evaluation, and self-hosted experimentation without depending on a single hosted endpoint.
  • Moderate model size: At approximately 7.8 billion parameters, it is more approachable for local deployment than much larger reasoning models, although exact hardware needs vary.
  • Korean and English support: It is relevant to bilingual research and applications involving Korean-language text.
  • Broad runtime compatibility: The documented ecosystem includes Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio.

Limitations

  • Non-commercial license: Commercial use requires separate authorization.
  • Text-only operation: It cannot directly analyze images or produce media.
  • No built-in tools: Web search, code execution, function calling, and real-time data access are not verified capabilities.
  • Self-hosting responsibility: Users must manage hardware, runtime configuration, scaling, monitoring, and security.
  • Unspecified knowledge cutoff: It should not be relied on for current information without external verification.
  • Reasoning latency: Extended generation can be slower and more expensive to run than a smaller non-reasoning model, particularly when producing long internal reasoning or multiple samples.

When to choose EXAONE-Deep-7.8B

Choose EXAONE-Deep-7.8B when the priority is local research into mathematical reasoning, science questions, coding evaluation, or English-Korean text generation. It is particularly suitable for a team that wants to inspect model behavior, reproduce published benchmarks, experiment with inference settings, or keep prompts and outputs within infrastructure it controls.

It is also a reasonable candidate when predictable access to open weights matters more than a polished user interface. The smaller parameter count can make experimentation more practical than using a much larger frontier model, although the actual speed and memory footprint depend on precision, quantization, batch size, and hardware. The editorial assessment supplied for this profile rates its reasoning and coding suitability highly and its cost position favorably relative to larger models; those scores are comparative editorial estimates, not LG AI Research measurements.

Another type of model is more appropriate when the application needs image or document understanding, voice, web search, current information, tool calling, hosted scaling, or commercial redistribution rights. A general-purpose hosted model may also be preferable for a consumer assistant because EXAONE-Deep-7.8B does not provide a mature subscription product or a documented first-party API experience. Within LG AI Research's wider EXAONE work, newer multimodal releases serve different purposes, but they should not be treated as interchangeable with this text-only reasoning checkpoint.

Practical verdict

EXAONE-Deep-7.8B is best understood as a focused research model rather than an all-purpose AI service. Its strongest case is local, text-based reasoning over mathematics, science, and code, especially for users who value open weights and Korean-language capability. Its 32,768-token context and broad local-runtime support make it useful for serious experimentation.

The trade-off is equally clear: the model has no verified hosted API pricing, no native multimodal or tool features, no stated knowledge cutoff, and a non-commercial license. Those constraints make it a strong candidate for research and self-hosted evaluation, but a poor default for a commercial assistant, real-time information product, or multimodal application without additional licensing and surrounding infrastructure.


Answers to Frequently Asked Questions

What hardware and model specifications should developers consider?
The model has approximately 7.8 billion parameters, uses a BF16 configuration, and includes 32 layers, grouped-query attention, and a 102,400-token vocabulary. Actual memory and speed requirements depend on precision, quantization, batch size, context length, and hardware, so developers should test their chosen runtime and configuration.
Is EXAONE-Deep-7.8B free for commercial use?
No. The model is released under the EXAONE AI Model License Agreement 1.1 - NC, which provides research-use rights and restricts commercial use unless LG AI Research grants separate permission. Businesses should review the full license and obtain authorization before commercial deployment.
What are the main capabilities and limitations of EXAONE-Deep-7.8B?
EXAONE-Deep-7.8B supports text input and text output, has a 32,768-token context length, and is optimized for multi-step reasoning in mathematics, science, and code. It does not provide verified built-in web search, real-time data access, code execution, function calling, structured-output mode, or multimodal input and output.
What is EXAONE-Deep-7.8B designed for?
EXAONE-Deep-7.8B is an open-weight, reasoning-focused causal language model from LG AI Research. It is designed mainly for mathematics, science problem solving, coding, and bilingual English-Korean text generation.
How can EXAONE-Deep-7.8B be deployed locally?
The model can be downloaded from Hugging Face and run with local inference tools including Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio. Users must provide compatible hardware and manage runtime configuration, serving, monitoring, and security.


Sources 4
Provider

About LG AI Research