EXAONE-Deep-7.8B is an open-weight causal language model from LG AI Research, built specifically for reasoning-heavy text tasks. Its main areas are mathematics, science problem solving, coding, and bilingual English-Korean generation. The model has approximately 7.8 billion parameters, or about 6.98 billion excluding embeddings, making it substantially smaller than many frontier hosted models and more practical for local experimentation.
Unlike a consumer chatbot, EXAONE-Deep-7.8B is primarily a model checkpoint for researchers and developers. It can be downloaded from Hugging Face and run with local inference software such as Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio. That gives users control over deployment, but also means they must provide compatible hardware, configure the runtime, and handle serving and safety considerations themselves.
What EXAONE-Deep-7.8B is
EXAONE-Deep-7.8B is a reasoning-specialized fine-tune of EXAONE-3.5-7.8B-Instruct. LG AI Research trained it with supervised fine-tuning, direct preference optimization, and online reinforcement learning. In practical terms, the model was adapted to spend more effort on multi-step problems instead of only producing short instruction-following responses.
The model belongs to LG AI Research's EXAONE family, but it has a narrower role than the organization's broader multimodal and enterprise-oriented releases. EXAONE-Deep-7.8B is text-only: it accepts text and produces text. It is not the appropriate EXAONE model for image understanding, document images, visual reasoning, or other multimodal workloads.
Its public release status is current and accessible as an open-weight research model. The model card and repository provide the primary usage information, including the model files, chat-template guidance, evaluation settings, and supported deployment approaches.
Core specifications and limits
| Specification | Verified detail |
|---|---|
| Provider | LG AI Research |
| Model type | Reasoning-focused causal language model |
| Parameters | Approximately 7.8 billion; approximately 6.98 billion excluding embeddings |
| Languages | English and Korean |
| Context length | 32,768 tokens |
| Maximum output | 32,768 tokens in the supplied model specification |
| Input | Text |
| Output | Text |
| Weights | BF16 configuration |
| License | EXAONE AI Model License Agreement 1.1 - NC |
A 32,768-token context window is large enough for lengthy prompts, source code, mathematical derivations, and substantial documents, although the usable amount depends on the inference configuration and available memory. The stated maximum output is also 32,768 tokens, but generating such a long response can be slow and memory-intensive. In ordinary use, shorter output limits are likely to be more practical.
The official configuration specifies 32 layers, 32 query-attention heads, eight key-value heads, and a 102,400-token vocabulary. Grouped-query attention uses fewer key-value heads than query heads, a design that can reduce inference memory requirements compared with some conventional attention configurations. These are model-architecture details rather than guarantees of a particular speed on a user's hardware.
Reasoning, mathematics and coding performance
Reasoning is the model's defining capability. LG AI Research reports results including 94.8 on MATH-500, 70.0 pass@1 on AIME 2024, 83.3 consensus@64 on AIME 2024, 59.6 pass@1 on AIME 2025, 76.7 consensus@64 on AIME 2025, 89.9 on CSAT Math 2025, 62.6 on GPQA Diamond, and 55.2 on LiveCodeBench.
These results indicate a model aimed at multi-step mathematical and scientific problems rather than simple conversational completion. Pass@1 measures whether a first sampled answer is correct, while consensus@64 reflects the result obtained when many samples are generated and the most consistent answer is selected. The two measures should not be treated as equivalent, and benchmark results do not guarantee the same performance on a particular prompt.
For coding, the reported LiveCodeBench result supports using the model for code-generation experiments, algorithmic problems, debugging assistance, and evaluation of reasoning over program logic. However, the supplied specifications mark tool use as unsupported. The model should therefore be treated as a text-based coding assistant, not as an agent that can independently execute code, inspect a repository, call external functions, or browse the web.
LG AI Research's recommended prompting guidance is especially relevant to reasoning workflows. The official instructions recommend using the supplied chat template, beginning reasoning with the <thought> marker, avoiding system prompts, and using temperature 0.6 with top_p 0.95 for evaluation. These recommendations should be followed when attempting to reproduce reported behavior. They are provider guidance, not universal requirements for every deployment.
Deployment and developer use
EXAONE-Deep-7.8B is intended for local or self-hosted inference. Supported deployment paths identified in the supplied research include Hugging Face Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio. The choice depends on whether the user prioritizes a research workflow, a high-throughput server, broad hardware compatibility, or a simpler desktop interface.
Transformers is a conventional starting point for researchers who want direct access to model configuration and generation settings. vLLM and SGLang are more relevant when serving the model to multiple requests or integrating it into an application. llama.cpp, Ollama, and LM Studio can make local experimentation more accessible, subject to the availability and quality of a compatible model conversion.
The model supports streaming according to the supplied specification, and fine-tuning is also listed as supported. Those flags describe the available model and deployment ecosystem rather than a hosted service commitment. Users still need to verify the exact runtime, quantization, hardware, and fine-tuning method they plan to use.
There is no identified official hosted API token price for EXAONE-Deep-7.8B. The model should not be presented as having a standard monthly plan or a first-party pay-as-you-go endpoint. A team that needs managed hosting must arrange its own infrastructure or use a third-party service if one supports the model and its license.
Modalities and feature boundaries
The model has text input and text output only. It does not provide image, audio, or video input; it does not generate images, audio, or video; and it does not produce embeddings or other non-text output according to the supplied model record.
It also lacks verified built-in web search, real-time data access, structured-output mode, batch API support, and function or tool calling. This makes it unsuitable as a ready-made web-grounded assistant or autonomous workflow agent. Developers can build surrounding application logic, but the model itself should not be assumed to call tools or return schema-constrained JSON reliably.
Its training-data knowledge cutoff has not been specified by LG AI Research. The model documentation warns that responses may not reflect the latest information. Even when the model gives a confident answer, current facts should be checked against authoritative sources rather than inferred from the model's fluency.
License, cost and commercial use
The EXAONE AI Model License Agreement 1.1 - NC provides research-use rights and restricts commercial use unless LG AI Research grants separate permission. This is the most important practical distinction between downloading the weights and using them in a commercial product. A business should review the complete license and obtain the necessary authorization before deploying the model for paid services, internal commercial operations, or other uses that may fall within the restriction.
The weights are available for research use through Hugging Face, but free access to the files does not mean unrestricted commercial use. There is no verified provider-advertised input or output price for this model, so cost is determined mainly by hardware, electricity, engineering time, storage, and any third-party hosting arrangement.
Main strengths and limitations
Strengths
- Reasoning specialization: The model is explicitly optimized for extended mathematical, scientific, and coding problems.
- Local control: Open weights allow research, evaluation, and self-hosted experimentation without depending on a single hosted endpoint.
- Moderate model size: At approximately 7.8 billion parameters, it is more approachable for local deployment than much larger reasoning models, although exact hardware needs vary.
- Korean and English support: It is relevant to bilingual research and applications involving Korean-language text.
- Broad runtime compatibility: The documented ecosystem includes Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, Ollama, and LM Studio.
Limitations
- Non-commercial license: Commercial use requires separate authorization.
- Text-only operation: It cannot directly analyze images or produce media.
- No built-in tools: Web search, code execution, function calling, and real-time data access are not verified capabilities.
- Self-hosting responsibility: Users must manage hardware, runtime configuration, scaling, monitoring, and security.
- Unspecified knowledge cutoff: It should not be relied on for current information without external verification.
- Reasoning latency: Extended generation can be slower and more expensive to run than a smaller non-reasoning model, particularly when producing long internal reasoning or multiple samples.
When to choose EXAONE-Deep-7.8B
Choose EXAONE-Deep-7.8B when the priority is local research into mathematical reasoning, science questions, coding evaluation, or English-Korean text generation. It is particularly suitable for a team that wants to inspect model behavior, reproduce published benchmarks, experiment with inference settings, or keep prompts and outputs within infrastructure it controls.
It is also a reasonable candidate when predictable access to open weights matters more than a polished user interface. The smaller parameter count can make experimentation more practical than using a much larger frontier model, although the actual speed and memory footprint depend on precision, quantization, batch size, and hardware. The editorial assessment supplied for this profile rates its reasoning and coding suitability highly and its cost position favorably relative to larger models; those scores are comparative editorial estimates, not LG AI Research measurements.
Another type of model is more appropriate when the application needs image or document understanding, voice, web search, current information, tool calling, hosted scaling, or commercial redistribution rights. A general-purpose hosted model may also be preferable for a consumer assistant because EXAONE-Deep-7.8B does not provide a mature subscription product or a documented first-party API experience. Within LG AI Research's wider EXAONE work, newer multimodal releases serve different purposes, but they should not be treated as interchangeable with this text-only reasoning checkpoint.
Practical verdict
EXAONE-Deep-7.8B is best understood as a focused research model rather than an all-purpose AI service. Its strongest case is local, text-based reasoning over mathematics, science, and code, especially for users who value open weights and Korean-language capability. Its 32,768-token context and broad local-runtime support make it useful for serious experimentation.
The trade-off is equally clear: the model has no verified hosted API pricing, no native multimodal or tool features, no stated knowledge cutoff, and a non-commercial license. Those constraints make it a strong candidate for research and self-hosted evaluation, but a poor default for a commercial assistant, real-time information product, or multimodal application without additional licensing and surrounding infrastructure.

