What EXAONE-3.5-7.8B-Instruct is
EXAONE-3.5-7.8B-Instruct is an instruction-tuned, decoder-only Transformer language model from LG AI Research. Instruction tuning means that the model has been adapted to respond to user requests rather than merely predict the next piece of text. In practical terms, it can be used for conversational responses, summaries, translations, document questions, code assistance, and mathematical problem solving.
The model is the 7.8-billion-parameter version of the EXAONE 3.5 family. The family also includes 2.4-billion- and 32-billion-parameter variants. The 7.8B model occupies a middle position: it is considerably more capable than the smallest family member for many demanding tasks, while remaining more practical to deploy locally than a 32B model. The marketed parameter count is 7.8 billion; the official model card reports 6.98 billion parameters when embeddings are excluded.
LG AI Research released EXAONE 3.5 on December 9, 2024. The exact model is distributed as open weights through the official Hugging Face model card and the EXAONE 3.5 repository.
Key specifications
| Specification | Verified detail |
|---|---|
| Provider | LG AI Research |
| Model family | EXAONE 3.5 |
| Model type | Instruction-tuned, decoder-only language model |
| Parameters | 7.8B marketed size; 6.98B excluding embeddings |
| Context length | 32,768 tokens |
| Languages | English and Korean |
| Architecture details | 32 layers, grouped-query attention, 32 query heads, 8 key-value heads |
| Vocabulary | 102,400 tokens |
| Input and output | Text input and text generation |
| Availability | Downloadable open weights for research use |
| Official per-token price | Not identified for this exact model |
The 32,768-token context window is the maximum context length identified in the supplied official materials. A context window includes the text supplied to the model and the text generated during a request, subject to the runtime and application configuration. The research does not specify a separate maximum output-token limit, so no independent output ceiling should be assumed.
Language and task capabilities
EXAONE-3.5-7.8B-Instruct is primarily a bilingual English-Korean model. Its focus makes it particularly relevant to applications that need Korean-language interaction, Korean-English translation, or workflows in which users and documents switch between both languages. The model can also support English-only tasks, but its most distinctive positioning is its combination of bilingual coverage and local deployment.
The model was evaluated for instruction following, long-context comprehension, mathematics, coding, and general knowledge. Those evaluation categories indicate the intended uses, but they should not be treated as a guarantee of a particular accuracy level. The supplied research does not provide a complete benchmark table for this page, and real-world performance will depend on prompting, language, quantization, runtime, and task complexity.
Long-context document work
A 32K context window allows an application to provide substantial reports, transcripts, policy documents, or retrieved passages in one request. Suitable tasks include summarization, document question answering, extracting information from long text, and comparing sections of a report. The model does not automatically retrieve current documents or browse the web. Search, retrieval, document ingestion, and citation handling must be added by the surrounding application.
Coding and mathematics
EXAONE-3.5-7.8B-Instruct is suitable for code generation, explanation, debugging assistance, and mathematical question answering. Its relatively compact size can make experimentation easier than with a larger model, particularly when a team wants to keep inference on local or private infrastructure. Generated code and calculations still require testing and verification; the model should not be treated as a compiler, proof system, or safety-critical calculator.
Deployment and inference options
The model is intended for users who want to control the inference environment. The official materials provide examples for Hugging Face Transformers and support deployment with vLLM and SGLang. Quantized AWQ and GGUF variants are also available, enabling use with llama.cpp, Ollama, LM Studio, and other compatible local runtimes.
In simple terms, quantization stores model weights with lower numerical precision to reduce memory use. This can make a model easier to run on available hardware, although the exact memory requirement depends on precision, context length, batch size, and runtime settings. Quantization may also affect output quality or speed. The research does not specify a single hardware requirement for every configuration, so deployment planning should be based on the chosen checkpoint and inference engine rather than the parameter count alone.
Local deployment can be useful for private document processing, internal experiments, and applications that cannot send prompts to a third-party hosted service. It also transfers more operational responsibility to the user: the deployment team must manage hardware, updates, access control, monitoring, prompt handling, and model safety.
Modalities, tools, and reasoning behavior
The official EXAONE-3.5-7.8B-Instruct release is text-only. It accepts text and generates text; it does not natively accept images, audio, or video, and it does not directly produce those media types. The model should therefore not be confused with multimodal EXAONE releases or with a separate vision-language system.
The supplied documentation does not identify native tool calling, function calling, enforced JSON or structured-output mode, provider-hosted web search, or built-in code execution for this exact model. An application can wrap the model with retrieval, tools, validators, or a structured prompting scheme, but those features would belong to the surrounding system rather than being verified native capabilities of the model.
Reasoning is best understood here as the model’s ability to work through instructions, mathematics, coding problems, and long documents. The supplied database assigns a reasoning score of 7 out of 10, but that is an editorial evaluation, not a score published by LG AI Research. The model may be a practical choice for ordinary analytical tasks, but users seeking the strongest available reasoning performance should test it against larger or more specialized alternatives.
Pricing and licensing
No official provider-hosted per-token input or output price was identified for EXAONE-3.5-7.8B-Instruct. Since the model is distributed as downloadable weights, the direct software cost is not presented as a conventional recurring subscription or API fee. Running it still creates infrastructure costs for hardware, hosting, electricity, storage, operations, and engineering.
Commercial use requires careful license review. Official materials describe the EXAONE 3.5 release as available for research purposes and refer users to the EXAONE AI Model License Agreement 1.1. Commercial deployment may require contacting LG AI Research or satisfying additional terms. Availability of downloadable weights should not be interpreted as unrestricted commercial permission.
Main strengths and limitations
Strengths
- English-Korean specialization: The model is designed for bilingual workflows rather than treating Korean as an incidental language.
- Long context: Its 32,768-token context length supports substantial documents, transcripts, and retrieved text.
- Local control: Open-weight distribution enables private or self-managed deployment instead of requiring a provider-hosted endpoint.
- Deployment flexibility: Transformers, vLLM, SGLang, llama.cpp, Ollama, and compatible tools can support different operational needs.
- Balanced model size: The 7.8B configuration offers a middle ground between small-model efficiency and larger-model capability.
Limitations
- Text only: Native image, audio, and video input or output are not supported.
- No verified built-in web access: Current information requires an external search or retrieval layer.
- No documented independent output limit: The context length is known, but a separate maximum output-token value was not identified.
- No confirmed native tool or JSON mode: Tool integration and strict structured output should not be assumed.
- Research-oriented licensing: Commercial use requires review of the applicable model agreement.
- Operational burden: Self-hosting requires users to provide suitable hardware and manage inference infrastructure.
- Known model risks: LG AI Research warns that responses may be inaccurate, biased, harmful, or outdated.
When to choose this model
Choose EXAONE-3.5-7.8B-Instruct when the main requirements are English-Korean text handling, long-context processing, local control, and an open-weight deployment path. It is a sensible candidate for a bilingual internal assistant, private document question-answering system, translation workflow, research prototype, coding helper, or Korean-language application that does not need native media understanding.
Its middle-sized configuration is especially relevant when a 32B model would impose too much hardware or latency overhead, while a smaller model would not provide enough quality for long documents, bilingual interaction, coding, or mathematics. Quantized versions may improve deployment efficiency, but the correct trade-off should be measured on the target workload rather than inferred from the model name.
Another option may be more appropriate when the application needs image or audio understanding, automatic web search, verified function calling, strict structured outputs, a managed API, or a mature consumer application with subscription support. A larger model may also be preferable for the hardest reasoning and coding tasks, while a smaller model may be better for high-throughput or severely resource-constrained applications. The EXAONE 3.5 2.4B and 32B variants provide family-level size alternatives, but the supplied research does not establish that either is universally better for a particular workload.
Bottom line
EXAONE-3.5-7.8B-Instruct is a practical open-weight choice for bilingual English-Korean text applications that benefit from a 32K context window and local inference. Its strongest differentiators are language focus, deployment control, and compatibility with several open-source serving and quantization tools. Its trade-offs are equally important: it is text-only, lacks a verified provider API and per-token price, does not document native tools or strict JSON output, and requires license and infrastructure decisions before production or commercial use.

