What is EXAONE-3.5-2.4B-Instruct?
EXAONE-3.5-2.4B-Instruct is an instruction-tuned causal language model developed by LG AI Research. Instruction tuning means the model has been trained to respond to natural-language requests, making it more suitable for tasks such as answering questions, rewriting text, summarizing documents, classifying content, and following structured prompts than a base text-completion model.
The model belongs to the EXAONE 3.5 family and is the smallest of its reported 2.4B, 7.8B, and 32B variants. LG AI Research positions this version for small or resource-constrained devices. In practical terms, it trades some capability headroom for lower hardware requirements and potentially faster local inference than larger models in the same family.
The model was released on December 9, 2024. It is available as downloadable open weights through the official model distribution channels, including Hugging Face. This is different from using a consumer chatbot or a conventional hosted API: the user or organization is responsible for selecting an inference environment, configuring the model, and complying with its license.
Capabilities and technical profile
EXAONE-3.5-2.4B-Instruct is designed for bilingual English and Korean text generation. Its documented maximum context length is 32,768 tokens. A context window is the amount of input and conversation history the model can consider in one request; it is not a guarantee that every long document will be understood equally well from beginning to end.
| Specification | Verified detail |
|---|---|
| Provider | LG AI Research |
| Model family | EXAONE 3.5 |
| Release date | December 9, 2024 |
| Model size | 2.14 billion parameters excluding embeddings; marketed family variant is 2.4B |
| Languages | English and Korean |
| Context length | 32,768 tokens |
| Output type | Text only |
| Access | Downloadable weights for local or research deployment |
| Hosted pricing | No official hosted API pricing verified |
The model card reports 30 layers, grouped-query attention with 32 query heads and 8 key-value heads, a vocabulary size of 102,400, and tied word embeddings. Grouped-query attention is an architectural approach that can reduce memory use during generation compared with storing separate key-value heads for every query head. These details are useful to deployment engineers, but they do not by themselves guarantee a particular speed on a specific computer.
Modalities and platform features
This model is text-only. It accepts text and produces text; it does not natively process images, audio, or video, and it does not generate non-text media. It should therefore be evaluated as a compact language model rather than as a vision-language or general multimodal system.
No authoritative model-specific documentation supplied for this page confirms native web search, function calling, tool use, enforced JSON mode, structured-output guarantees, batch processing, prompt caching, or a model-specific fine-tuning workflow. A developer may be able to build surrounding application logic that performs some of these tasks, but that would not make them native, verified capabilities of EXAONE-3.5-2.4B-Instruct.
There is also no documented maximum output-token limit in the supplied sources. The 32,768-token context length should not be interpreted as a separate promise that the model can generate 32,768 output tokens. Actual generation limits depend on the inference framework, memory available, and the amount of input included in the request.
Performance, reasoning, and coding
LG AI Research reports the following benchmark results for the model: 7.81 on MT-Bench, 33.0 on LiveBench, 48.2 on Arena-Hard, 37.1 on AlpacaEval, 73.6 on IFEval, 7.24 on KoMT-Bench, and 8.51 on LogicKor. These are provider-reported results from the published comparison material. Benchmark scores depend on the test design, prompts, evaluators, and comparison models, so they should not be treated as universal predictions of performance in a particular application.
The model’s instruction tuning and bilingual coverage make it suitable for ordinary reasoning tasks such as extracting information, comparing supplied text, following multi-step instructions, and producing concise explanations. However, it has substantially fewer parameters than the larger EXAONE 3.5 variants. The supplied evaluation classifies its reasoning capability as moderate rather than frontier-level, and its coding capability as limited-to-moderate. It may be useful for code explanation, simple generation, transformation, and lightweight programming assistance, but demanding software engineering or complex reasoning may benefit from a larger model.
The editorial assessment supplied with the model gives it a reasoning score of 5 out of 10, coding score of 4 out of 10, speed score of 8 out of 10, and cost score of 9 out of 10. These scores are comparative editorial estimates, not specifications published by LG AI Research. The high speed and cost assessments reflect the model’s small size and downloadable local-deployment format, but actual results vary with quantization, hardware, batch size, and serving software.
Deployment options and license
Official guidance supports loading the model with Transformers and provides deployment guidance for vLLM, SGLang, TensorRT-LLM, llama.cpp, and Ollama. Quantized AWQ and GGUF variants are available through the EXAONE 3.5 model collection. Quantization reduces the numerical precision used to store or run a model and can lower memory requirements, although it may introduce a quality or compatibility trade-off.
The official license is the EXAONE AI Model License Agreement 1.1 - NC. According to the supplied research, it permits research use but prohibits commercial use unless a separate commercial license is obtained. This restriction is central to deployment decisions: downloading the weights does not automatically grant permission to embed the model in a paid product, commercial service, or business workflow.
There is no verified official hosted API price for this exact model. The practical cost is therefore tied to the hardware, cloud instance, inference framework, and operational setup used to run the weights. Organizations should also review the license and any applicable terms before deploying a quantized or modified version.
Main strengths and limitations
Strengths
- Compact deployment profile: Its smaller size makes it more suitable than larger language models for local inference and resource-constrained environments.
- English-Korean coverage: It is specifically designed for bilingual text generation, which can be valuable for translation-adjacent workflows, bilingual assistants, rewriting, and Korean-language applications.
- Long context for its size: A 32,768-token context window supports substantial prompts and documents, subject to the quality and memory limits of the chosen runtime.
- Multiple deployment paths: Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, and Ollama are identified in the official deployment guidance, with AWQ and GGUF variants available.
- Local control: Downloadable weights can support applications that need local or on-premise processing rather than sending every prompt to a third-party hosted endpoint.
Limitations
- Research-focused licensing: Commercial use requires separate authorization under the supplied interpretation of the EXAONE AI Model License Agreement 1.1 - NC.
- Text-only operation: The model cannot natively understand images, audio, or video.
- No verified hosted API: No official hosted API price or broad first-party API offering was verified for this exact model.
- No documented native tools: Web search, function calling, JSON mode, structured-output enforcement, batch processing, caching, and a model-specific fine-tuning workflow are not confirmed.
- Smaller capability ceiling: The 2.4B-class model is better suited to lightweight tasks than to demanding reasoning, complex coding, or knowledge-intensive workloads.
- Outdated information risk: LG AI Research warns that responses may be inaccurate, biased, harmful, or outdated. No current web-search capability is documented.
Best use cases
EXAONE-3.5-2.4B-Instruct is a reasonable candidate when the application needs a relatively small bilingual text model and can run inference locally. Suitable examples include:
- English-Korean question answering over supplied information
- Summarization and rewriting of internal or user-provided text
- Text classification and lightweight content-routing workflows
- Drafting short responses, descriptions, or bilingual communications
- Local experiments on resource-constrained devices
- Research and educational projects involving open-weight language models
- Private document-processing prototypes where local execution is preferred
For production use, developers should test the model on representative Korean and English prompts, measure latency and memory usage on the intended hardware, and add application-level validation. The model should not be relied on as an authoritative source of current facts.
When to choose this model
Choose EXAONE-3.5-2.4B-Instruct when compact size, local deployment, bilingual English-Korean generation, and research access matter more than maximum reasoning or coding performance. It is especially attractive for experiments where a hosted API is unnecessary and where the license permits the intended non-commercial use.
A larger model is likely more appropriate for complex multi-step reasoning, difficult programming tasks, high-stakes analysis, or workloads where response quality is more important than memory and inference efficiency. Within the EXAONE 3.5 family, the 7.8B or 32B variants may offer greater capability at the cost of higher resource requirements, although the supplied research does not provide a direct, complete benchmark comparison for every task.
A multimodal model should be selected instead when image, audio, or video understanding is required. A hosted commercial model may also be a better fit when an organization needs a documented API, managed scaling, current-information tools, enforced structured output, or clearer commercial service terms. EXAONE-3.5-2.4B-Instruct’s advantage is narrower and more practical: it is a small, bilingual, downloadable text model for local and research-oriented use.

