EXAONE 3.5

EXAONE-3.5-32B-Instruct

by LG AI Research · Available open-weight model; older EXAONE generation

EXAONE-3.5-32B-Instruct is LG AI Research’s largest EXAONE 3.5 instruction model, with approximately 30.95 billion parameters and a 32,768-token context window. It generates text in English and Korean and supports coding, mathematics, long-context processing, and local deployment through frameworks such as Transformers, vLLM, SGLang, llama.cpp, and Ollama. It is text-only, has no verified hosted token pricing or native web search, and uses a license that requires careful commercial-use review.

Text Reasoning Coding
EXAONE-3.5-32B-Instruct is the largest model in LG AI Research’s EXAONE 3.5 instruction-tuned family. Released as an open-weight model in December 2024, it focuses on English and Korean text generation, instruction following, long-context tasks, coding, mathematics, and general-purpose language work. Its downloadable weights make it a potential alternative to hosted AI services for organizations that can provide the required hardware and manage deployment themselves.
Outputs

What EXAONE-3.5-32B-Instruct can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

7/10 Reasoning
7/10 Coding
4/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family EXAONE 3.5
Model type General Purpose
Context window 33K tokens
Release date 2024-12-09
Status Available open-weight model; older EXAONE generation
Knowledge cutoff notes

The official model card warns that the model does not reflect the latest information but does not publish a precise training-data or knowledge-cutoff date.

Model notes

The canonical Hugging Face identifier is LGAI-EXAONE/EXAONE-3.5-32B-Instruct. The model has approximately 30.95B parameters excluding embeddings and supports a 32,768-token context window. Official documentation lists Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, and Ollama deployment paths. AWQ and GGUF quantized versions are available. The official license is EXAONE AI Model License Agreement 1.1 - NC, so commercial and other usage conditions must be reviewed carefully. No official hosted token pricing or exact knowledge-cutoff date was identified. Editorial scores are comparative estimates, not vendor specifications.

Model guide

EXAONE-3.5-32B-Instruct: LG AI Research’s Bilingual Open-Weight Model

EXAONE-3.5-32B-Instruct is a 32-billion-parameter bilingual English-Korean instruction-tuned language model from LG AI Research. It is designed for local or self-hosted text generation, with a 32,768-token context window, coding and mathematics capability, and deployment support across several inference frameworks.

What is EXAONE-3.5-32B-Instruct?

EXAONE-3.5-32B-Instruct is an instruction-tuned causal language model developed by LG AI Research. In practical terms, it is a text-generation model trained to respond to written instructions rather than merely continue text. It can be used for conversations, drafting, summarization, translation, coding assistance, mathematics, and other language tasks.

The model is part of the EXAONE 3.5 family, which also includes smaller 2.4B and 7.8B variants. The 32B version is the family’s largest configuration and is intended for users who prioritize output quality and capability over lightweight hardware requirements. Its canonical model identifier is LGAI-EXAONE/EXAONE-3.5-32B-Instruct, and the weights are distributed through the LGAI-EXAONE organization on Hugging Face.

LG AI Research released the model on December 9, 2024. It is an older generation within the provider’s broader EXAONE catalog, but it remains relevant for users specifically seeking an open-weight bilingual English-Korean model rather than a consumer chatbot or a conventional hosted API.

Key specifications and context limit

SpecificationVerified detail
Model typeInstruction-tuned causal language model
ParametersApproximately 30.95 billion, excluding embeddings
LanguagesEnglish and Korean
Context length32,768 tokens
Transformer layers64
Attention designGrouped-query attention with 40 query heads and 8 key-value heads
Vocabulary102,400 tokens
Input and outputText input and text output

The 32,768-token context window is the amount of text the model can consider in one request, including the prompt and the generated response. This is useful for long documents, extended conversations, source-code files, and bilingual material. The published information does not specify a separate maximum output-token limit, so users should not assume that the entire context window is available for generated output alone.

Language and core capabilities

The clearest specialization of EXAONE-3.5-32B-Instruct is bilingual English-Korean text generation. It is designed for instruction following in both languages and can support tasks such as drafting, rewriting, summarizing, question answering, translation-oriented workflows, and document analysis based on text supplied by the user.

Its larger parameter count gives it a higher-capacity position within the EXAONE 3.5 family, although the supplied research does not establish a universal quality ranking against every current commercial or open-weight model. LG AI Research reports strong performance in instruction following, real-world usability, long-context understanding, and Korean-language evaluations. Those are provider or research claims; actual results will depend on prompting, hardware, quantization, language, and the particular evaluation task.

The model is also intended for coding and mathematics. It can help explain code, generate functions, inspect errors, work through calculations, and produce structured answers in response to suitable prompts. These capabilities do not mean that it includes a built-in code execution environment or that generated code is automatically correct. Developers should test code and validate mathematical outputs, especially in production or high-impact settings.

Reasoning, coding, and tool support

EXAONE-3.5-32B-Instruct is an instruction model rather than a separately documented reasoning-mode model. The available research does not verify a dedicated reasoning switch, hidden chain-of-thought feature, provider-managed reasoning budget, or reasoning-token limit. It can perform multi-step reasoning in ordinary text generation, but the quality and reliability of that reasoning should be evaluated for the intended workload.

Coding is one of its supported use cases, alongside mathematics and general text generation. However, tool use, function calling, structured-output guarantees, JSON mode, streaming, batch processing, caching, and fine-tuning are not verified in the supplied model-specific research. The model should therefore be treated as a downloadable text-generation model, not as a complete hosted agent platform.

There is no provider-hosted web-search system attached to the model. It does not automatically retrieve current information, and its model card warns that it does not reflect the latest information. Applications that need current facts can add their own retrieval system, but that would be an external application feature rather than a native EXAONE-3.5-32B-Instruct capability.

Supported modalities and practical limits

This model accepts text and produces text. It does not natively accept images, audio, or video, and it does not directly generate images, audio, or video. This distinction matters because LG AI Research’s wider EXAONE lineup includes newer multimodal research and model releases; those capabilities should not be attributed to this specific 3.5 32B instruction model.

For text-only applications, the 32,768-token context is substantial enough for many long-document and multi-turn use cases. Nevertheless, a long context does not guarantee accurate recall of every detail. Important passages may still need to be highlighted, summarized, or retrieved in stages. Since no exact maximum output-token value is published in the supplied sources, deployment settings should be tested rather than based on an assumed output ceiling.

Deployment options and hardware trade-offs

EXAONE-3.5-32B-Instruct is aimed at local, private, or self-hosted inference. Official documentation provides deployment paths for Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, and Ollama. AWQ and GGUF quantized versions are also available. Quantization reduces the memory needed to load a model by using lower-precision representations, although it can introduce quality or compatibility trade-offs.

A 32B model is considerably more demanding than the 2.4B and 7.8B members of the same family. The larger model may offer better capability for difficult language, coding, and long-context tasks, but it generally requires more memory, compute, and operational work. Smaller variants may be more appropriate for laptops, constrained servers, faster interactive applications, or deployments where infrastructure cost matters more than maximum quality.

The editorial assessment supplied with this profile rates the model’s reasoning and coding capability at 7 out of 10, speed at 4 out of 10, and cost at 8 out of 10. These are comparative editorial estimates, not measurements published by LG AI Research and not guarantees of runtime performance. Actual speed and cost depend on hardware, quantization, batching, context length, and inference software.

Pricing and licensing

No official hosted token pricing or recurring subscription price is identified for EXAONE-3.5-32B-Instruct. Because the model is distributed as open weights, the main financial considerations are hardware, hosting, electricity, engineering, storage, and operational support rather than a documented provider API price. Partner-mediated commercial access may exist elsewhere in the EXAONE ecosystem, but no model-specific hosted price is verified here.

The model uses the EXAONE AI Model License Agreement 1.1 - NC. The “NC” designation means users should not treat it as an unrestricted commercial open-source license. Before using the weights in a commercial product, review the official license and determine whether the planned use complies with its non-commercial and other conditions. Quantized copies and third-party deployment packages should also be checked for their own distribution terms.

Main strengths and limitations

Strengths

  • English-Korean bilingual focus, including support for Korean-language workloads.
  • Large 32B configuration for users seeking more capability than smaller EXAONE 3.5 variants.
  • 32,768-token context window for long prompts and document-oriented tasks.
  • Downloadable weights that support local and self-hosted deployment.
  • Documented compatibility with several widely used inference frameworks.
  • Quantized AWQ and GGUF options for more constrained deployments.
  • Useful target tasks including instruction following, coding, mathematics, and general text generation.

Limitations

  • Text-only input and output; there is no native image, audio, or video support.
  • No verified provider-managed web search, current-information retrieval, or code execution.
  • No verified model-specific support for function calling, JSON mode, streaming, batch APIs, caching, or fine-tuning.
  • No published exact maximum output-token limit in the supplied documentation.
  • The 32B size can make inference slower and more hardware-intensive than smaller models.
  • The license requires careful review and may not suit unrestricted commercial deployment.
  • Responses can contain factual errors, bias, harmful content, or outdated information.

When to choose EXAONE-3.5-32B-Instruct

Choose this model when you need a downloadable English-Korean language model for local or self-hosted inference and can support its infrastructure requirements. It is a reasonable candidate for Korean-language applications, bilingual document workflows, private experimentation, coding assistance, long-context text processing, and organizations that want more control over where inference takes place.

It is less suitable when you need a polished consumer chatbot, a first-party managed API with clearly documented token pricing, built-in web search, native multimodal input, guaranteed structured outputs, or integrated tools and agents. A smaller EXAONE 3.5 variant may be preferable when speed and hardware efficiency are more important than the capacity of the 32B model. A newer multimodal model may be more appropriate for image or document-image understanding, while a managed commercial service may reduce the operational burden for teams that do not want to run inference themselves.

For high-impact use, treat the model as an assistive component rather than an authority. Add application-level retrieval when current information matters, validate generated code and calculations, apply safety filtering, and include human review for decisions that affect people, finances, health, security, or legal rights.


Answers to Frequently Asked Questions

What are the main limitations of EXAONE-3.5-32B-Instruct?
The model accepts and produces text only; it does not natively support images, audio, or video. It has no verified built-in web search, current-information retrieval, code execution, function calling, JSON mode, streaming, batch APIs, caching, or fine-tuning support. Its 32B size also requires substantially more memory and compute than smaller EXAONE 3.5 variants, and its outputs may contain errors or outdated information.
Is EXAONE-3.5-32B-Instruct free for commercial use?
Commercial use should not be assumed to be unrestricted. The model is provided under the EXAONE AI Model License Agreement 1.1 - NC, so users should review the official license and confirm that their intended use complies with its non-commercial and other conditions.
How can EXAONE-3.5-32B-Instruct be deployed?
EXAONE-3.5-32B-Instruct is distributed as open weights through Hugging Face under the identifier LGAI-EXAONE/EXAONE-3.5-32B-Instruct. It can be deployed locally or in self-hosted environments using Transformers, vLLM, SGLang, TensorRT-LLM, llama.cpp, or Ollama. AWQ and GGUF quantized versions are also available.
What is EXAONE-3.5-32B-Instruct?
EXAONE-3.5-32B-Instruct is an instruction-tuned causal language model developed by LG AI Research. It is designed for English and Korean text tasks such as conversation, summarization, translation-oriented workflows, coding assistance, mathematics, document analysis, and general text generation.
What are the context length and language capabilities of EXAONE-3.5-32B-Instruct?
The model supports English and Korean and has a 32,768-token context window, allowing it to process long prompts, documents, source-code files, and extended conversations. The context limit includes both the input prompt and generated response.


Sources 5
Provider

About LG AI Research