1. What the current EXAONE API is
EXAONE is not currently exposed through a broadly documented LG AI Research API with its own public endpoint, pricing page, SDK package, or developer console. The most direct managed option is an EXAONE 4.0 API delivered through FriendliAI.
FriendliAI provides an OpenAI-compatible Model API. That means developers can use familiar chat-completion request patterns while changing the API key, base URL, and model name. The current serverless base URL is https://api.friendli.ai/serverless/v1, and the documented EXAONE 4.0 32B model identifier is LGAI-EXAONE/EXAONE-4.0-32B.
The other route is self-hosting. Official EXAONE model releases are available through the LG-AI-EXAONE organization on Hugging Face and can be run with tools such as Transformers, vLLM, or SGLang. With self-hosting, you control the serving environment, but you also take responsibility for GPUs, authentication, scaling, monitoring, latency, quotas, and application reliability.
Who should use this API?
The FriendliAI integration is intended for developers who want hosted access to EXAONE 4.0 without provisioning GPU infrastructure. It may be especially relevant for Korean-language and multilingual applications, professional or document-oriented workflows, agent experiments, and teams evaluating EXAONE without operating a model-serving stack.
Self-hosting is more appropriate when data locality, infrastructure control, customization, or predictable deployment ownership matters more than managed operations. Neither route should be treated as a conventional LG-hosted consumer chatbot API.
2. Getting access and obtaining credentials
For the managed route, create or use a FriendliAI account and obtain a Friendli API token through the FriendliAI platform. Store the token as a secret, preferably in an environment variable named FRIENDLI_TOKEN. Do not place it in browser code, public repositories, or client-side mobile applications.
The API request uses standard bearer authentication:
Authorization: Bearer $FRIENDLI_TOKENLG AI Research does not currently publish a separate first-party key system for the commercial EXAONE endpoint. The account, credential, billing, quota, and access relationship for the hosted API are therefore with FriendliAI.
Self-hosted access
A self-hosted deployment does not use FriendliAI billing or FriendliAI credentials unless you separately place it behind your own gateway or serving platform. You must implement the authentication and access controls needed by your application. The model license and deployment terms should also be reviewed before commercial use.
3. Choosing the appropriate API or model route
| Route | Best for | Main responsibility |
|---|---|---|
| FriendliAI Serverless Endpoint | Fastest hosted integration and standard API applications | FriendliAI account terms, quotas, pricing, and service conditions |
| FriendliAI Dedicated Endpoint | Workloads needing allocated deployment capacity | Capacity planning, deployment cost, and operational configuration |
| Self-hosted EXAONE | Data locality, infrastructure control, or custom serving | GPU capacity, serving software, scaling, security, and reliability |
The examples in this guide use LGAI-EXAONE/EXAONE-4.0-32B through FriendliAI’s Serverless Endpoint. EXAONE 4.0 supports reasoning and non-reasoning modes and covers Korean, English, and Spanish. Confirm the exact behavior and supported parameters for the deployed model before depending on them in production.
4. Making your first request
The simplest request is a chat completion. It sends a list of messages and receives an assistant response. First, install the current OpenAI Python package and set the Friendli token in your environment:
pip install openai
export FRIENDLI_TOKEN="your-token"
Then configure the OpenAI client with FriendliAI’s compatible base URL:
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["FRIENDLI_TOKEN"],
base_url="https://api.friendli.ai/serverless/v1",
)
response = client.chat.completions.create(
model="LGAI-EXAONE/EXAONE-4.0-32B",
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain what an API is in one sentence."},
],
temperature=0.2,
max_tokens=128,
)
print(response.choices[0].message.content)
The same operation can be made with raw HTTP:
curl --fail-with-body --silent --show-error
https://api.friendli.ai/serverless/v1/chat/completions
-H "Authorization: Bearer ${FRIENDLI_TOKEN}"
-H "Content-Type: application/json"
-d '{
"model": "LGAI-EXAONE/EXAONE-4.0-32B",
"messages": [
{"role": "user", "content": "Explain what an API is in one sentence."}
],
"temperature": 0.2,
"max_tokens": 128
}'5. Understanding the response
A successful chat-completion response follows the familiar OpenAI-compatible structure. The generated text is normally found at response.choices[0].message.content in the Python SDK. The response can also contain metadata such as an identifier, model information, finish reason, and usage information when provided by the service.
Applications should not assume that every optional field is present. Check for missing content, unexpected finish reasons, API errors, and malformed tool arguments. For user-facing applications, add request timeouts and a clear error path instead of treating every non-successful request as an empty answer.
6. How pricing works
LG AI Research does not publish a standalone EXAONE API price list. The FriendliAI integration is a hosted inference service, so pricing and account terms are determined by FriendliAI. No EXAONE-specific public price was verified in the supplied documentation.
FriendliAI also offers Dedicated Endpoints, which are based on allocated deployment capacity rather than only per-request serverless usage. The correct choice depends on workload shape, availability requirements, and the capacity terms offered to your account.
Before committing to production, confirm the current pricing model, quotas, regional availability, rate limits, billing unit, and any minimum commitments directly with FriendliAI. Do not estimate project cost from a generic FriendliAI price description because the EXAONE-specific commercial terms may differ.
7. Important capabilities and unavailable features
Text generation and languages
The public EXAONE 4.0 API path is text-oriented. The model supports reasoning and non-reasoning modes and is documented for Korean, English, and Spanish use. The available behavior depends on the deployed model and endpoint configuration.
Streaming responses
Streaming sends generated text in increments rather than waiting for the complete answer. It is useful for chat interfaces because the application can display partial output sooner.
stream = client.chat.completions.create(
model="LGAI-EXAONE/EXAONE-4.0-32B",
messages=[
{"role": "user", "content": "List three practical uses of structured data."}
],
stream=True,
max_tokens=256,
)
for chunk in stream:
text = chunk.choices[0].delta.content or ""
print(text, end="", flush=True)
print()
Streaming requires the application to handle partial chunks, connection interruptions, and the possibility that a final response is not received. Store the completed result only after the stream has ended successfully.
Tool and function calling
EXAONE 4.0 supports agentic tool use and function calling. A tool definition tells the model what an application-controlled function can do; the model does not execute that function itself. Your server must validate the requested arguments, perform the action, and optionally send the result back in a follow-up conversation turn.
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}]
tool_response = client.chat.completions.create(
model="LGAI-EXAONE/EXAONE-4.0-32B",
messages=[{"role": "user", "content": "What is the weather in Seoul?"}],
tools=tools,
tool_choice="auto",
)
message = tool_response.choices[0].message
if message.tool_calls:
call = message.tool_calls[0]
print(call.function.name)
print(call.function.arguments)
else:
print(message.content)
Never execute arbitrary tool arguments without validation. Restrict tools to the actions the user is authorized to request, and apply normal application security controls.
Structured outputs and JSON
FriendliAI advertises JSON mode and schema-guided outputs across its Model API platform. These features can help when an application needs machine-readable data instead of free-form text. However, support should be confirmed for the selected EXAONE endpoint and model combination before production use. A platform-level feature description is not proof that every deployed model accepts every structured-output option.
Files, images, and other hosted features
The current public EXAONE 4.0 API documentation does not provide reliable evidence for hosted file uploads, image input, hosted web search, persistent assistants, or provider-managed fine-tuning through a first-party LG AI Research API. Treat the endpoint as text-focused unless FriendliAI confirms a specific capability for your account and model.
8. SDKs, playground, and developer tools
Because the endpoint is OpenAI-compatible, developers can use the current official OpenAI Python or JavaScript SDKs by supplying FriendliAI’s custom base URL. Raw HTTP is also suitable when you want direct control over requests. Compatibility does not mean that every OpenAI feature is automatically available; test each parameter against the FriendliAI deployment.
FriendliAI provides a web console and a Serverless Endpoint overview for the EXAONE 4.0 deployment. This is the relevant playground-style experience for the hosted API. It is not evidence of a separate LG AI Research developer console or a native LG SDK.
9. Advanced integration and production considerations
Configuration, timeouts, and retries
Keep the base URL, model identifier, and token outside application code so they can be changed without a redeployment. Configure explicit request timeouts. Retry only errors that are safe to retry, use backoff, and avoid sending duplicate tool actions when a request may have completed but the connection was interrupted.
Quotas and latency
No EXAONE-specific public rate-limit table, guaranteed latency target, or verified latency benchmark was found. FriendliAI describes its platform as optimized for low-latency inference and offers serverless and dedicated deployment options, but actual performance depends on the account, endpoint, request size, traffic, and deployment configuration. Confirm quotas and service-level commitments before promising response times to users.
Privacy, retention, and data handling
No current EXAONE-specific retention period or comprehensive policy was verified for prompts and outputs sent through the hosted FriendliAI endpoint. Review FriendliAI’s privacy, data-processing, and service terms before sending confidential or regulated information. Self-hosting gives the operator more control over inference data, but application logs, monitoring systems, backups, and infrastructure providers still need to be governed.
Model and service changes
The hosted API is partner-mediated. A future direct LG service, a new FriendliAI deployment, or a different EXAONE model may use different identifiers, URLs, authentication rules, quotas, or supported parameters. Pin the model configuration where possible and test changes before switching production traffic.
10. When this API is a good or poor choice
A good choice when
- You want managed access to EXAONE 4.0 without running GPUs.
- Your application benefits from Korean, English, or Spanish language support.
- You need OpenAI-compatible chat completions, streaming, or tool calling.
- You are evaluating EXAONE for professional, document-oriented, or agentic workflows and can accept partner-managed access.
- You need self-hosting as an alternative and are prepared to operate the infrastructure yourself.
A poor choice when
- You require a clearly documented first-party LG API with its own stable pricing, SDKs, quotas, and support process.
- Your application depends on hosted file uploads, web search, persistent assistants, or image input through the public endpoint.
- You need a publicly verified EXAONE-specific rate-limit table, retention period, or latency SLA before evaluation.
- You want a mature consumer application ecosystem rather than a model API and deployment platform.
For a first experiment, FriendliAI’s OpenAI-compatible Serverless Endpoint is the shortest path: obtain a token, configure the base URL, select the EXAONE 4.0 model, and make a chat-completion request. For sensitive or infrastructure-controlled workloads, compare that managed route with a properly secured self-hosted deployment.
