What is HCX-003?
HCX-003 is a text-generation model provided by NAVER through NAVER Cloud's CLOVA Studio. It belongs to the HyperCLOVA X model family and is intended for applications that send text prompts and receive text responses. Typical tasks include summarization, classification, information extraction, rewriting, question answering, conversational responses, and Korean-language content generation.
Unlike a multimodal model, HCX-003 does not process images, audio, or video. Its role is narrower and more conventional: it provides language generation for applications that do not require visual understanding, native tool orchestration, or extended reasoning features.
NAVER lists HCX-003 alongside newer or differently positioned models such as HCX-007, HCX-005, HCX-DASH-002, and HCX-DASH-001 in the CLOVA Studio catalog. HCX-003 remains available in both Classic and VPC environments according to the supplied documentation.
Context window and technical limits
The model has an 8,192-token combined context limit. This total includes the input and the requested output, so a long prompt leaves less room for the response. NAVER documents a maximum input length of 7,600 tokens and a maximum requested output of 4,096 tokens.
| Specification | HCX-003 |
|---|---|
| Model family | HyperCLOVA X |
| Input | Text only |
| Output | Text only |
| Combined context limit | 8,192 tokens |
| Maximum documented input | 7,600 tokens |
| Maximum requested output | 4,096 tokens |
| Release date | October 17, 2023 |
| Availability | Current in the supplied CLOVA Studio catalog |
These limits make HCX-003 suitable for ordinary documents, messages, support interactions, and structured text-processing jobs. It is less suitable for very long documents or prompts that require a large amount of source material and a long response in the same request. The supplied research does not provide a model-specific knowledge-cutoff date.
Modalities and API capabilities
HCX-003 accepts text and produces text. It does not support image, audio, or video input, and it does not generate non-text media. Applications should therefore convert non-text material into text elsewhere before sending it to the model, if that workflow is appropriate.
NAVER documents access through CLOVA Studio chat APIs, including the standard Chat Completions API and an OpenAI-compatible access path. Streaming is supported, allowing an application to receive generated text progressively rather than waiting for the entire response. This can improve the perceived responsiveness of chat and writing interfaces.
The model supports tuning. In this context, tuning means creating a customized version using task-specific examples or datasets. This can be useful when a general prompt is not sufficient to produce consistent classifications, formats, terminology, or domain-specific writing. Tuning does not turn HCX-003 into a multimodal or tool-using model; it remains a text-focused model with the same general capability boundaries described in the catalog.
Important unsupported features
The supplied NAVER model comparison information marks HCX-003 as unsupported for function calling, structured outputs, and reasoning mode.
- Function calling: The model does not provide native function-calling support for selecting and invoking external tools through a defined interface.
- Structured outputs: It is not documented as supporting schema-constrained output. Prompts may request JSON or another format, but that should not be treated as the same as provider-enforced structured output.
- Reasoning mode: HCX-003 does not offer the dedicated reasoning feature identified for some other models in the lineup.
- Multimodal understanding: It cannot directly analyze images, audio, or video.
These limitations matter when designing software. A developer can still build application logic around ordinary text responses, but tool selection, validation, retries, and output parsing would need to be handled outside the model. Workflows that depend on reliable schema compliance or native agent actions may be better served by another model or by a separate orchestration layer.
Performance, cost, and lineup positioning
The supplied editorial data rates HCX-003's speed and cost efficiency at 7 out of 10, with reasoning and coding both rated at 5 out of 10. These are comparative editorial estimates, not NAVER-published benchmarks. They suggest a model positioned for practical, focused text workloads rather than maximum reasoning depth or specialized software-development performance.
HCX-003's main trade-off is capability breadth versus operational simplicity. Its text-only design and moderate context window can be sufficient for applications that do not need images or complex tool use. Tuning and batch support also make it relevant for repeatable domain-specific processing. However, applications requiring long-context analysis, multimodal input, advanced reasoning, or native action-taking should evaluate other models in NAVER's lineup.
The supplied public documentation does not expose a stable model-specific input or output price. Pricing is provided through the NAVER Cloud service portal, so no verified numeric price is stated here. Actual cost and service limits may depend on the account, environment, deployment configuration, and applicable NAVER Cloud policies.
Best use cases for HCX-003
HCX-003 is a practical fit when the primary requirement is dependable text transformation or generation rather than broad media understanding. Suitable examples include:
- Korean-language article, product, marketing, or support-content generation
- Summarizing messages, reports, or other text within the context limit
- Extracting named fields, facts, or categories from text through prompt-based workflows
- Text classification and routing
- Rewriting content into a specified tone or format
- Domain-specific generation using tuned models
- Batch data generation and dataset expansion
- Conversational applications that need streamed text responses
Its Korean-language focus is especially relevant for teams building around Korean content and terminology. The research describes HCX-003 as suitable for Korean, English, and Japanese use cases, but the strongest stated positioning is Korean-focused text generation and transformation.
When should you choose another model?
Choose another option when the application depends on capabilities that HCX-003 does not provide. A multimodal model is more appropriate for image, audio, or video understanding. A model with a larger context window is preferable for long documents or large collections of source material. A model with native function calling or structured outputs is a better fit for applications that must reliably invoke tools or return schema-validated data.
Within NAVER's current lineup, the supplied research identifies HCX-007 as offering a reasoning mode, while HCX-005 and HCX-DASH-002 provide larger context windows than HCX-003. Those models may be more appropriate when deeper reasoning or longer inputs are more important than using the smaller model's focused text-processing profile. The exact cost and performance differences are not established by the supplied sources, so they should be checked in the current CLOVA Studio documentation or service portal before deployment.
Overall assessment
HCX-003 is best understood as a compact, conventional language model for text-only workloads. Its useful strengths are its availability through CLOVA Studio, Korean-focused positioning, tuning support, streaming chat access, and compatibility with batch-oriented text tools. Its 8,192-token combined context limit is adequate for many routine tasks, but it constrains large-document workflows.
The model is a sensible choice when an application needs generated or transformed text and does not require native tools, schema enforcement, multimodal input, or advanced reasoning. Its limitations are clear enough to support straightforward architecture decisions: use HCX-003 for focused text operations, and consider a different model when the workload depends on broader modalities, longer context, or more sophisticated agent behavior.

