What is HCX-005?
HCX-005 is a multimodal language model provided by NAVER Cloud as part of the HyperCLOVA X model family. The model was introduced on April 17, 2025, and is available through CLOVA Studio and related developer interfaces. In practical terms, “multimodal” means that HCX-005 can work with more than text: it accepts images as well as written prompts, then returns a text response.
This makes HCX-005 different from a conventional text-only language model, but its output scope remains deliberately narrow. The supplied specifications identify text as its output type; it does not natively generate images, audio, video, music, embeddings, or speech. Its visual capability is therefore primarily for understanding and analysis rather than media creation.
NAVER Cloud positions HCX-005 for Korean-language business applications, image understanding, document and visual analysis, instruction following, and API-based assistants. Those are provider-oriented use cases. The assessment that it is a good fit for a particular application should still depend on the quality of the application’s prompts, images, retrieval system, and tool integrations.
Where HCX-005 fits in the HyperCLOVA X family
HCX-005 belongs to NAVER Cloud’s current HyperCLOVA X lineup and is specifically identified in the supplied documentation as a vision-capable model. Its role is not to cover every possible AI workload, but to provide language generation with image input and developer-oriented interaction features.
The model documentation identifies HCX-007 as a separate thinking model. That distinction matters when selecting a model: HCX-005 is aimed at multimodal understanding and practical assistant workflows, while a dedicated thinking model may be more appropriate for applications whose central requirement is deeper reasoning. The available research does not provide a direct benchmark comparison between the two models, so this is a positioning distinction rather than a claim that one model is universally more accurate.
HCX-005 is also more specialized than a general-purpose media model. It can interpret visual material, but the supplied specifications do not support using it for image generation, video creation, voice synthesis, or music production.
Core capabilities and technical limits
| Specification | HCX-005 |
|---|---|
| Provider | NAVER Cloud |
| Model family | HyperCLOVA X |
| Release date | April 17, 2025 |
| Input | Text and images |
| Output | Text |
| Context length | 128,000 tokens |
| Maximum output | 4,096 tokens |
| Image limit | Up to five images per request, with at most one image per turn |
| Image formats | BMP, PNG, JPG, JPEG, and WEBP |
| Image delivery | Public URL or data URI |
| Streaming | Supported |
| Function calling | Supported |
| Fine-tuning | PEFT tuning supported |
| Structured Outputs | Not supported according to the supplied documentation |
The 128,000-token context window is useful for long prompts, document collections, or conversations that need substantial background material. A context window is the amount of input and generated text the model can consider within a request; it is not the same as a guarantee that every long document will be understood equally well. The maximum generated response is 4,096 tokens, so applications that need very long reports may need to request multiple sections or manage continuation carefully.
HCX-005 can receive up to five images in one request, but the documented restriction of at most one image per turn is important when designing a conversation. Images may be supplied through a public URL or a data URI, and the documented formats include BMP, PNG, JPG, JPEG, and WEBP. These details make it suitable for workflows such as analyzing a small group of page images, product photographs, charts, or scanned business material, subject to the model’s actual visual interpretation quality.
Tools, API access, and output formats
HCX-005 supports Chat Completions v3, OpenAI-compatible access, streaming responses, and function calling. Function calling allows the model to request an application-defined operation, such as looking up an order, querying an internal database, or creating a reservation. The model does not perform those external actions by itself; the surrounding application must validate the request, execute the function, and return the result.
Streaming is useful when an application wants to display generated text progressively rather than waiting for the complete response. This can improve perceived responsiveness, particularly for assistants and interactive document tools.
Structured Outputs are documented as unsupported for HCX-005. A separate legacy JSON-mode capability is not verified in the supplied research, so developers should not assume that requesting JSON guarantees schema-valid output. If an application needs machine-readable responses, it should use careful prompting, validation, error handling, and potentially function calling rather than treating ordinary text generation as a formal schema guarantee.
Image input, tuning, function calling, and Structured Outputs have documented feature-combination restrictions. Teams should check the current CLOVA Studio compatibility rules before combining these features in one request, rather than assuming every supported feature can be enabled simultaneously.
Reasoning, coding, speed, and cost profile
HCX-005 is best understood as an instruction-following multimodal assistant rather than a specialist reasoning model. The supplied editorial evaluation gives it a reasoning score of 6 out of 10 and a coding score of 6 out of 10. These are comparative editorial scores, not provider-published benchmark results. They suggest a middle-of-the-range expectation for general reasoning and coding, but they should not be treated as measured accuracy or as a substitute for testing on representative tasks.
The same evaluation assigns HCX-005 a speed score of 7 out of 10 and a cost score of 6 out of 10. Again, these are editorial assessments rather than official performance or price claims. They indicate a reasonable practical balance for applications that need visual input, text generation, and interactive responses, but the actual experience will depend on prompt length, image processing, traffic, deployment configuration, and billing.
HCX-005 is not described as a web-search model, and the supplied research does not verify a built-in web-search capability. If an assistant must answer using current external information, the application should provide an appropriately controlled retrieval or search layer instead of assuming that the model can browse the web.
Pricing and commercial details
CLOVA Studio pricing for HCX-005 is stated in the supplied research as being calculated per 1,000 input tokens and per 1,000 output tokens. However, the current numeric amounts are not displayed in the accessible official pricing table used for this record. There is therefore no verified price figure to quote responsibly.
This means that HCX-005 should not be compared using an invented per-token rate. Before deployment, teams should confirm the current input and output prices in NAVER Cloud’s official commercial documentation or account console. They should also account for the fact that image inputs, long contexts, tool calls, and repeated retries can affect total usage even when the nominal billing unit is tokens.
Main strengths of HCX-005
- Visual understanding with text generation: It can interpret image inputs and explain, summarize, or reason about their contents in text.
- Large context capacity: The 128,000-token context window supports long prompts and substantial document-oriented workflows.
- Korean business fit: Its positioning is especially relevant to Korean-language applications and organizations evaluating NAVER Cloud services.
- Developer integration: Chat Completions v3, OpenAI-compatible access, streaming, function calling, and PEFT tuning support a range of application architectures.
- Practical visual limits: Support for common image formats, URLs, and data URIs makes it possible to integrate images without requiring a single fixed upload method.
Main limitations to consider
- Text-only output: HCX-005 does not generate images, audio, video, music, speech, or embeddings.
- Limited maximum response size: The output limit is 4,096 tokens, which may require multi-step generation for long deliverables.
- No verified built-in web search: Current web research must be supplied through an external retrieval or search system if required.
- No documented Structured Outputs: Applications requiring strict schemas need validation and fallback handling.
- Image-conversation restrictions: Up to five images may be included per request, but no more than one image is allowed per turn under the documented rules.
- Unclear current numeric pricing: The billing units are known, but the accessible official pricing source does not expose a current amount.
- Not the obvious choice for deep reasoning: Applications centered on complex reasoning should evaluate a dedicated thinking model, including the separately documented HCX-007, rather than assuming HCX-005 is optimized for that role.
Best use cases
HCX-005 is a strong candidate for Korean-language business assistants that need to combine written instructions with visual material. Examples include an assistant that explains an uploaded form, a customer-support workflow that classifies product photographs, a document tool that summarizes page images, or an internal application that sends visual evidence to a function-calling workflow.
It can also suit product catalog and commerce applications where images need to be interpreted alongside text, provided the surrounding system handles product data, authentication, and business rules. Its long context is useful when an application must provide policy documents, product information, or several related pieces of reference material in the same request.
PEFT tuning may be useful when an organization needs domain adaptation without changing the model’s entire parameter set. The supplied research confirms PEFT support but does not specify a guaranteed improvement for any particular dataset, so tuning should be evaluated against a representative validation set.
When to choose HCX-005
Choose HCX-005 when the central requirement is text generation informed by images, especially in Korean-language or NAVER Cloud-oriented business software. It is also a sensible option when streaming, function calling, OpenAI-compatible access, or PEFT tuning are important parts of the integration.
Consider another option when the application needs native media generation, speech output, audio or video understanding, guaranteed structured JSON, documented web search, or especially demanding multi-step reasoning. A text-only model may be simpler and more economical for workloads that never use images, while a dedicated reasoning model may be better for difficult analytical tasks. Conversely, a media-generation model is more appropriate when the desired output is an image, video, or audio asset rather than a textual explanation.
Overall, HCX-005 occupies a practical middle position: it adds image understanding and useful application controls without claiming to be a universal multimodal generator or a dedicated deep-reasoning system. Its suitability depends on whether those boundaries match the workflow being built.

