HyperCLOVA X

HCX-005

by NAVER AI · current

HCX-005 is NAVER Cloud’s HyperCLOVA X multimodal model for text and image understanding. It offers a 128,000-token context window, 4,096-token maximum output, streaming, function calling, OpenAI-compatible access, and PEFT tuning. The model is suited to Korean-language business assistants, document analysis, and visual workflows, but it produces text only, lacks verified built-in web search and Structured Outputs, and has no publicly verified numeric price in the accessible official pricing table.

Text Reasoning Coding
HCX-005 is a multimodal model in NAVER Cloud’s HyperCLOVA X family, designed for applications that combine language understanding with image analysis. It can process text and images and produce text responses, making it suitable for visual question answering, document interpretation, Korean-language business assistants, and workflows that need structured interaction with external functions. The model offers a 128,000-token context window, a maximum output of 4,096 tokens, streaming responses, function calling, OpenAI-compatible access, and PEFT tuning. Its main trade-off is scope: HCX-005 is a text-output vision model, not a native image, audio, or video generator, and the accessible official pricing information confirms the billing unit but not a current numeric rate.
Outputs

What HCX-005 can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Streaming Fine-tuning
Model profile

Performance characteristics

6/10 Reasoning
6/10 Coding
7/10 Speed
6/10 Cost efficiency
Specifications

Technical details

Model family HyperCLOVA X
Model type Multimodal
Context window 128K tokens
Maximum output 4K tokens
Release date 2025-04-17
Status current
Knowledge cutoff notes

NAVER's official HCX-005 documentation does not provide a specific knowledge-cutoff date.

Model notes

HCX-005 is a HyperCLOVA X multimodal vision model provided through CLOVA Studio. It accepts text and image inputs and returns text. The model supports Chat Completions v3, OpenAI-compatible access, function calling, streaming responses, and PEFT tuning. Image input, tuning, function calling, and Structured Outputs are subject to documented feature-combination restrictions. Images can be supplied by public URL or data URI; supported formats include BMP, PNG, JPG, JPEG, and WEBP. Up to five images may be included per request, with at most one image per turn. Structured Outputs are documented as unsupported for HCX-005, while a separate legacy JSON-mode capability is not verified. The model was introduced on April 17, 2025, with a 128,000-token context length and 4,096-token maximum output.

Cost

Model pricing

Input Priced per 1,000 input tokens; current amount not displayed in the accessible official pricing table.
Output Priced per 1,000 output tokens; current amount not displayed in the accessible official pricing table.
Model guide

HCX-005: NAVER’s Multimodal HyperCLOVA X Model for Korean Business Applications

HCX-005 is a current HyperCLOVA X multimodal language model from NAVER Cloud. It accepts text and images, returns text, supports a 128,000-token context window, and can generate up to 4,096 tokens per response. Its main strengths are Korean-language applications, image and document understanding, instruction following, function calling, streaming, and PEFT tuning. It is a practical choice for API-based assistants that need visual analysis without native image, audio, or video generation, but its exact current token prices are not publicly displayed in the accessible official pricing table.

What is HCX-005?

HCX-005 is a multimodal language model provided by NAVER Cloud as part of the HyperCLOVA X model family. The model was introduced on April 17, 2025, and is available through CLOVA Studio and related developer interfaces. In practical terms, “multimodal” means that HCX-005 can work with more than text: it accepts images as well as written prompts, then returns a text response.

This makes HCX-005 different from a conventional text-only language model, but its output scope remains deliberately narrow. The supplied specifications identify text as its output type; it does not natively generate images, audio, video, music, embeddings, or speech. Its visual capability is therefore primarily for understanding and analysis rather than media creation.

NAVER Cloud positions HCX-005 for Korean-language business applications, image understanding, document and visual analysis, instruction following, and API-based assistants. Those are provider-oriented use cases. The assessment that it is a good fit for a particular application should still depend on the quality of the application’s prompts, images, retrieval system, and tool integrations.

Where HCX-005 fits in the HyperCLOVA X family

HCX-005 belongs to NAVER Cloud’s current HyperCLOVA X lineup and is specifically identified in the supplied documentation as a vision-capable model. Its role is not to cover every possible AI workload, but to provide language generation with image input and developer-oriented interaction features.

The model documentation identifies HCX-007 as a separate thinking model. That distinction matters when selecting a model: HCX-005 is aimed at multimodal understanding and practical assistant workflows, while a dedicated thinking model may be more appropriate for applications whose central requirement is deeper reasoning. The available research does not provide a direct benchmark comparison between the two models, so this is a positioning distinction rather than a claim that one model is universally more accurate.

HCX-005 is also more specialized than a general-purpose media model. It can interpret visual material, but the supplied specifications do not support using it for image generation, video creation, voice synthesis, or music production.

Core capabilities and technical limits

SpecificationHCX-005
ProviderNAVER Cloud
Model familyHyperCLOVA X
Release dateApril 17, 2025
InputText and images
OutputText
Context length128,000 tokens
Maximum output4,096 tokens
Image limitUp to five images per request, with at most one image per turn
Image formatsBMP, PNG, JPG, JPEG, and WEBP
Image deliveryPublic URL or data URI
StreamingSupported
Function callingSupported
Fine-tuningPEFT tuning supported
Structured OutputsNot supported according to the supplied documentation

The 128,000-token context window is useful for long prompts, document collections, or conversations that need substantial background material. A context window is the amount of input and generated text the model can consider within a request; it is not the same as a guarantee that every long document will be understood equally well. The maximum generated response is 4,096 tokens, so applications that need very long reports may need to request multiple sections or manage continuation carefully.

HCX-005 can receive up to five images in one request, but the documented restriction of at most one image per turn is important when designing a conversation. Images may be supplied through a public URL or a data URI, and the documented formats include BMP, PNG, JPG, JPEG, and WEBP. These details make it suitable for workflows such as analyzing a small group of page images, product photographs, charts, or scanned business material, subject to the model’s actual visual interpretation quality.

Tools, API access, and output formats

HCX-005 supports Chat Completions v3, OpenAI-compatible access, streaming responses, and function calling. Function calling allows the model to request an application-defined operation, such as looking up an order, querying an internal database, or creating a reservation. The model does not perform those external actions by itself; the surrounding application must validate the request, execute the function, and return the result.

Streaming is useful when an application wants to display generated text progressively rather than waiting for the complete response. This can improve perceived responsiveness, particularly for assistants and interactive document tools.

Structured Outputs are documented as unsupported for HCX-005. A separate legacy JSON-mode capability is not verified in the supplied research, so developers should not assume that requesting JSON guarantees schema-valid output. If an application needs machine-readable responses, it should use careful prompting, validation, error handling, and potentially function calling rather than treating ordinary text generation as a formal schema guarantee.

Image input, tuning, function calling, and Structured Outputs have documented feature-combination restrictions. Teams should check the current CLOVA Studio compatibility rules before combining these features in one request, rather than assuming every supported feature can be enabled simultaneously.

Reasoning, coding, speed, and cost profile

HCX-005 is best understood as an instruction-following multimodal assistant rather than a specialist reasoning model. The supplied editorial evaluation gives it a reasoning score of 6 out of 10 and a coding score of 6 out of 10. These are comparative editorial scores, not provider-published benchmark results. They suggest a middle-of-the-range expectation for general reasoning and coding, but they should not be treated as measured accuracy or as a substitute for testing on representative tasks.

The same evaluation assigns HCX-005 a speed score of 7 out of 10 and a cost score of 6 out of 10. Again, these are editorial assessments rather than official performance or price claims. They indicate a reasonable practical balance for applications that need visual input, text generation, and interactive responses, but the actual experience will depend on prompt length, image processing, traffic, deployment configuration, and billing.

HCX-005 is not described as a web-search model, and the supplied research does not verify a built-in web-search capability. If an assistant must answer using current external information, the application should provide an appropriately controlled retrieval or search layer instead of assuming that the model can browse the web.

Pricing and commercial details

CLOVA Studio pricing for HCX-005 is stated in the supplied research as being calculated per 1,000 input tokens and per 1,000 output tokens. However, the current numeric amounts are not displayed in the accessible official pricing table used for this record. There is therefore no verified price figure to quote responsibly.

This means that HCX-005 should not be compared using an invented per-token rate. Before deployment, teams should confirm the current input and output prices in NAVER Cloud’s official commercial documentation or account console. They should also account for the fact that image inputs, long contexts, tool calls, and repeated retries can affect total usage even when the nominal billing unit is tokens.

Main strengths of HCX-005

  • Visual understanding with text generation: It can interpret image inputs and explain, summarize, or reason about their contents in text.
  • Large context capacity: The 128,000-token context window supports long prompts and substantial document-oriented workflows.
  • Korean business fit: Its positioning is especially relevant to Korean-language applications and organizations evaluating NAVER Cloud services.
  • Developer integration: Chat Completions v3, OpenAI-compatible access, streaming, function calling, and PEFT tuning support a range of application architectures.
  • Practical visual limits: Support for common image formats, URLs, and data URIs makes it possible to integrate images without requiring a single fixed upload method.

Main limitations to consider

  • Text-only output: HCX-005 does not generate images, audio, video, music, speech, or embeddings.
  • Limited maximum response size: The output limit is 4,096 tokens, which may require multi-step generation for long deliverables.
  • No verified built-in web search: Current web research must be supplied through an external retrieval or search system if required.
  • No documented Structured Outputs: Applications requiring strict schemas need validation and fallback handling.
  • Image-conversation restrictions: Up to five images may be included per request, but no more than one image is allowed per turn under the documented rules.
  • Unclear current numeric pricing: The billing units are known, but the accessible official pricing source does not expose a current amount.
  • Not the obvious choice for deep reasoning: Applications centered on complex reasoning should evaluate a dedicated thinking model, including the separately documented HCX-007, rather than assuming HCX-005 is optimized for that role.

Best use cases

HCX-005 is a strong candidate for Korean-language business assistants that need to combine written instructions with visual material. Examples include an assistant that explains an uploaded form, a customer-support workflow that classifies product photographs, a document tool that summarizes page images, or an internal application that sends visual evidence to a function-calling workflow.

It can also suit product catalog and commerce applications where images need to be interpreted alongside text, provided the surrounding system handles product data, authentication, and business rules. Its long context is useful when an application must provide policy documents, product information, or several related pieces of reference material in the same request.

PEFT tuning may be useful when an organization needs domain adaptation without changing the model’s entire parameter set. The supplied research confirms PEFT support but does not specify a guaranteed improvement for any particular dataset, so tuning should be evaluated against a representative validation set.

When to choose HCX-005

Choose HCX-005 when the central requirement is text generation informed by images, especially in Korean-language or NAVER Cloud-oriented business software. It is also a sensible option when streaming, function calling, OpenAI-compatible access, or PEFT tuning are important parts of the integration.

Consider another option when the application needs native media generation, speech output, audio or video understanding, guaranteed structured JSON, documented web search, or especially demanding multi-step reasoning. A text-only model may be simpler and more economical for workloads that never use images, while a dedicated reasoning model may be better for difficult analytical tasks. Conversely, a media-generation model is more appropriate when the desired output is an image, video, or audio asset rather than a textual explanation.

Overall, HCX-005 occupies a practical middle position: it adds image understanding and useful application controls without claiming to be a universal multimodal generator or a dedicated deep-reasoning system. Its suitability depends on whether those boundaries match the workflow being built.


Answers to Frequently Asked Questions

When should a business choose HCX-005?
HCX-005 is a good choice when an application needs Korean-language text generation informed by images, particularly for business assistants, document analysis, product-image workflows, and visual customer support. Another model may be more suitable for deep reasoning, native media generation, speech, guaranteed structured JSON, or built-in web search.
Does HCX-005 support function calling and structured outputs?
HCX-005 supports function calling, streaming, Chat Completions v3, OpenAI-compatible access, and PEFT tuning. Structured Outputs are not supported according to the supplied documentation, so applications requiring strict schemas should use validation, error handling, and carefully designed prompts.
Does HCX-005 generate images, audio, video, or speech?
No. HCX-005 accepts images for understanding and analysis, but its output is text only. It does not natively generate images, audio, video, music, speech, or embeddings.
What is HCX-005 and what can it do?
HCX-005 is a multimodal language model from NAVER Cloud’s HyperCLOVA X family. It accepts text and images as input and generates text responses, making it suitable for Korean-language business assistants, document analysis, image understanding, and visual question-answering workflows.
What are the main technical specifications of HCX-005?
HCX-005 supports a 128,000-token context window, generates up to 4,096 tokens, and accepts up to five images per request with a maximum of one image per turn. It supports BMP, PNG, JPG, JPEG, and WEBP images delivered through public URLs or data URIs.


Sources 7
Provider

About NAVER AI