Parakeet CTC

Parakeet CTC 0.6B zh-CN

by NVIDIA AI · Current and downloadable; available through NVIDIA NIM and Riva interfaces

A specialized 600-million-parameter NVIDIA speech recognition model for Mandarin Chinese and American English, including code-switched audio. It supports streaming and offline transcription through downloadable, NIM, and Riva deployment paths, with maximum audio duration determined by available GPU memory.

Text Reasoning Coding
NVIDIA Parakeet CTC 0.6B zh-CN, also identified as Parakeet-CTC-0.6B-Unified, is a specialized speech recognition model for audio containing Mandarin Chinese, American English, or both languages in the same recording. It accepts mono audio and returns transcription text, making it suitable for meetings, interviews, contact centers, lectures, and other bilingual speech-to-text workflows.
Outputs

What Parakeet CTC 0.6B zh-CN can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Parakeet CTC
Model type Other
Status Current and downloadable; available through NVIDIA NIM and Riva interfaces
Knowledge cutoff notes

Not applicable to this speech recognition model. NVIDIA's official documentation describes training data and model architecture but does not publish a conventional knowledge-cutoff date.

Model notes

The exact NVIDIA model identifier is parakeet-ctc-0.6b-zh-cn. The model card names the model Parakeet-CTC-0.6B-Unified and lists version Parakeet-CTC-XL-unified-0.6b_spe7k_zh-CN_3.0. It is an approximately 600-million-parameter FastConformer-CTC ASR model trained on more than 17,000 hours of Mandarin Chinese and American English speech. The model accepts mono audio and returns transcription text. The official card states that maximum input duration depends on GPU memory and that no preprocessing is needed for the documented model input. It supports streaming and offline speech-to-text deployment. Reasoning, coding, context-window, maximum-output-token, JSON-mode, and general tool-use fields are not applicable to this specialist ASR model. Editorial scores reflect its specialist role rather than comparison with general-purpose language models.

Cost

Model pricing

Input No public per-token price; downloadable model and NVIDIA service terms apply
Output No public per-token price; downloadable model and NVIDIA service terms apply
Model guide

NVIDIA Parakeet CTC 0.6B zh-CN for Mandarin-English Speech Recognition

NVIDIA Parakeet CTC 0.6B zh-CN is a downloadable 600-million-parameter automatic speech recognition model for Mandarin Chinese and American English code-switching. Built on the FastConformer-CTC architecture, it supports streaming and offline transcription and is intended for production speech-to-text deployments through NVIDIA infrastructure.

What is NVIDIA Parakeet CTC 0.6B zh-CN?

NVIDIA Parakeet CTC 0.6B zh-CN is an automatic speech recognition (ASR) model. ASR systems convert spoken language in an audio recording into written text. This model is designed specifically for Mandarin Chinese and American English, including code-switching situations where a speaker changes between the two languages during one conversation.

The model contains approximately 600 million parameters and is available as a downloadable NVIDIA model. NVIDIA also makes it available through NVIDIA NIM and Riva-related interfaces, giving organizations options for self-hosted, containerized, or service-connected deployment. Its official model identifier is parakeet-ctc-0.6b-zh-cn. The model card identifies the underlying version as Parakeet-CTC-XL-unified-0.6b_spe7k_zh-CN_3.0.

This is not a general-purpose conversational model. It does not answer questions, summarize conversations by itself, generate code, or produce speech. Its job is narrower and more practical: transforming supported speech recordings into Mandarin-English transcription text.

Architecture and training data

Parakeet CTC 0.6B zh-CN uses the Parakeet-CTC architecture, also described by NVIDIA as FastConformer-CTC. Conformer architectures combine transformer-style attention with convolutional processing, which is useful for modeling both the broader context of speech and local acoustic patterns. CTC, or connectionist temporal classification, is a training approach commonly used to align audio frames with text without requiring every frame to have a manually assigned character or word label.

NVIDIA states that the model was trained on more than 17,000 hours of Mandarin Chinese and American English speech. The training mixture includes public and proprietary datasets. According to the model documentation, transcripts were normalized to support casing, punctuation, and spoken-form variations. The resulting output can include mixed-case text and punctuation such as periods, commas, question marks, spaces, and apostrophes.

The training-hour figure and architecture description are provider-published information. They should not be interpreted as a guarantee of accuracy for every accent, recording environment, or specialist vocabulary. Real-world performance can vary with microphone quality, background noise, overlapping speakers, speaking style, and terminology that is uncommon in the training distribution.

Supported inputs and outputs

The underlying model accepts one-dimensional mono audio. WAV is the primary format described in the model card, and NVIDIA's Riva interface documentation also demonstrates audio delivered in WAV, OGG, or OPUS containers. The documented NVIDIA deployment path does not require users to perform additional audio preprocessing.

The output is a one-dimensional text string containing Mandarin Chinese and English transcription. The model does not directly produce audio, images, video, embeddings, structured tool calls, or general-purpose text responses. Its output is transcription rather than translation, summarization, sentiment analysis, speaker labeling, or dialogue management.

CapabilityParakeet CTC 0.6B zh-CN
Primary taskAutomatic speech recognition
Audio inputYes; mono audio, with WAV specified for the model
Supported languagesMandarin Chinese and American English
Code-switchingDesigned for Mandarin-English mixed speech
Text outputYes; transcription text
Audio, image, or video outputNo
StreamingSupported
Tool or function callingNot applicable to the model itself

Audio duration and output limits

NVIDIA does not publish a conventional token context window or maximum output-token count for this ASR model. Those language-model fields are not directly applicable because the model consumes audio and emits a transcription rather than continuing a text prompt.

The maximum usable audio duration depends on available GPU memory. This means that there is no single duration limit that applies to every deployment. A larger or better-configured GPU may accommodate longer inputs, while a constrained environment may require the recording to be divided into shorter segments. The model documentation does not provide one universal maximum number of minutes that can be treated as a guaranteed limit.

For long recordings, applications may therefore need to manage chunking, streaming, storage, and the joining of transcription segments. Those surrounding application responsibilities are separate from the core model and may affect the final usability of a transcription workflow.

Deployment through NVIDIA infrastructure

Parakeet CTC 0.6B zh-CN is listed as a downloadable model in NVIDIA's model catalog. It is also exposed through NVIDIA NIM and Riva-related documentation. NIM provides a containerized deployment path for NVIDIA AI models, while Riva APIs provide speech-focused interfaces for transcription and streaming use cases.

NVIDIA's documentation describes Linux and Docker deployment options for the NIM path. The Riva API documentation uses the model-specific function identifier together with the zh-CN language code. This makes the model relevant to organizations that want to integrate transcription into an existing NVIDIA-based application, rather than only to users looking for a standalone desktop transcription tool.

The model card lists tested or supported NVIDIA hardware ranging from accelerators such as the A2, A10, A16, and L4 to higher-end data-center GPUs including the A100 and H100. Selected GeForce RTX systems are also listed. Actual throughput, concurrent request capacity, and maximum audio duration depend on the GPU, memory, software configuration, and deployment architecture.

Main strengths and trade-offs

The clearest strength of this model is specialization. Instead of trying to perform conversation, reasoning, coding, and many unrelated tasks, it focuses on Mandarin-English speech recognition. That makes its downloadable and self-hosted deployment model potentially useful where organizations need control over infrastructure or want to process speech within an NVIDIA environment.

Support for code-switching is another important distinction. A recording that alternates between Mandarin and American English is a more specific target than a monolingual transcription workload. The model also supports both streaming and offline deployment, allowing it to serve live or near-live applications as well as previously recorded files.

Its main trade-off is that it is not a complete speech analytics system. It returns transcription text, but additional components may be needed for speaker diarization, translation, summarization, keyword extraction, redaction, sentiment analysis, or workflow automation. The supplied research does not establish built-in support for those functions.

NVIDIA does not publish a public per-token price for this model. The downloadable model and NVIDIA NIM or Riva usage are governed by the applicable NVIDIA model, service, API, licensing, and infrastructure terms. Infrastructure costs can include GPU hardware, cloud GPU time, storage, networking, and operational support, but the supplied information does not provide a fixed cost estimate.

Reasoning, coding, and tool capabilities

Reasoning and coding are not meaningful primary capabilities for Parakeet CTC 0.6B zh-CN. It recognizes speech patterns and maps them to text; it is not intended to solve multi-step problems, write software, follow general conversational instructions, or produce a textual analysis from an arbitrary prompt.

The model also has no documented native tool or function-calling capability. A surrounding application can use the transcription as an input to other systems, but that orchestration belongs to the application layer. For example, a separate service could receive the transcript and then create a meeting summary, but that would not be a function performed by Parakeet itself.

Best use cases

  • Mandarin Chinese and American English transcription
  • Meetings, interviews, lectures, and recorded media containing bilingual speech
  • Code-switched customer-support or contact-center audio
  • Streaming speech-to-text applications
  • Self-hosted enterprise transcription on NVIDIA hardware
  • Offline processing of audio files where the organization controls the deployment environment

For example, a contact center could use the model to convert Mandarin-English calls into searchable text before applying separate analytics tools. A meeting application could use streaming recognition to display a live transcript, while an offline pipeline could process stored interviews in batches.

Limitations to consider

Parakeet CTC 0.6B zh-CN is specialized for speech recognition and should not be selected as a general-purpose AI assistant. It requires audio input and returns text. It does not natively generate spoken responses, images, video, embeddings, or conversational answers.

The official model card notes that the model does not handle special characters. Accuracy may also decline with strong accents, heavy background noise, poor microphones, overlapping speakers, unusual names, domain-specific terminology, or speech conditions that differ from the training data. The model card does not establish a conventional knowledge-cutoff date because this is not a knowledge-based language model.

Speaker diarization is not identified as a built-in capability in the supplied research. If an application needs reliable separation of multiple speakers, it may need an additional diarization system. Similarly, transcription output should not automatically be treated as a verified record: important business, medical, legal, or customer-service use cases may require human review or domain-specific quality checks.

When to choose this model

Choose Parakeet CTC 0.6B zh-CN when the central requirement is Mandarin-English speech recognition, especially when code-switching, streaming, downloadable deployment, or NVIDIA-based infrastructure matters. It is a sensible fit for teams that want a focused ASR model rather than paying for or operating a broader generative model with capabilities they do not need.

Another speech recognition option may be more appropriate if the required languages fall outside Mandarin Chinese and American English, if the workflow depends on a published universal duration limit, or if the application needs integrated speaker diarization, translation, summarization, or other speech analytics. A general-purpose language model may be better for reasoning over an existing transcript, but it would not replace this model's dedicated audio-recognition role.

In cost and speed terms, the supplied research does not provide benchmark numbers or a fixed service price. The practical trade-off is therefore deployment-dependent: local or self-hosted use can provide control over data and throughput, while it requires compatible NVIDIA hardware and operational work. Actual speed and cost should be measured on the target GPU, audio mix, concurrency level, and application pipeline rather than inferred from the parameter count alone.

Licensing and current catalog position

Parakeet CTC 0.6B zh-CN is a current, downloadable NVIDIA model available alongside NIM and Riva deployment paths. Use is governed by the NVIDIA Community Model License and any applicable NVIDIA service or API terms. Organizations should review those terms before redistributing the model or placing it into a commercial production workflow.

Overall, this model is best understood as a focused Mandarin-English transcription component. Its value comes from its language pairing, code-switching orientation, streaming support, and fit with NVIDIA deployment infrastructure—not from general reasoning, coding, or multimodal generation.


Answers to Frequently Asked Questions

What are the main limitations of NVIDIA Parakeet CTC 0.6B zh-CN?
The model is a focused speech-to-text system rather than a general-purpose AI assistant. Accuracy may decrease with strong accents, background noise, poor microphones, overlapping speakers, unusual names, specialized terminology, and unsupported special characters. Speaker diarization, translation, summarization, and other speech analytics require additional systems, and important transcripts may need human review.
Can NVIDIA Parakeet CTC 0.6B zh-CN be deployed for streaming and self-hosted use?
Yes. The model supports streaming and offline transcription and is available as a downloadable NVIDIA model through NVIDIA NIM and Riva-related deployment paths. It can be run on supported NVIDIA hardware, although performance, concurrency, and maximum audio duration depend on the GPU, memory, software configuration, and deployment architecture.
What audio formats and outputs does Parakeet CTC 0.6B zh-CN support?
The underlying model accepts one-dimensional mono audio, with WAV specified in the model card. NVIDIA Riva documentation also demonstrates WAV, OGG, and OPUS inputs. The output is transcription text; the model does not directly produce audio, translations, summaries, speaker labels, images, or video.
What is NVIDIA Parakeet CTC 0.6B zh-CN used for?
NVIDIA Parakeet CTC 0.6B zh-CN is an automatic speech recognition model that converts Mandarin Chinese and American English speech into text. It is designed for bilingual and code-switched audio, including meetings, interviews, contact-center calls, lectures, and streaming transcription applications.
Does NVIDIA Parakeet CTC 0.6B zh-CN support Mandarin-English code-switching?
Yes. The model is specifically designed for speech that switches between Mandarin Chinese and American English during the same conversation or recording.


Sources 5
Provider

About NVIDIA AI