Parakeet

Parakeet CTC 1.1B

by NVIDIA AI · Available open-weight checkpoint

NVIDIA Parakeet CTC 1.1B is an approximately 1.1-billion-parameter English speech recognition checkpoint using FastConformer CTC. It accepts 16 kHz mono audio, produces lowercase transcripts, and supports local inference, fine-tuning, offline processing, and buffered streaming through compatible NVIDIA runtimes.

Text Reasoning Coding
NVIDIA Parakeet CTC 1.1B is a roughly 1.1-billion-parameter speech-to-text model developed jointly by NVIDIA NeMo and Suno.ai. It is designed for English transcription rather than conversation or general text generation. The open checkpoint can be used with NVIDIA NeMo and Hugging Face Transformers, while NeMo-Speech.cpp provides lightweight local and buffered-streaming deployment options.
Outputs

What Parakeet CTC 1.1B can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Parakeet
Model type Other
Status Available open-weight checkpoint
Knowledge cutoff notes

A model knowledge cutoff is not applicable or documented for this speech recognition checkpoint. It transcribes supplied audio rather than answering questions from a stated training-data cutoff.

Model notes

Canonical model identifier: nvidia/parakeet-ctc-1.1b. The model is an approximately 1.1-billion-parameter FastConformer CTC ASR checkpoint jointly developed by NVIDIA NeMo and Suno.ai. It accepts 16 kHz mono audio and produces lowercase English transcription text. The checkpoint is available for inference and fine-tuning through NVIDIA NeMo and Transformers. NeMo-Speech.cpp supports whole-file recognition and overlapping buffered streaming, while optional external components can provide punctuation, capitalization, inverse text normalization, word boosting, diarization, or language-model rescoring. No standard hosted token pricing applies to the open checkpoint itself. The model card lists CC-BY-4.0 licensing.

Model guide

NVIDIA Parakeet CTC 1.1B: Fast English ASR for Local and Offline Transcription

NVIDIA Parakeet CTC 1.1B is an open-weight English automatic speech recognition model built with a FastConformer CTC architecture. It converts 16 kHz mono audio into lowercase English text and is suited to local inference, offline transcription, fine-tuning, and buffered streaming through compatible NVIDIA runtimes.

What is NVIDIA Parakeet CTC 1.1B?

NVIDIA Parakeet CTC 1.1B is an open-weight automatic speech recognition (ASR) model. ASR systems take spoken audio as input and produce a written transcript. This model is published under the nvidia/parakeet-ctc-1.1b identifier and contains approximately 1.1 billion parameters.

The model was developed jointly by the NVIDIA NeMo and Suno.ai teams. Its architecture combines a FastConformer encoder with a connectionist temporal classification (CTC) output head. In practical terms, FastConformer processes the acoustic features of speech efficiently, while CTC provides a direct way to align the audio with the characters or tokens in the resulting transcript.

Parakeet CTC 1.1B is a focused speech-recognition checkpoint, not a conversational language model. It does not answer questions about a recording, generate a dialogue, or perform general reasoning. Its primary job is to turn English speech into text.

Purpose and position in NVIDIA's lineup

The checkpoint sits in NVIDIA's speech and developer ecosystem rather than in a consumer chatbot product. It can be loaded locally through NVIDIA NeMo or used with Hugging Face Transformers. NVIDIA's NeMo-Speech.cpp project supplies a native C++ runtime and GGUF distributions intended for local inference. NVIDIA also documents related Parakeet deployment through speech-focused NIM and Riva tooling, although packaging, hardware requirements, and runtime behavior can differ between those products and the open checkpoint.

This positioning makes Parakeet CTC 1.1B most relevant to developers building transcription pipelines, applications that extract text from recorded audio, local voice interfaces, and customized speech models. It is not presented as a metered first-party token API model. Users typically provide their own hardware or choose a separate hosted NVIDIA deployment product.

Inputs, outputs, and supported modalities

The documented model input is 16 kHz mono-channel audio. WAV files are supported in the documented NeMo workflow. The native output is a text transcription string consisting of lowercase English alphabet output.

CapabilityParakeet CTC 1.1B
Primary taskEnglish automatic speech recognition
Audio inputSupported; 16 kHz mono audio
Text outputSupported; lowercase English transcription
Image input or outputNot supported
Video input or outputNot supported
Audio output or speech synthesisNot supported
Native tool or function callingNot supported
Structured JSON outputNot a native model capability

The model does not natively add full capitalization or punctuation. NeMo-Speech.cpp documentation describes optional components for punctuation and capitalization restoration, inverse text normalization, word boosting, diarization, and language-model rescoring. These are separate processing or decoding components, so their results should not be confused with the raw output of the acoustic ASR checkpoint.

Recognition accuracy and training data

The model card reports training on approximately 64,000 hours of English speech. The training mixture includes private NVIDIA and Suno.ai data along with public sources such as LibriSpeech, Fisher, Switchboard, VCTK, VoxPopuli, Europarl-ASR, Mozilla Common Voice, and People's Speech.

Reported greedy-decoding word error rates include 1.83% on LibriSpeech test-clean, 3.54% on VoxPopuli, 4.20% on TEDLIUM-v3, 10.27% on GigaSpeech, 13.69% on Earnings-22, 15.62% on AMI, 3.54% on SPGI Speech, and 6.53% on Common Voice. These are benchmark figures from the model card, using CTC greedy decoding without an external language model. They are useful for comparison, but they do not guarantee the same performance on every microphone, accent, speaker group, or industry vocabulary.

In real deployments, accuracy can change substantially with background noise, overlapping speakers, recording quality, pronunciation, and domain-specific terms. A test on representative recordings is therefore more informative than relying on one benchmark number.

Local, offline, and streaming deployment

With NVIDIA NeMo, the model can be loaded using the nvidia/parakeet-ctc-1.1b pretrained identifier. Hugging Face Transformers also supports the checkpoint through an automatic speech recognition pipeline or AutoModelForCTC.

NeMo-Speech.cpp is the most relevant option when a smaller native runtime or offline processing is important. Its documentation describes whole-file recognition and overlapping buffered streaming for Parakeet CTC. Whole-file recognition is appropriate for recordings such as interviews, meetings, lectures, and call archives. Buffered streaming can reduce the wait before partial processing in live or near-live applications, although the exact latency depends on the runtime, audio buffering, and hardware.

CTC decoding can use greedy recognition by default. An optional Flashlight and KenLM decoder can perform beam-search rescoring, which may improve recognition of domain-specific language when configured appropriately. That extra decoding stage adds complexity and should be evaluated against the speed requirements of the application.

Main strengths and trade-offs

  • Focused design: The model is specialized for English speech recognition rather than burdened with unrelated text-generation features.
  • Local control: The checkpoint and compatible runtimes support local or offline processing, which can be useful when audio should remain on an organization's hardware.
  • Multiple deployment paths: Developers can choose NeMo, Transformers, NeMo-Speech.cpp, or a separate NVIDIA speech deployment product depending on their environment.
  • Streaming support through compatible runtime software: NeMo-Speech.cpp supports overlapping buffered streaming, not just batch transcription.
  • Fine-tuning potential: The checkpoint can be used as a starting point for fine-tuning with NVIDIA NeMo or compatible Transformers workflows.

The principal trade-off is specialization. Parakeet CTC 1.1B is not a multilingual transcription model, a general-purpose language model, or a speech-to-speech assistant. It produces lowercase, normally unpunctuated text unless separate post-processing is added. A deployment that needs polished documents may therefore require punctuation restoration, capitalization, inverse text normalization, diarization, or custom vocabulary handling.

Its approximate 1.1-billion-parameter size may also require more computing resources than very small edge-oriented speech models, while offering a stronger accuracy target on the documented English benchmarks. The supplied research does not provide a universal hardware requirement or a fixed real-time factor, so speed should be measured on the target device rather than inferred from the parameter count alone.

Reasoning, coding, tools, and limits

Parakeet CTC 1.1B does not provide a conversational reasoning capability. It does not analyze the meaning of a transcript, write software, call external tools, browse the web, or generate structured responses through a native JSON mode. Its output is transcription text, and its useful input is supplied audio rather than text prompts.

No context-window size or maximum output-token limit is documented for this checkpoint in the supplied research. For an ASR model, practical processing limits depend on the selected runtime, audio duration, memory, and streaming or batching configuration. These should not be replaced with an invented token limit.

Pricing, license, and availability

The open checkpoint has no standard input-token or output-token price. It is downloaded and run through a compatible framework or runtime, so the direct cost depends on the user's hardware, hosting arrangement, and any separate NVIDIA deployment product used.

The Hugging Face model card identifies the checkpoint as available under the CC-BY-4.0 license. Users should review that license and the terms of any surrounding runtime or hosted service before commercial deployment. The open model's availability should also be distinguished from paid NVIDIA enterprise or hosted offerings, whose licensing and pricing may be separate.

When to choose Parakeet CTC 1.1B

Choose Parakeet CTC 1.1B when the priority is English speech-to-text with local control, offline operation, NeMo integration, or the ability to fine-tune a substantial open checkpoint. It is a sensible candidate for recorded-audio transcription, searchable media archives, audio extraction for retrieval systems, internal meeting processing, and customized NVIDIA speech pipelines.

Another option may be more appropriate when the application needs multilingual recognition, built-in punctuation-rich transcripts, speaker diarization without extra components, a managed pay-as-you-go API, or a conversational voice assistant. A smaller ASR model may be preferable on constrained hardware, while a hosted speech service may reduce operational work. Conversely, a general-purpose language or multimodal model is a better fit when the task begins after transcription—for example, summarizing, answering questions about, or extracting structured information from the recognized text.

Overall, Parakeet CTC 1.1B is best understood as an open, English-focused transcription engine. Its value comes from the combination of documented recognition performance, local deployment paths, and fine-tuning support—not from general reasoning, generation, or broad multimodal interaction.


Answers to Frequently Asked Questions

What are the license and pricing terms for NVIDIA Parakeet CTC 1.1B?
The Hugging Face model card identifies the checkpoint as available under the CC-BY-4.0 license. It has no standard input- or output-token pricing because users download and run it through their own hardware, hosting setup, or a separate NVIDIA deployment product. Runtime, hosted-service, and enterprise licensing terms may be separate.
How accurate is NVIDIA Parakeet CTC 1.1B?
The model card reports greedy-decoding word error rates such as 1.83% on LibriSpeech test-clean, 3.54% on VoxPopuli, 4.20% on TEDLIUM-v3, and 10.27% on GigaSpeech. Actual performance depends on recording quality, background noise, accents, overlapping speakers, and domain-specific vocabulary, so testing representative audio is recommended.
Can NVIDIA Parakeet CTC 1.1B run locally and offline?
Yes. The model can be run locally through NVIDIA NeMo, Hugging Face Transformers, or the native C++ NeMo-Speech.cpp runtime. NeMo-Speech.cpp supports whole-file transcription and overlapping buffered streaming, allowing offline or near-real-time processing on compatible hardware.
Does Parakeet CTC 1.1B support punctuation, capitalization, and speaker diarization?
The raw model output is lowercase English transcription without full capitalization or punctuation. Separate processing or decoding components can provide punctuation and capitalization restoration, inverse text normalization, word boosting, language-model rescoring, and diarization.
What is NVIDIA Parakeet CTC 1.1B used for?
NVIDIA Parakeet CTC 1.1B is an open-weight automatic speech recognition model designed to convert 16 kHz mono English audio into lowercase text transcripts. It is suitable for recorded-audio transcription, meeting and lecture processing, searchable media archives, local voice interfaces, and customized speech pipelines.


Sources 5
Provider

About NVIDIA AI