Parakeet TDT

Parakeet TDT 0.6B v2

by NVIDIA AI · Current open-weight model; publicly available through Hugging Face and NVIDIA NeMo

NVIDIA Parakeet TDT 0.6B v2 is an open-weight English speech-recognition model built with FastConformer-TDT. It focuses on fast transcription, punctuation, word and segment timestamps, song-to-lyrics use cases, local inference, NeMo workflows, and fine-tuning. It is not a general-purpose language, reasoning, coding, or media-generation model.

Text Reasoning Coding
NVIDIA Parakeet TDT 0.6B v2 is a specialized English speech-to-text model for turning supplied audio into punctuated text. It is aimed at users who need high transcription speed, timestamps, and local or NeMo-based deployment rather than general conversation, text generation, or multilingual speech processing.
Outputs

What Parakeet TDT 0.6B v2 can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming Fine-tuning
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
10/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Parakeet TDT
Model type Other
Release date 2025-06-04
Status Current open-weight model; publicly available through Hugging Face and NVIDIA NeMo
Knowledge cutoff notes

Knowledge cutoff is not applicable or published for this automatic speech recognition model. It transcribes supplied audio rather than answering from a fixed text knowledge corpus.

Model notes

The topic name was normalized to the verified canonical model identity nvidia/parakeet-tdt-0.6b-v2. NVIDIA describes it as a 600-million-parameter English ASR model with FastConformer-TDT architecture, punctuation, word-level timestamps, and song-to-lyrics transcription. The Hugging Face model card lists NeMo loading and fine-tuning workflows and identifies the license as CC BY 4.0. NVIDIA reported a 6.05% WER and 3386.02 real-time factor in a June 4, 2025 evaluation context. No model-specific token pricing applies to the open-weight checkpoint; hosted NVIDIA API pricing was not verified in the authoritative model documentation. Context length and maximum output-token fields are not applicable or not published for this speech-transcription checkpoint. Streaming is supported through compatible NeMo workflows, while hosted-service behavior can vary by deployment.

Model guide

NVIDIA Parakeet TDT 0.6B v2: Fast English Speech Transcription

NVIDIA Parakeet TDT 0.6B v2 is an open-weight, 600-million-parameter automatic speech recognition model for English audio. Its FastConformer-TDT architecture is designed for rapid transcription and supports punctuation, character-, word-, and segment-level timestamps, song-to-lyrics transcription, local inference, streaming-oriented workflows, and fine-tuning through NVIDIA NeMo.

What is NVIDIA Parakeet TDT 0.6B v2?

NVIDIA Parakeet TDT 0.6B v2 is an open-weight automatic speech recognition (ASR) model. ASR systems listen to supplied audio and produce a text transcription; unlike a general-purpose language model, Parakeet is not designed to answer questions, write prose, or reason over a broad knowledge base.

The model contains approximately 600 million parameters and is intended specifically for English speech. Its output can include punctuation and timing information at the character, word, and segment levels. That makes it useful not only for plain transcripts but also for captions, searchable media, and workflows that need to align text with the original audio.

Parakeet TDT 0.6B v2 is provided by NVIDIA and is publicly available through the nvidia/parakeet-tdt-0.6b-v2 model repository on Hugging Face. It fits into NVIDIA's speech AI and NeMo ecosystem as a focused, high-speed transcription checkpoint rather than as a general-purpose conversational model.

Architecture and primary purpose

The model uses a FastConformer-TDT architecture. FastConformer is an efficient variant of the Conformer architecture commonly used for speech recognition. TDT stands for Token-and-Duration Transducer, a transducer approach that predicts transcription tokens together with duration information. In practical terms, this design is intended to support efficient decoding while retaining useful alignment information.

Its primary purpose is high-throughput English transcription. Typical applications include meeting and interview transcripts, caption generation, media indexing, voice-interface preprocessing, audio search, and song-to-lyrics transcription. The model can be used for offline processing of local files, while compatible NVIDIA NeMo workflows also support streaming-oriented or buffered inference.

Capabilities and supported inputs and outputs

The model accepts audio and produces text transcription. Audio is its relevant input modality, and text is its direct output modality. It does not natively generate images, video, speech, music, or other media.

  • English automatic speech recognition.
  • Punctuated text transcription.
  • Character-level, word-level, and segment-level timestamps through NeMo transcription workflows.
  • Song-to-lyrics transcription.
  • Offline transcription of local audio.
  • Streaming-oriented and buffered inference through compatible NeMo tooling.
  • Fine-tuning and adaptation through NVIDIA NeMo.

Timestamp support is especially useful when the transcript must remain synchronized with audio. For example, a captioning pipeline can use word or segment timing to place text on a video, while an audio-search system can use timestamps to jump from a search result to the relevant point in a recording.

Performance, speed, and cost positioning

NVIDIA reported a 6.05% word error rate and a real-time factor of 3386.02 in June 2025 technical coverage. These are provider-reported results from a particular evaluation context, not guarantees for every recording, accent, microphone, language variety, noise condition, or hardware setup. Word error rate measures transcription mistakes, while real-time factor compares processing speed with the duration of the audio; the exact practical result depends on the complete inference environment.

The model's main trade-off is specialization. A 600-million-parameter ASR checkpoint can be a more suitable and potentially more economical choice for transcription than using a larger general-purpose multimodal model for the same task, particularly when the workload can run locally on available NVIDIA hardware. However, the supplied documentation does not provide a universal hosted price or a model-specific token price. The open-weight checkpoint itself is available for local use under the stated license, while any infrastructure, hardware, or hosted NVIDIA service costs are separate.

Users should therefore distinguish between model availability and total operating cost. Local deployment may avoid per-request API charges but can require compatible compute, storage, setup, and maintenance. Hosted deployment may simplify operations, but NVIDIA's model-specific hosted pricing was not verified in the supplied sources.

How to access and deploy it

The canonical model identifier is nvidia/parakeet-tdt-0.6b-v2. NVIDIA documents local loading and inference through NVIDIA NeMo, its open-source framework for speech and other generative AI workflows. NeMo also provides the relevant transcription and fine-tuning workflows.

NVIDIA's model documentation additionally describes access through a hosted NVIDIA API using the Riva client. Hosted behavior can differ from local behavior: availability, latency, chunking, timestamp handling, and streaming support may depend on the service and deployment version. The supplied research does not establish a single universal context window or maximum output-token limit for this speech-transcription checkpoint. Those language-model limits are not directly applicable in the same way as they are for a text-generation API.

The Hugging Face model card identifies the model as being distributed under the CC BY 4.0 license. Anyone planning commercial use should review the license terms and separately verify the terms that apply to NVIDIA-hosted services, infrastructure, and any associated data handling.

Reasoning, coding, and tool support

Parakeet TDT 0.6B v2 is not a reasoning model in the usual language-model sense. It does not independently analyze a question, plan a multi-step answer, or use a knowledge corpus to produce explanations. Its task is to decode supplied speech into text.

It also has no native coding capability. It can transcribe spoken programming terms or dictated code as audio, but it does not generate, test, execute, or edit software as a coding assistant.

No native function calling, tool use, web search, or code execution capability is documented for the checkpoint. Developers can place it inside a larger application that performs actions after transcription—for example, a voice interface can send the recognized text to another system—but that orchestration belongs to the surrounding application, not to Parakeet itself.

Important limitations

  • English focus: The model is intended for English transcription. It should not be selected when broad multilingual coverage is a core requirement.
  • Not speaker diarization: The model should not be treated as a complete system for identifying which person spoke when.
  • Not a general-purpose assistant: It does not provide open-ended conversation, text generation, web research, image generation, or question answering.
  • No native media generation: It transcribes audio but does not generate speech, music, images, or video.
  • Streaming depends on the stack: NeMo supports streaming-oriented workflows, but buffering, latency, chunking, and timestamp behavior may vary with the selected configuration and deployment.
  • Hardware and setup requirements: Local inference performance depends on the available environment and compatible NVIDIA software and hardware. The reported benchmark figures should not be assumed for every system.

Because this is an ASR model, a conventional knowledge-cutoff date is not applicable or published in the supplied documentation. It transcribes the audio provided at inference time rather than answering from a fixed text knowledge corpus.

When to choose Parakeet TDT 0.6B v2

Choose Parakeet TDT 0.6B v2 when the central requirement is fast English transcription and you want control over local or NeMo-based deployment. It is a strong fit for batch media processing, captions, meeting records, audio archives, voice-interface pipelines, and applications that benefit from word or segment timestamps. Its open-weight distribution can also be attractive when a team wants to integrate the checkpoint into its own infrastructure rather than depend entirely on a general-purpose hosted assistant.

Another ASR option may be more appropriate when the project requires languages beyond English, built-in speaker diarization, a managed service with clearly published usage pricing, or a simpler turnkey workflow. A general multimodal language model may be preferable when transcription is only the first step and the same system must then summarize, reason over, extract structured information from, or act on the transcript. That choice adds broader capabilities but may involve different cost, latency, privacy, and deployment trade-offs.

Parakeet TDT 0.6B v2 is best understood as a focused transcription engine: its value comes from converting English speech into accurately aligned text quickly, not from trying to replace a conversational model or a complete audio-processing platform.


Answers to Frequently Asked Questions

What are the main limitations of NVIDIA Parakeet TDT 0.6B v2?
The model is intended primarily for English speech and is not a general-purpose conversational, reasoning, or coding model. It does not natively provide speaker diarization, web search, tool use, question answering, or media generation. Its performance also depends on audio quality, accents, noise conditions, hardware, and the complete inference setup.
How can I access and deploy NVIDIA Parakeet TDT 0.6B v2?
The model is available on Hugging Face under the identifier nvidia/parakeet-tdt-0.6b-v2. It can be loaded locally through NVIDIA NeMo, and NVIDIA documentation also describes access through a hosted NVIDIA API using the Riva client. Local deployment requires compatible NVIDIA hardware and software, while hosted availability and behavior may vary.
What is NVIDIA Parakeet TDT 0.6B v2 used for?
NVIDIA Parakeet TDT 0.6B v2 is an open-weight automatic speech recognition model for fast English transcription. It can be used for meeting and interview transcripts, captions, media indexing, audio search, voice-interface preprocessing, and song-to-lyrics transcription.
Does NVIDIA Parakeet TDT 0.6B v2 support timestamps and streaming?
Yes. Through compatible NVIDIA NeMo transcription workflows, the model can provide character-level, word-level, and segment-level timestamps. NeMo also supports streaming-oriented and buffered inference, although latency, chunking, and timestamp behavior depend on the deployment configuration.


Sources 4
Provider

About NVIDIA AI