Canary

Canary-1B

by NVIDIA AI · Active; open-weight checkpoint and NVIDIA deployment options available

NVIDIA Canary-1B is an approximately 1-billion-parameter encoder-decoder model that converts speech into text for multilingual transcription and speech translation. It supports local NeMo use and GPU-accelerated Speech NIM deployment, but the documented NIM path is offline-only and the released checkpoint uses a CC BY-NC 4.0 license.

Text Reasoning Coding
NVIDIA Canary-1B is built for speech transcription and translation rather than general-purpose chat, coding, or text generation. Its Fast-Conformer encoder and autoregressive decoder support multilingual audio-to-text workloads, including offline transcription and translation between English and supported languages. The released checkpoint is available under the CC BY-NC 4.0 license, making licensing review essential for commercial use.
Outputs

What Canary-1B can produce

Text
Inputs

What it can understand

Audio
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
6/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Canary
Model type Other
Release date 2024-04-18
Status Active; open-weight checkpoint and NVIDIA deployment options available
Knowledge cutoff notes

No authoritative model-specific knowledge cutoff was identified. This is a speech recognition and translation model rather than a conventional knowledge-grounded language model.

Model notes

Canonical public model identifier is nvidia/canary-1b. The model is an approximately 1-billion-parameter Fast-Conformer encoder-decoder model for ASR and automatic speech translation. NVIDIA's original release described English, German, Spanish, and French transcription plus bidirectional translation involving English. Current NVIDIA Speech NIM documentation lists Canary 1b Multilingual for 26 transcription language variants and describes the NIM deployment as offline-only. The released Hugging Face checkpoint is licensed CC BY-NC 4.0, so commercial use and redistribution require careful license review. Editorial scores are comparative estimates for a specialized ASR model, not vendor specifications.

Model guide

NVIDIA Canary-1B: Open Multilingual Speech Recognition and Translation

NVIDIA Canary-1B is an approximately 1-billion-parameter open-weight encoder-decoder speech model for multilingual automatic speech recognition and bidirectional speech translation. It accepts speech audio and produces text, with local deployment options through NVIDIA NeMo and GPU-accelerated deployment through NVIDIA Speech NIM.

What is NVIDIA Canary-1B?

NVIDIA Canary-1B is a multilingual automatic speech recognition (ASR) and speech translation model from NVIDIA's NeMo team. Its canonical public identifier is nvidia/canary-1b. The model contains approximately 1 billion parameters and uses an encoder-decoder design: an audio encoder processes speech features, while a decoder generates the resulting text.

In practical terms, Canary-1B turns spoken audio into written language. Depending on the task and deployment configuration, that output can be a transcription in the source language or a translation into another supported language. It is not a general-purpose conversational model, and it should not be evaluated as an alternative to a text-focused large language model.

The model was released on April 18, 2024, and remains available as an open-weight checkpoint alongside NVIDIA deployment options. Current NVIDIA documentation places Canary within its speech model and Speech NIM deployment catalog.

Primary purpose and positioning

Canary-1B is intended for applications where the input is speech and the required result is text. Typical examples include multilingual meeting transcription, recorded-audio indexing, research into speech recognition, and batch translation of spoken content.

Its main distinction is the combination of multilingual ASR and speech translation in one encoder-decoder model. The original release described transcription in English, German, Spanish, and French, together with bidirectional translation involving English. More recent NVIDIA Speech NIM documentation lists Canary 1b Multilingual with 26 supported transcription language variants. Because language coverage depends on the specific release and deployment path, users should check the current support matrix before building a production pipeline.

Canary is therefore best understood as a specialized speech model, not as a smaller general language model. It does not natively produce audio, images, video, embeddings, or executable actions.

Architecture and supported inputs and outputs

The model uses a Fast-Conformer encoder to analyze log-mel spectrogram features derived from speech audio. A log-mel spectrogram is a compact representation of how sound energy changes across frequencies over time. The decoder then generates text tokens autoregressively, meaning it produces the output sequence step by step.

Special task and language tokens indicate whether the system should perform automatic speech recognition or automatic speech translation. The resulting output is text with punctuation and capitalization support.

CapabilityCanary-1B support
Audio inputYes; speech audio is the model's primary input
Text inputNot identified as a supported model input
Text outputYes; transcription or translation
Audio outputNo
Image or video inputNo
Tool or function callingNo
Structured JSON outputNo documented native support

There is no documented context-window size or maximum output-token limit for this checkpoint in the supplied NVIDIA materials. These limits should not be inferred from the model's parameter count. In an operational system, practical audio-duration limits may instead be determined by the selected framework, GPU memory, batching configuration, and deployment container.

Deployment through NeMo and Speech NIM

Local users can load the nvidia/canary-1b checkpoint through NVIDIA NeMo for research, experimentation, and custom speech pipelines. NVIDIA also documents a Canary-1B deployment through NVIDIA Speech NIM, a GPU-accelerated packaging and serving option for speech models.

The documented Canary NIM deployment is described as offline-only. That distinction matters: offline processing can be appropriate for recorded calls, uploaded media, archives, or batch translation, but it should not automatically be treated as a low-latency streaming transcription service. The model's autoregressive decoder may also require more computation than simpler CTC or transducer-based ASR systems, particularly when response latency is the primary requirement.

Efficient inference generally requires suitable NVIDIA GPU infrastructure. The supplied research does not identify a universal minimum GPU, memory requirement, throughput figure, or deployment cost for every NeMo or NIM configuration, so those values should be validated against the target hardware and current NVIDIA documentation.

Main strengths

  • Multilingual speech processing: Canary-1B supports transcription across a broad set of language variants in current Speech NIM documentation, while the original release also included speech translation involving English.
  • Combined ASR and translation tasks: The encoder-decoder design supports both transcription and speech-to-text translation rather than only source-language transcription.
  • Open-weight access: The checkpoint can be used with NVIDIA NeMo, giving researchers and developers more control than a hosted-only transcription service.
  • Readable transcription output: Punctuation and capitalization support make the output more useful for transcripts and downstream text processing.
  • NVIDIA deployment integration: NeMo and Speech NIM provide documented paths for local or GPU-accelerated deployment within NVIDIA's speech ecosystem.

NVIDIA presents Canary-1B as a high-accuracy model among similarly sized open speech models. That is a provider positioning claim rather than an independent conclusion. Actual accuracy will vary with language, accents, recording conditions, background noise, microphone quality, and the chosen deployment version.

Limitations and trade-offs

The model's specialization is also its main limitation. Canary-1B does not provide general text reasoning, coding, chat, document analysis, web search, retrieval, speech synthesis, or speech-to-speech generation. A system that needs those capabilities would require additional models or a different model type.

Latency is another consideration. Canary-1B's autoregressive decoder can support strong transcription and translation quality, but it generally involves more sequential computation than some lower-latency ASR architectures. The documented Canary NIM path is offline-only, so a streaming-focused application may be better served by a model and serving stack explicitly designed for real-time incremental transcription.

Language support is not identical across every release. The original checkpoint documentation and current Speech NIM support matrix describe different language scopes, so developers should confirm that both the required source language and translation direction are supported by the exact version they plan to deploy.

Finally, the released Hugging Face checkpoint uses the CC BY-NC 4.0 license. This permits research and non-commercial use subject to the license terms, but it should not be assumed to permit commercial redistribution or commercial production use. Organizations should review the license and any separate NVIDIA deployment terms before adoption.

Pricing, licensing, and availability

No per-minute, per-request, or subscription price is identified for the Canary-1B checkpoint in the supplied research. The open-weight model itself is available through NVIDIA's model distribution channels, but using it locally still carries infrastructure, storage, engineering, and GPU operating costs.

NVIDIA Speech NIM may involve separate software, infrastructure, entitlement, or enterprise terms. No universal Canary-1B NIM price is established here, so pricing should be confirmed with NVIDIA for the intended deployment. The absence of a listed model price does not mean that a production NIM deployment is cost-free.

The most important availability condition is licensing: the released checkpoint's CC BY-NC 4.0 terms make it a better fit for research, evaluation, and non-commercial prototyping unless an organization confirms an appropriate commercial arrangement.

Reasoning, coding, and tool capabilities

Canary-1B is not designed for reasoning or code generation. It converts audio into text according to a speech recognition or translation task; it does not offer a general instruction-following interface for multi-step analysis.

There is no documented native tool calling, function calling, web browsing, code execution, file analysis, or structured-output mode. Developers can place the model inside a larger application—for example, transcribing a recording before sending the text to another system—but those surrounding capabilities do not belong to Canary-1B itself.

Best use cases

  • Offline transcription of multilingual recordings.
  • Speech translation between English and supported languages.
  • Research and prototyping with NVIDIA NeMo ASR models.
  • GPU-accelerated batch processing of interviews, meetings, lectures, or media archives.
  • Enterprise speech pipelines where local or controlled deployment is more important than a hosted API workflow.

When to choose Canary-1B

Choose Canary-1B when the central problem is converting multilingual speech into text, especially when you want an open-weight checkpoint and are comfortable operating NVIDIA-oriented GPU infrastructure. It is particularly suitable when offline processing is acceptable and transcription or translation quality matters more than the lowest possible latency.

Consider another option when you need real-time streaming, a simple hosted per-minute service, commercial-friendly licensing, native speech synthesis, or a general-purpose assistant. A streaming-specialized ASR model may provide a better latency trade-off, while a text language model may be more appropriate after transcription if the application requires summarization, reasoning, coding, or tool use. Those alternatives solve different problems; they are not direct replacements for Canary-1B's multilingual audio-to-text role.

Bottom line

NVIDIA Canary-1B is a focused, open-weight model for multilingual speech recognition and speech translation. Its Fast-Conformer encoder, text-generating decoder, NeMo integration, and Speech NIM deployment path make it useful for quality-oriented offline speech workflows. It is less suitable for streaming applications, general-purpose language tasks, or commercial deployments that cannot accommodate the CC BY-NC 4.0 license. The most important evaluation questions are whether the required language and translation direction are supported, whether offline GPU inference fits the application, and whether the licensing terms match the intended use.


Answers to Frequently Asked Questions

Does NVIDIA Canary-1B support text generation, tool calling, or speech synthesis?
No. Canary-1B is specialized for converting speech audio into transcription or translation text. It does not natively provide general reasoning, coding, tool or function calling, web browsing, structured JSON output, speech synthesis, or speech-to-speech generation.
What are the licensing terms for NVIDIA Canary-1B?
The released Hugging Face checkpoint uses the CC BY-NC 4.0 license, which supports research and non-commercial use subject to its terms. Commercial users should review the license and any separate NVIDIA deployment terms before using Canary-1B in production.
Can NVIDIA Canary-1B perform real-time streaming transcription?
Canary-1B is better suited to offline processing than real-time streaming. The documented Canary Speech NIM deployment is offline-only, and its autoregressive decoder may require more sequential computation than streaming-focused ASR architectures.
Which languages does NVIDIA Canary-1B support?
Language coverage depends on the specific release and deployment path. The original release supported transcription in English, German, Spanish, and French, with bidirectional translation involving English. Current NVIDIA Speech NIM documentation lists Canary 1b Multilingual with 26 supported transcription language variants, so users should verify the current support matrix.
What is NVIDIA Canary-1B used for?
NVIDIA Canary-1B is an open-weight multilingual speech model used to transcribe spoken audio into text and translate speech into another supported language. Common applications include offline meeting transcription, recorded-audio indexing, research, and batch speech translation.


Sources 6
Provider

About NVIDIA AI