Tiny Aya

Tiny Aya Fire

by Cohere · Live

Cohere Tiny Aya Fire is a 3.35B instruction-tuned multilingual text model focused on South Asian languages. It covers 70 languages, supports 8K context and output limits, and is available through Cohere’s Chat endpoint or a gated CC-BY-NC-4.0 Hugging Face release. Its compact size favors translation, language-access applications, and local or edge deployment, while its reasoning, coding, multimodal, and commercial self-hosting limitations require careful evaluation.

Text Reasoning Coding
Cohere Tiny Aya Fire is a compact multilingual model built for language access rather than frontier-scale reasoning or coding. Its main distinction is regional specialization: it is designed to handle South Asian languages while retaining coverage across a broader 70-language set. With 3.35 billion parameters, an 8K context limit, and an 8K maximum output, it targets efficient cloud, local, edge, and potentially offline deployments. Developers can use it through Cohere’s Chat API or evaluate the gated open-weight release on Hugging Face, subject to its license and acceptable-use requirements.
Outputs

What Tiny Aya Fire can produce

Text
Inputs

What it can understand

Text
Model profile

Performance characteristics

3/10 Reasoning
3/10 Coding
8/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Tiny Aya
Model type Lightweight
Context window 8K tokens
Maximum output 8K tokens
Release date 2026-02-17
Status Live
Knowledge cutoff notes

No exact knowledge cutoff was stated in the current Cohere model documentation or Tiny Aya Fire model card reviewed.

Model notes

Tiny Aya Fire is the South Asia-focused instruction-tuned variant in the Tiny Aya family. The model identifier is tiny-aya-fire; the Hugging Face model card also identifies the underlying model as tiny-aya-it-fire. It has 3.35 billion parameters, supports 70 languages, and is documented with 8K input context and 8K maximum output. Cohere provides API access through the Chat endpoint, while the open-weight Hugging Face release is gated and licensed CC-BY-NC-4.0 with additional acceptable-use requirements. A model-specific public API price was not identified in the reviewed official pricing sources. Comparative scores are editorial estimates, not provider benchmarks.

Model guide

Tiny Aya Fire: A Compact Multilingual Model for South Asian Languages

Tiny Aya Fire is Cohere’s 3.35-billion-parameter, instruction-tuned multilingual language model focused on South Asian languages. It covers 70 languages, supports text generation and translation, provides an 8K context window and 8K maximum output, and is available through Cohere’s Chat endpoint as well as a gated open-weight Hugging Face release.

What is Tiny Aya Fire?

Tiny Aya Fire is an instruction-tuned multilingual language model from Cohere’s Tiny Aya family. Instruction tuning means the model has been adapted to follow natural-language requests, making it more suitable for conversation, translation, question answering, and other practical text-generation tasks than a purely pretrained language model.

The model is specifically positioned around South Asian language capability. It is intended to generate and understand text across languages such as Hindi, Marathi, Bengali, Gujarati, Punjabi, Tamil, Telugu, Nepali, and Urdu, while the broader model documentation describes coverage across 70 languages. This makes Tiny Aya Fire different from a small general-purpose model whose primary optimization target is English-language use.

Its 3.35 billion parameters place it in the lightweight segment of contemporary language models. That smaller size can make deployment and inference more manageable, but it also means users should not expect the same depth of reasoning, coding ability, or broad task performance associated with much larger frontier models.

Where Tiny Aya Fire fits in Cohere’s lineup

Tiny Aya Fire is the South Asia-focused instruction-tuned variant in Cohere’s Tiny Aya family. It sits alongside Cohere’s broader multilingual Aya work, but its practical purpose is narrower and more deployment-oriented: provide useful multilingual text generation with a relatively small model footprint.

Within Cohere’s wider catalog, Tiny Aya Fire is not a replacement for the company’s general Command models, retrieval models, or document-focused systems. It is a specialized language model for applications where multilingual coverage, lower resource requirements, and regional language performance matter more than maximum reasoning or coding capability.

Cohere lists the model as live and exposes it through the Chat endpoint using the model identifier tiny-aya-fire. The published model repository is CohereLabs/tiny-aya-fire. The Hugging Face release is gated, so access requires accepting the repository’s conditions rather than downloading the weights anonymously.

Verified specifications and supported modalities

The documented input and output modality is text. Tiny Aya Fire accepts text prompts and produces text responses; the supplied model information does not document native image, audio, or video input or output. It is therefore best evaluated as a multilingual text model, not as a multimodal assistant.

SpecificationDocumented value
ProviderCohere
Model familyTiny Aya
Parameters3.35 billion
Language coverage70 languages
Context length8,192 tokens
Maximum output8,192 tokens
InputText
OutputText
Model typeLightweight, instruction-tuned language model
Open-weight licenseCC-BY-NC-4.0, with additional acceptable-use requirements

Cohere’s documentation describes an autoregressive transformer architecture using sliding-window and global-attention layers. The model card also identifies supervised fine-tuning and preference training for helpfulness and safety. These details explain its basic operating design, but they do not constitute a guarantee of quality for every supported language or application.

Language strengths and practical uses

Tiny Aya Fire’s clearest advantage is its combination of regional focus and compact size. It may be a practical choice when an application needs text generation or translation in South Asian languages but cannot justify the infrastructure, latency, or cost of a much larger model.

Potential use cases include multilingual customer-support drafts, translation assistance, educational tools, language-access interfaces, regional content generation, and text classification or transformation workflows. It can also be considered for privacy-sensitive environments where a local model is preferable to sending every prompt to a hosted service.

For example, a team could use Tiny Aya Fire to create a first-pass translation between English and a supported South Asian language, generate a response in a user’s preferred language, or provide a lightweight multilingual interface for an offline or intermittently connected application. Human review remains important for legal, medical, financial, safety-critical, or culturally sensitive material.

The 8K context window is sufficient for many individual prompts, short conversations, and moderate documents. It is not unlimited: long reports, large collections of source material, or extended conversation histories may need to be split, summarized, or retrieved in sections before being sent to the model.

Reasoning, coding, and tool support

Tiny Aya Fire is not positioned as a frontier reasoning model. The supplied evaluation data gives it an editorial reasoning score of 3 out of 10, but that score is an internal comparative estimate rather than a provider-published benchmark. In practical terms, users should treat it as a model for multilingual language tasks, not as the default choice for complex mathematical reasoning, long chains of analysis, or difficult planning problems.

The same distinction applies to coding. Its editorial coding score is 3 out of 10, and the research does not identify it as a specialist programming model. It may still generate simple code snippets or help transform text containing code, but developers seeking advanced software generation, debugging, repository-level work, or code reasoning should consider a model designed and evaluated for those tasks.

Tool use, function calling, streaming, fine-tuning, caching, batch processing, and structured-output support are not verified in the supplied model record. They should not be assumed merely because the model is available through a chat endpoint. Applications that depend on any of these features should confirm them in the current Cohere documentation before implementation.

Speed, cost, and deployment trade-offs

The model’s relatively small parameter count is its main deployment advantage. Smaller models generally require fewer computing resources than large models, which can make them more attractive for local, edge, or cost-sensitive inference. The research assigns Tiny Aya Fire an editorial speed score of 8 out of 10 and a cost score of 9 out of 10. These are subjective estimates, not published Cohere performance or pricing benchmarks.

The open-weight release can be used with Transformers and compatible serving tools such as vLLM, SGLang, and GGUF-based runtimes, according to the supplied research. This gives technically capable teams more control over infrastructure and data handling. However, the CC-BY-NC-4.0 license is non-commercial, so organizations must review the license carefully before using the open-weight release in a commercial product or service.

Cohere also provides hosted access through its Chat endpoint. A model-specific public token price was not identified in the official pricing material reviewed. Because no verified input or output price is available here, the cost of API use should be confirmed directly with Cohere rather than estimated from the model’s small size.

Limitations to consider

Tiny Aya Fire’s specialization does not remove the usual quality differences between languages. The model is designed for 70 languages, but the research does not provide language-by-language benchmark results or establish that performance is equal across the entire set. Teams should test the exact languages, dialects, writing systems, and domains that matter to their users.

The model also has several important boundaries:

  • It produces text only and does not have documented native image, audio, or video capabilities.
  • Its 8K context and 8K maximum output limit may be restrictive for very long documents or extended conversations.
  • It is not intended to compete with frontier models on difficult reasoning or advanced coding.
  • The open-weight release is gated and licensed CC-BY-NC-4.0, which limits commercial self-hosting.
  • No exact knowledge cutoff was stated in the reviewed documentation.
  • Factual accuracy, safety, and language-specific quality should be validated before high-stakes deployment.

These limitations matter particularly when selecting between a compact specialist model and a larger general-purpose model. A larger option may be more appropriate when the task requires complex reasoning, reliable code generation, broad tool integration, or consistently strong performance across many unrelated domains.

When to choose Tiny Aya Fire

Choose Tiny Aya Fire when the central requirement is multilingual text generation with an emphasis on South Asian languages, and when a compact model is preferable to a larger general-purpose system. It is especially suitable for translation prototypes, multilingual assistants, regional education applications, language-access tools, and local or edge deployments where resource use and data control are important.

It is also a sensible candidate for teams that want to compare hosted inference with self-managed deployment. Cohere’s Chat endpoint offers a managed access path, while the gated open-weight release allows qualified users to explore local serving options. The two routes have different operational, legal, and cost implications, so they should not be treated as interchangeable.

Consider another model when the primary requirement is advanced reasoning, sophisticated programming, multimodal input, image or audio processing, extensive tool calling, or commercial self-hosting under a permissive license. Tiny Aya Fire’s value is its focused multilingual and lightweight profile; choosing it for tasks outside that profile may produce weaker results than choosing a model built for those capabilities.

Bottom line

Tiny Aya Fire is a focused 3.35-billion-parameter multilingual model rather than an all-purpose frontier assistant. Its strongest case is South Asian language work, especially when 70-language coverage, text generation, translation, and a relatively small deployment footprint are more important than maximum reasoning or coding performance. The 8K context and output limits make it usable for many everyday text workflows, while the hosted Cohere option and gated open-weight release provide two different access paths.

Before adopting it, verify quality in the specific languages and domains that matter, confirm any required API features with Cohere, and review the non-commercial open-weight license. For the right multilingual and resource-conscious application, those checks are more important than generalized model rankings.


Answers to Frequently Asked Questions

Can Tiny Aya Fire be self-hosted commercially?
The gated open-weight release can be used with tools such as Transformers, vLLM, SGLang, and GGUF-based runtimes, but it is licensed under CC-BY-NC-4.0 with additional acceptable-use requirements. Organizations should review the license carefully because it is non-commercial and may not permit use in a commercial product or service.
What is Tiny Aya Fire best used for?
The model is well suited to multilingual customer-support drafts, translation assistance, educational tools, regional content generation, language-access interfaces, text transformation, and lightweight local or edge deployments. Human review is recommended for legal, medical, financial, safety-critical, and culturally sensitive content.
What are Tiny Aya Fire’s main specifications and capabilities?
Tiny Aya Fire accepts text and produces text, has an 8,192-token context length and an 8,192-token maximum output, and uses an autoregressive transformer architecture. It does not have documented native image, audio, or video capabilities.
What is Tiny Aya Fire?
Tiny Aya Fire is a 3.35-billion-parameter, instruction-tuned multilingual language model from Cohere’s Tiny Aya family. It is designed for text generation, translation, question answering, and other language tasks, with a particular focus on South Asian languages.
Which languages does Tiny Aya Fire support?
Tiny Aya Fire is documented as supporting 70 languages, including Hindi, Marathi, Bengali, Gujarati, Punjabi, Tamil, Telugu, Nepali, and Urdu. Performance may vary by language, dialect, writing system, and domain, so users should test the specific languages they need.


Sources 7
Provider

About Cohere