Falcon-H1-Tiny

Falcon-H1-Tiny-Tool-Calling-90M

by Technology Innovation Institute (TII) · Current downloadable open-weight model

A compact English Falcon-H1-Tiny model specialized for lightweight function calling and local automation. It has approximately 90 million parameters, a 262,144-token configured context window, a hybrid Transformer–Mamba architecture, and downloadable weights for use with several open-source inference runtimes.

Text Reasoning Coding
Falcon-H1-Tiny-Tool-Calling-90M is a specialized member of TII's Falcon-H1-Tiny family. It is an English text-generation model with approximately 91.1 million parameters and a built-in chat template for producing function calls in structured JSON enclosed by tool-call tags. The model is distributed as downloadable weights under the Falcon-LLM License and can be run with Transformers, vLLM, SGLang, llama.cpp, Ollama, or MLX.
Outputs

What Falcon-H1-Tiny-Tool-Calling-90M can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Tool use Streaming
Model profile

Performance characteristics

2/10 Reasoning
2/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Falcon-H1-Tiny
Model type Lightweight
Context window 262K tokens
Release date 2026-01-15
Status Current downloadable open-weight model
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was published in the verified model card or configuration.

Model notes

The official Hugging Face model card identifies this as an English causal decoder-only model with approximately 90 million parameters, a hybrid Transformer and Mamba architecture, and the Falcon-LLM License. Its tokenizer chat template is specialized for function calling and instructs the model to emit JSON arrays inside <tool_call> tags. This is structured text generated by the model, not a separately documented provider-managed JSON-schema or hosted function-calling API. The model can be run locally with Transformers, vLLM, SGLang, llama.cpp, Ollama, or MLX. Hugging Face currently indicates that the model is not deployed by an Inference Provider. The release date uses the official repository's January 15, 2026 collection/update date; a more precise launch timestamp was not verified. Editorial scores reflect the model's tiny parameter count and specialization rather than an official benchmark rating.

Model guide

Falcon-H1-Tiny-Tool-Calling-90M: A Tiny Local Model for Function Calling

Falcon-H1-Tiny-Tool-Calling-90M is a 90-million-parameter English causal language model from the Technology Innovation Institute, designed for lightweight text generation and function-calling workflows. Its hybrid Transformer–Mamba architecture, 262,144-token configured context window, downloadable weights, and tool-call chat template make it suitable for local, edge, and resource-constrained applications.

What is Falcon-H1-Tiny-Tool-Calling-90M?

Falcon-H1-Tiny-Tool-Calling-90M is a compact open-weight language model developed by the Technology Innovation Institute (TII). It belongs to TII's Falcon-H1-Tiny family and is designed for text generation with a particular emphasis on function calling, sometimes called tool calling.

In a tool-calling workflow, a language model does not directly perform an outside action. Instead, it produces a structured request naming a function and supplying its arguments. The surrounding application can then validate that request and decide whether to call an API, query a database, control a device, or run another local operation. This makes the model useful for small automation systems without requiring a large hosted model.

The model has approximately 90 million parameters; the official repository reports 91.1 million parameters. That is very small compared with most general-purpose language models. The trade-off is lower expected capability for broad knowledge, complex reasoning, sophisticated coding, and long multi-step planning. Its value is primarily in low-resource deployment and narrowly defined tasks.

Position in the Falcon family

TII presents Falcon-H1-Tiny-Tool-Calling-90M as part of the Falcon-H1-Tiny collection. It is not a consumer chatbot subscription or a provider-managed agent platform. The model is distributed as downloadable weights through its official Hugging Face repository, allowing users to run it on their own infrastructure or through compatible third-party services.

The tool-calling variant is more specialized than a small general text model because its tokenizer configuration includes a chat template for formatting tools and function calls. That specialization can make it a more suitable starting point for simple local agents than a similarly small model without a documented tool-call format. However, the supplied research does not establish that it outperforms other small models on a standardized benchmark.

Verified specifications

SpecificationVerified detail
ProviderTechnology Innovation Institute
Model familyFalcon-H1-Tiny
ParametersApproximately 90 million; 91.1 million reported in the repository
ArchitectureHybrid Transformer and Mamba causal language model
Primary languageEnglish
Configured context length262,144 tokens
Model configuration24 hidden layers, 512 hidden size, 32,768-token vocabulary
WeightsSafetensors, BF16
LicenseFalcon-LLM License
Output typeText only

The 262,144-token figure is the configured maximum context length reported by the model configuration. It should not be interpreted as a guarantee that every runtime, hardware setup, or application can process that many tokens efficiently. The supplied research does not specify a maximum generated-output token count.

How its tool calling works

The model's documented chat template instructs it to place function-call data inside <tool_call> and </tool_call> tags. The content is a JSON array containing a function name and an arguments object. A simplified example has this shape:

<tool_call>
[
  {"name": "function_name", "arguments": {"arg1": "value"}}
]
</tool_call>

This format gives an application a predictable place to look for a tool request, but it is still model-generated text. The supplied documentation does not describe a separate provider-managed function-calling service, hosted agent runtime, or guaranteed JSON-schema enforcement. It also does not establish a distinct JSON mode. Applications should parse the output, validate the JSON, check the requested function and arguments against an allowlist, and handle malformed or unsafe requests.

The model can also produce an ordinary text response when no tool call is needed. A simple application could therefore ask it to classify a request, call a local function when appropriate, return the function result to the model, and request a final response. Because the model is very small, developers should keep tool descriptions concise and workflows narrow rather than expecting reliable long-horizon planning.

Deployment and availability

Falcon-H1-Tiny-Tool-Calling-90M is intended for local or self-managed use. The official model card documents compatibility with Transformers, vLLM, SGLang, llama.cpp, Ollama, and Apple's MLX ecosystem. A separate GGUF repository provides quantized variants for runtimes that support that format.

Hugging Face currently indicates that the model is not deployed by an Inference Provider. As a result, users should not assume that an official hosted endpoint is available for immediate API calls. Typical deployment involves downloading the model files, selecting a compatible runtime, and providing the application layer that manages prompts, tool definitions, validation, execution, and error handling.

The model is distributed as downloadable weights rather than as a documented paid subscription product. No official per-input-token, per-output-token, or hosted endpoint price is identified in the supplied research. Download availability does not remove the need to review the Falcon-LLM License, hardware costs, hosting costs, and any restrictions that may apply to redistribution or commercial deployment.

Modalities and capability profile

This is an English text-in, text-out causal language model. It does not natively accept images, audio, or video, and it does not generate images, audio, or video. Files can only be handled if an application first converts their contents into text; the model itself is not documented as a vision, speech, OCR, or general multimodal model.

Tool use is the model's defining capability. Its documented template supports generating function names and argument objects, which is useful for routing requests to simple APIs, utilities, local software functions, or device controls. Tool calls should remain under the host application's control. The model should not be allowed to execute arbitrary commands or access sensitive systems solely because it generated a plausible-looking function request.

Reasoning and coding are possible in the general sense that the model generates text, but the supplied research does not report specialist reasoning or coding benchmarks. Editorial evaluations rate its reasoning and coding suitability as limited relative to larger models. Those are subjective assessments based on its small parameter count and intended role, not provider-published scores. It should be treated as a lightweight specialist rather than a general reasoning or advanced programming model.

Strengths and trade-offs

  • Small resource requirement: Approximately 90 million parameters make the model substantially easier to store and run than billion-parameter alternatives, although actual memory use depends on the runtime, precision, and context.
  • Tool-oriented design: The included chat template provides a documented convention for emitting function calls.
  • Local control: Downloadable weights allow deployment on local, edge, or private infrastructure instead of requiring a specific official hosted endpoint.
  • Runtime flexibility: The documented ecosystem includes several common open-source inference runtimes, with quantized options available through a separate GGUF repository.
  • Long configured context: The configuration specifies 262,144 tokens, although practical performance at that length depends on hardware and software.
  • Limited general intelligence: The tiny parameter count makes it a poor choice for difficult research, broad knowledge work, nuanced writing, advanced coding, or complex agent planning.
  • English focus: The official model description identifies English as its primary language, so it is not the best-supported choice for multilingual workloads based on the supplied information.
  • License and operations remain the user's responsibility: Self-hosting requires attention to licensing, security, monitoring, validation, and hardware or hosting costs.

When to choose this model

Choose Falcon-H1-Tiny-Tool-Calling-90M when the main requirement is a small, fast, inexpensive-to-operate text model that can produce structured requests for a limited set of tools. Appropriate examples include:

  • Routing short user requests to a small collection of known functions.
  • Extracting API arguments from simple natural-language commands.
  • Operating lightweight offline assistants on constrained hardware.
  • Controlling local utilities or edge devices through an allowlisted function layer.
  • Building prototypes where downloading and modifying open model weights is preferable to depending on a hosted API.

Its speed and cost advantages are practical trade-offs rather than guarantees of a particular latency. A smaller model generally requires fewer resources than a large model, but measured performance will depend on hardware, quantization, batch size, context length, and runtime. The supplied research does not provide a benchmark latency or throughput figure.

When another option may be better

Use a larger language model when the task requires reliable multi-step reasoning, complex code generation, broad factual coverage, sophisticated planning, or robust interpretation of ambiguous instructions. A larger tool-capable model may also be preferable when the application has many tools, complicated schemas, or costly consequences for selecting the wrong function.

Use a multimodal model when users need to submit images, audio, video, or scanned documents directly. Falcon-H1-Tiny-Tool-Calling-90M is text-only and does not replace TII's separate perception or multimodal model work. Use a hosted service instead when operational simplicity, managed scaling, usage monitoring, and a provider-supported endpoint matter more than local control.

For any deployment, the surrounding application is as important as the model. Keep the available tools narrow, validate every argument, enforce permissions independently of model output, set timeouts, log failures safely, and provide a fallback when the model emits invalid JSON or an unsuitable function call.

Bottom line

Falcon-H1-Tiny-Tool-Calling-90M is best understood as a compact building block for local function-calling automation, not as a general-purpose conversational assistant. Its approximately 90-million-parameter footprint, hybrid Transformer–Mamba design, very long configured context, and documented tool-call template make it attractive for constrained deployments. Its limitations—English focus, text-only operation, no documented hosted inference provider, unspecified output limit, and modest expected reasoning and coding ability—are equally important. It is a sensible choice when low resource use and local control come first, but larger or hosted models are more appropriate for demanding, broad, or high-stakes work.


Answers to Frequently Asked Questions

What are the main limitations of Falcon-H1-Tiny-Tool-Calling-90M?
It is an English-focused, text-only model with limited expected performance for complex reasoning, advanced coding, broad factual tasks, and long-horizon planning. It does not natively process images, audio, or video, and there is no documented official hosted inference endpoint or guaranteed JSON-schema enforcement.
Can Falcon-H1-Tiny-Tool-Calling-90M run locally?
Yes. The model is distributed as downloadable weights for local or self-managed deployment and is documented as compatible with Transformers, vLLM, SGLang, llama.cpp, Ollama, and Apple's MLX ecosystem. Quantized GGUF variants are also available through a separate repository.
What is Falcon-H1-Tiny-Tool-Calling-90M?
Falcon-H1-Tiny-Tool-Calling-90M is a compact open-weight language model from the Technology Innovation Institute (TII), designed primarily for local text generation and function calling. It has approximately 90 million parameters, with 91.1 million reported in the official repository.
How does Falcon-H1-Tiny-Tool-Calling-90M perform function or tool calling?
The model generates function-call data inside and tags as a JSON array containing a function name and arguments. The host application must parse and validate this output, check the function and arguments against an allowlist, and execute the approved operation.


Sources 4
Provider

About Technology Innovation Institute (TII)