Developer platform

TII Falcon Developer Platform

Developer overview for Technology Innovation Institute (TII), including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API Amazon Bedrock Runtime Converse API for managed Falcon deployments; self-hosted inference APIs for locally deployed Falcon models
SDK support No verified first-party TII SDK. Managed integrations can use AWS SDKs such as boto3 and AWS SDK for JavaScript, or compatible HTTP clients.
Rate limits No universal TII rate limit was verified. Managed limits depend on the host; Amazon Bedrock applies account, region, model, and throughput quotas, while SageMaker depends on endpoint capacity and scaling.
Platform

API overview

Endpoints

API access

Primary API Amazon Bedrock Runtime Converse API for managed Falcon deployments; self-hosted inference APIs for locally deployed Falcon models
Pricing

API pricing

Pricing model Open-weight and self-hosted deployment; managed inference through partner cloud platforms using their billing and authentication

No universal TII-hosted API price was verified. Falcon model access may be free under the applicable license, while Amazon Bedrock and SageMaker usage is billed by AWS.

Developer experience

SDKs & usability

SDKs No verified first-party TII SDK. Managed integrations can use AWS SDKs such as boto3 and AWS SDK for JavaScript, or compatible HTTP clients.
Ease of use Moderate. Self-hosting requires model-serving and GPU expertise; managed AWS deployment is easier but requires AWS configuration and model-specific setup.
Documentation Moderate. TII provides model pages, licenses, research, and release information, while detailed API operations are documented by the selected hosting provider.
Latency Deployment-dependent. Latency varies with model size, region, GPU or managed endpoint capacity, queueing, prompt length, and serving configuration.
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Fine-tuning
✓ Image input
✓ Playground
Feature notes

TII does not appear to publish a universal first-party hosted inference endpoint or API-key service. The current developer offering consists of official Falcon model releases, model repositories, licenses, and deployment through third-party infrastructure. Amazon Bedrock Marketplace and SageMaker JumpStart provide managed access to selected Falcon models. Streaming is supported by the Bedrock ConverseStream API when supported by the selected model and deployment. Function calling is available for some newer TII/Falcon model variants or serving configurations, but is not universal across the catalog. Image input is available for selected multimodal TII models such as Falcon Perception, not for all Falcon text models. Structured output, file upload, web search, and persistent assistants are hosting-layer or model-specific features and should be verified for the selected deployment. TII does not provide a verified first-party persistent assistant or agent runtime.

Examples

API examples

# Requires an Amazon Bedrock API key with access to a compatible OpenAI-style Bedrock endpoint.
# Set AWS_REGION, BEDROCK_API_KEY, and FALCON_MODEL_ID before running.

curl --fail-with-body "https://bedrock-runtime.${AWS_REGION}.amazonaws.com/openai/v1/chat/completions" \
  -H "Authorization: Bearer ${BEDROCK_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "'"${FALCON_MODEL_ID}"'",
    "messages": [
      {
        "role": "system",
        "content": "You are a concise technical assistant."
      },
      {
        "role": "user",
        "content": "Explain what Falcon models are in two sentences."
      }
    ],
    "temperature": 0.2,
    "max_tokens": 200
  }' | jq -r '.choices[0].message.content'
# Install: pip install boto3

import os
import time
import boto3
from botocore.exceptions import ClientError

region = os.environ.get("AWS_REGION", "us-east-1")
model_id = os.environ["FALCON_MODEL_ID"]

client = boto3.client("bedrock-runtime", region_name=region)

messages = [
    {
        "role": "user",
        "content": [
            {
                "text": "Give a short explanation of Falcon models."
            }
        ],
    }
]

try:
    response = client.converse(
        modelId=model_id,
        system=[
            {
                "text": "You are a precise technical assistant."
            }
        ],
        messages=messages,
        inferenceConfig={
            "temperature": 0.2,
            "maxTokens": 200,
        },
    )
    text = response["output"]["message"]["content"][0]["text"]
    print(text)
except ClientError as error:
    code = error.response.get("Error", {}).get("Code", "Unknown")
    message = error.response.get("Error", {}).get("Message", str(error))
    if code in {"ThrottlingException", "ServiceUnavailableException", "ModelTimeoutException"}:
        time.sleep(2)
        raise RuntimeError(f"Transient Bedrock error; retry the request: {message}") from error
    raise RuntimeError(f"Bedrock request failed: {message}") from error

# Advanced example: streaming output through the Bedrock ConverseStream API.
stream = client.converse_stream(
    modelId=model_id,
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "text": "List three practical Falcon deployment options."
                }
            ],
        }
    ],
    inferenceConfig={"maxTokens": 250},
)

for event in stream.get("stream", []):
    delta = event.get("contentBlockDelta", {}).get("delta", {})
    if "text" in delta:
        print(delta["text"], end="", flush=True)
print()
// Install: npm install @aws-sdk/client-bedrock-runtime

import {
  BedrockRuntimeClient,
  ConverseCommand,
  ConverseStreamCommand,
} from "@aws-sdk/client-bedrock-runtime";

const region = process.env.AWS_REGION || "us-east-1";
const modelId = process.env.FALCON_MODEL_ID;

if (!modelId) {
  throw new Error("Set FALCON_MODEL_ID to the selected Falcon deployment identifier.");
}

const client = new BedrockRuntimeClient({ region });

try {
  const response = await client.send(
    new ConverseCommand({
      modelId,
      system: [{ text: "You are a concise technical assistant." }],
      messages: [
        {
          role: "user",
          content: [{ text: "Explain Falcon models in two sentences." }],
        },
      ],
      inferenceConfig: {
        temperature: 0.2,
        maxTokens: 200,
      },
    }),
  );

  const text = response.output?.message?.content
    ?.map((item) => item.text || "")
    .join("");
  console.log(text || "No text returned");
} catch (error) {
  console.error("Falcon request failed:", error instanceof Error ? error.message : error);
  process.exitCode = 1;
}

const stream = await client.send(
  new ConverseStreamCommand({
    modelId,
    messages: [
      {
        role: "user",
        content: [{ text: "List three Falcon deployment options." }],
      },
    ],
    inferenceConfig: { maxTokens: 250 },
  }),
);

for await (const event of stream.stream || []) {
  const text = event.contentBlockDelta?.delta?.text;
  if (text) process.stdout.write(text);
}
process.stdout.write("\n");
<?php

$region = getenv('AWS_REGION') ?: 'us-east-1';
$apiKey = getenv('BEDROCK_API_KEY');
$modelId = getenv('FALCON_MODEL_ID');

if (!$apiKey || !$modelId) {
    fwrite(STDERR, "Set BEDROCK_API_KEY and FALCON_MODEL_ID.\n");
    exit(1);
}

$url = "https://bedrock-runtime.{$region}.amazonaws.com/openai/v1/chat/completions";
$payload = [
    'model' => $modelId,
    'messages' => [
        [
            'role' => 'system',
            'content' => 'You are a concise technical assistant.'
        ],
        [
            'role' => 'user',
            'content' => 'Explain Falcon models in two sentences.'
        ]
    ],
    'temperature' => 0.2,
    'max_tokens' => 200
];

$ch = curl_init($url);
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $apiKey,
        'Content-Type: application/json',
        'Accept: application/json'
    ],
    CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR),
    CURLOPT_TIMEOUT => 120
]);

$body = curl_exec($ch);
if ($body === false) {
    $error = curl_error($ch);
    curl_close($ch);
    throw new RuntimeException('Network error: ' . $error);
}

$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
curl_close($ch);

$data = json_decode($body, true);
if ($status < 200 || $status >= 300) {
    $message = $data['error']['message'] ?? $body;
    throw new RuntimeException("HTTP {$status}: {$message}");
}

if (!isset($data['choices'][0]['message']['content'])) {
    throw new RuntimeException('The response did not contain assistant text.');
}

echo $data['choices'][0]['message']['content'] . PHP_EOL;
Policies

Data & usage

Data training

TII's open-model releases are distributed under model-specific licenses and terms. For managed inference, submitted prompts and outputs are governed primarily by the selected hosting provider's data-processing and service terms. No single TII-wide API training policy for a universal hosted endpoint was verified. Organizations should review the exact TII model license and the provider policy for Amazon Bedrock, SageMaker, Hugging Face, or another selected host.

Data retention

No universal TII-managed API retention policy was verified because TII does not appear to operate a single public hosted inference API. Retention, logging, storage, and regional processing depend on whether the model is self-hosted or invoked through a provider such as Amazon Bedrock or SageMaker. Self-hosting allows the operator to configure retention and logging controls.

Rate limits

No universal TII rate limit was verified. Managed limits depend on the host; Amazon Bedrock applies account, region, model, and throughput quotas, while SageMaker depends on endpoint capacity and scaling.

Developer guide

TII Falcon Developer API: Access, Models, and Deployment Guide

TII's Falcon developer platform is primarily an open-model and deployment ecosystem, not a universal first-party hosted API. Developers can download Falcon weights and self-host them, use official model repositories, or access selected Falcon deployments through partner services such as Amazon Bedrock Marketplace and SageMaker JumpStart. Authentication, billing, quotas, endpoints, and data handling depend on the selected hosting method.
TII Falcon is aimed at developers and organizations that want to deploy open or open-access language and multimodal models rather than consume a single closed chatbot API. The practical options are self-hosting official Falcon releases or using managed cloud infrastructure, especially Amazon Bedrock and SageMaker. This guide explains how those options differ, how to make a first managed request, which capabilities vary by model, and what to verify before production use.
TII Falcon is primarily an open-model developer ecosystem rather than a universal first-party hosted API. Developers can self-host official Falcon releases or use selected models through Amazon Bedrock Marketplace and SageMaker JumpStart. The guide covers access, AWS authentication, model selection, Python requests, streaming, pricing, tool calling, files, structured outputs, SDKs, quotas, and production trade-offs.

What the TII Falcon developer API is

Technology Innovation Institute (TII) does not appear to provide a universal, first-party hosted inference API with one TII base URL, one API key, and one account-wide quota system. Instead, its developer offering is centered on the Falcon family of open and open-access models, official model repositories, model licenses, research releases, and deployment guidance.

There are two main ways to use Falcon in an application:

  • Self-hosting: download a permitted model from TII's official repositories and serve it with an inference framework such as Transformers, vLLM, or another compatible server.
  • Managed hosting: deploy or invoke a supported Falcon model through a cloud provider. Amazon Bedrock Marketplace and Amazon SageMaker JumpStart are the clearest documented managed routes in the supplied research.

This distinction is important. TII supplies the model and applicable license, while the selected host normally supplies the endpoint, authentication, quotas, billing, scaling, and operational controls.

Who should use Falcon

Falcon is most suitable for developers, researchers, and organizations that need control over model deployment, data location, infrastructure, or model selection. It is also relevant when a team wants to evaluate multilingual, Arabic-focused, efficient, edge-oriented, or multimodal models.

It is a less direct choice for a beginner who expects a mature consumer-style API with a single subscription, universal feature set, persistent assistants, built-in web research, and guaranteed hosted availability. Those features are not universal across the Falcon ecosystem and may not be provided by TII itself.

Getting access and obtaining credentials

For self-hosting, start with the exact Falcon model repository and model card, then review its license, hardware requirements, supported software, and usage restrictions. Downloading model weights does not create a TII API account or provide a TII API key.

For managed AWS access, create or use an AWS account, select a supported Falcon listing in Amazon Bedrock Marketplace or SageMaker JumpStart, and confirm that the model is available in the required AWS Region. AWS credentials, permissions, account quotas, and any marketplace or service terms apply. The model identifier is deployment-specific and should be copied from the AWS console or listing rather than hard-coded from an unrelated example.

Do not assume that a TII account, the Falcon web application, or a downloaded model supplies credentials for Bedrock or another hosting provider. There is no verified universal TII API key or TII-managed base URL in the current materials.

Choosing the model and API path

Choose the serving method first, then verify the capabilities of the exact Falcon model. Text-oriented releases are different from multimodal models such as Falcon Perception, and support can also change between a model and its serving layer.

RequirementLikely pathWhat to verify
Managed conversational inferenceAmazon Bedrock RuntimeModel ID, Region, Converse support, quotas, and AWS pricing
Custom data control or private deploymentSelf-hosted inferenceLicense, GPU capacity, serving framework, security, and scaling
Managed endpoint deploymentAmazon SageMaker JumpStartInstance type, endpoint capacity, scaling, and deployment cost
Vision or OCR workloadsA supported multimodal Falcon releaseImage format, input schema, model availability, and host support

Falcon-H1, Falcon-H1-Tiny, Falcon 3, Falcon Perception, Falcon Arabic, Falcon Mamba, and other TII releases do not necessarily expose the same inputs or output features. Confirm whether the selected deployment supports image input, audio or video analysis, tool calling, streaming, and the required message format.

Making a first request with Amazon Bedrock

For a managed AWS integration, the current recommended pattern is the Bedrock Runtime Converse operation. It provides a common message-oriented interface for supported models, while AWS handles authentication and the managed endpoint. The exact Falcon model ID must be supplied by the selected AWS deployment.

Install the AWS SDK for Python and configure standard AWS credentials using the normal AWS credential chain. Set AWS_REGION and FALCON_MODEL_ID before running this example:

pip install boto3
import os
import boto3

region = os.environ.get("AWS_REGION", "us-east-1")
model_id = os.environ["FALCON_MODEL_ID"]

client = boto3.client("bedrock-runtime", region_name=region)

response = client.converse(
    modelId=model_id,
    system=[
        {"text": "You are a concise technical assistant."}
    ],
    messages=[
        {
            "role": "user",
            "content": [
                {"text": "Explain what Falcon models are in two sentences."}
            ],
        }
    ],
    inferenceConfig={
        "temperature": 0.2,
        "maxTokens": 200,
    },
)

text = response["output"]["message"]["content"][0]["text"]
print(text)

This example uses the current boto3 Bedrock Runtime client and the native Converse request shape. It does not use a TII-specific API key because the request is authenticated through AWS.

Understanding the response

The Converse response contains an output message with one or more content blocks. For a basic text response, the generated text is available at response["output"]["message"]["content"]. Production code should not assume that every response contains exactly one text block: the selected model and API features may produce different content structures.

Applications should also inspect response metadata and errors, preserve request identifiers where available, and handle throttling or temporary service failures. Keep the model ID configurable so that a deployment can be changed without rewriting application logic.

How pricing generally works

There is no single TII-wide public API price covering all Falcon usage. Open model weights may be downloadable without a TII per-token API charge, subject to the relevant model license and terms. Downloading a model is not the same as operating it: self-hosting still creates infrastructure, storage, networking, monitoring, and engineering costs.

Managed access is billed by the selected provider. On Amazon Bedrock, costs and availability depend on the model listing, usage or throughput arrangement, AWS Region, and applicable AWS or marketplace terms. SageMaker deployments can add endpoint compute, storage, and scaling costs. Check the current AWS listing and console before estimating production spend; the research does not establish a universal Falcon price or fixed per-token rate.

Capabilities developers can access

Falcon capabilities are model- and host-dependent rather than guaranteed platform-wide.

  • Text generation and chat: supported by conversational Falcon models when the serving layer accepts messages or an equivalent prompt format.
  • Multilingual and Arabic use: available in relevant Falcon releases, but language quality and supported languages vary by model.
  • Image input and OCR: available for selected releases such as Falcon Perception and related vision-oriented models, not for the entire Falcon catalog.
  • Audio and video analysis: described for selected Falcon 3 materials, but the exact input format and hosted availability must be confirmed.
  • Code generation and reasoning: advertised by some newer Falcon variants, including smaller models designed for efficient deployment.
  • Tool or function calling: available for some newer models or serving configurations, but not universal across Falcon.

File upload, web search, persistent assistants, agent runtimes, and structured JSON output should not be treated as default TII API features. They may be implemented by a host, application layer, or model-specific serving stack and must be verified for the chosen deployment.

Streaming and advanced requests

Amazon Bedrock provides the ConverseStream operation for streaming supported model responses. Streaming returns partial content events as generation proceeds, allowing an interface to display text progressively instead of waiting for the complete response.

stream = client.converse_stream(
    modelId=model_id,
    messages=[
        {
            "role": "user",
            "content": [
                {"text": "List three practical Falcon deployment options."}
            ],
        }
    ],
    inferenceConfig={"maxTokens": 250},
)

for event in stream.get("stream", []):
    delta = event.get("contentBlockDelta", {}).get("delta", {})
    if "text" in delta:
        print(delta["text"], end="", flush=True)
print()

Streaming support depends on the selected model and deployment. Treat the event structure as part of the Bedrock API contract and test it with the exact Falcon listing you plan to use.

Tool calling, files, and structured output

Some Falcon variants or serving configurations may support function or tool calling. However, the research does not establish a universal TII tool-calling schema. Verify the selected model's input and output contract before building an application around tool calls.

There is no verified universal TII file-upload API. Multimodal or document-processing models may accept particular content types, but file handling can also belong to the host or your own application. A common integration pattern is to validate files in your service, convert them into the format required by the selected model, and apply size and content restrictions before inference.

Structured outputs are likewise not confirmed as a platform-wide Falcon feature. If an application requires strict JSON, validate the generated response in your application and confirm whether the selected host and model offer a supported constrained-output mechanism. Do not assume that a normal text response is schema-valid merely because it resembles JSON.

SDKs, playgrounds, and developer tools

TII does not have a verified first-party SDK for a universal hosted Falcon API. For AWS-managed access, use the current AWS SDK for the chosen language, such as boto3 for Python or @aws-sdk/client-bedrock-runtime for JavaScript. These SDKs use AWS authentication and the Bedrock Runtime API.

The Amazon Bedrock console can be used to inspect available models and experiment with managed deployments, subject to account access and regional availability. TII also provides official model pages and repositories for self-hosting. A local deployment may expose an API supplied by the chosen inference framework, but that endpoint is created and operated by the deploying organization rather than by TII.

Important limits and production considerations

There is no universal TII rate limit, latency guarantee, or uptime commitment identified in the supplied research. Managed limits depend on the provider: Bedrock applies account, Region, model, and throughput quotas, while SageMaker behavior depends on endpoint capacity and scaling configuration. Self-hosted latency depends on hardware, model size, batching, queueing, prompt length, and serving configuration.

Before production deployment, verify all of the following for the exact model and host:

  • license permissions, including restrictions on shared hosted inference or fine-tuning;
  • model availability in the required Region;
  • supported message, image, audio, video, or document formats;
  • streaming, tool calling, and any structured-output behavior;
  • quotas, concurrency, timeout behavior, and retry guidance;
  • billing for inference, endpoint capacity, storage, and marketplace services;
  • data retention, logging, regional processing, and training or improvement terms;
  • safety evaluation, monitoring, authentication, and abuse controls.

Self-hosting offers more control over data location and retention, but the operator becomes responsible for GPU capacity, scaling, network security, access control, observability, patching, safety controls, and model updates. Managed hosting reduces infrastructure work but adds provider dependencies and provider-specific costs and policies.

When Falcon is a good or poor choice

Falcon is a good choice when a team wants downloadable model weights, self-hosting options, control over deployment, or access to specific multilingual, Arabic, efficient, edge, or multimodal research models. It can also fit organizations that already operate AWS infrastructure and want selected Falcon releases through Bedrock or SageMaker.

It is a poorer choice when the primary requirement is a single polished API with stable cross-model behavior, one pricing schedule, universal JSON enforcement, built-in web search, persistent assistants, or broad first-party support. Those characteristics are not established for the TII Falcon ecosystem as a whole.

The safest implementation approach is to treat the model, host, and serving API as one deployment contract. Select the exact Falcon release, verify its license and capabilities, use the host's current native SDK or API, and keep model identifiers and provider-specific settings configurable.

Sources 10