What the TII Falcon developer API is
Technology Innovation Institute (TII) does not appear to provide a universal, first-party hosted inference API with one TII base URL, one API key, and one account-wide quota system. Instead, its developer offering is centered on the Falcon family of open and open-access models, official model repositories, model licenses, research releases, and deployment guidance.
There are two main ways to use Falcon in an application:
- Self-hosting: download a permitted model from TII's official repositories and serve it with an inference framework such as Transformers, vLLM, or another compatible server.
- Managed hosting: deploy or invoke a supported Falcon model through a cloud provider. Amazon Bedrock Marketplace and Amazon SageMaker JumpStart are the clearest documented managed routes in the supplied research.
This distinction is important. TII supplies the model and applicable license, while the selected host normally supplies the endpoint, authentication, quotas, billing, scaling, and operational controls.
Who should use Falcon
Falcon is most suitable for developers, researchers, and organizations that need control over model deployment, data location, infrastructure, or model selection. It is also relevant when a team wants to evaluate multilingual, Arabic-focused, efficient, edge-oriented, or multimodal models.
It is a less direct choice for a beginner who expects a mature consumer-style API with a single subscription, universal feature set, persistent assistants, built-in web research, and guaranteed hosted availability. Those features are not universal across the Falcon ecosystem and may not be provided by TII itself.
Getting access and obtaining credentials
For self-hosting, start with the exact Falcon model repository and model card, then review its license, hardware requirements, supported software, and usage restrictions. Downloading model weights does not create a TII API account or provide a TII API key.
For managed AWS access, create or use an AWS account, select a supported Falcon listing in Amazon Bedrock Marketplace or SageMaker JumpStart, and confirm that the model is available in the required AWS Region. AWS credentials, permissions, account quotas, and any marketplace or service terms apply. The model identifier is deployment-specific and should be copied from the AWS console or listing rather than hard-coded from an unrelated example.
Do not assume that a TII account, the Falcon web application, or a downloaded model supplies credentials for Bedrock or another hosting provider. There is no verified universal TII API key or TII-managed base URL in the current materials.
Choosing the model and API path
Choose the serving method first, then verify the capabilities of the exact Falcon model. Text-oriented releases are different from multimodal models such as Falcon Perception, and support can also change between a model and its serving layer.
| Requirement | Likely path | What to verify |
|---|---|---|
| Managed conversational inference | Amazon Bedrock Runtime | Model ID, Region, Converse support, quotas, and AWS pricing |
| Custom data control or private deployment | Self-hosted inference | License, GPU capacity, serving framework, security, and scaling |
| Managed endpoint deployment | Amazon SageMaker JumpStart | Instance type, endpoint capacity, scaling, and deployment cost |
| Vision or OCR workloads | A supported multimodal Falcon release | Image format, input schema, model availability, and host support |
Falcon-H1, Falcon-H1-Tiny, Falcon 3, Falcon Perception, Falcon Arabic, Falcon Mamba, and other TII releases do not necessarily expose the same inputs or output features. Confirm whether the selected deployment supports image input, audio or video analysis, tool calling, streaming, and the required message format.
Making a first request with Amazon Bedrock
For a managed AWS integration, the current recommended pattern is the Bedrock Runtime Converse operation. It provides a common message-oriented interface for supported models, while AWS handles authentication and the managed endpoint. The exact Falcon model ID must be supplied by the selected AWS deployment.
Install the AWS SDK for Python and configure standard AWS credentials using the normal AWS credential chain. Set AWS_REGION and FALCON_MODEL_ID before running this example:
pip install boto3import os
import boto3
region = os.environ.get("AWS_REGION", "us-east-1")
model_id = os.environ["FALCON_MODEL_ID"]
client = boto3.client("bedrock-runtime", region_name=region)
response = client.converse(
modelId=model_id,
system=[
{"text": "You are a concise technical assistant."}
],
messages=[
{
"role": "user",
"content": [
{"text": "Explain what Falcon models are in two sentences."}
],
}
],
inferenceConfig={
"temperature": 0.2,
"maxTokens": 200,
},
)
text = response["output"]["message"]["content"][0]["text"]
print(text)This example uses the current boto3 Bedrock Runtime client and the native Converse request shape. It does not use a TII-specific API key because the request is authenticated through AWS.
Understanding the response
The Converse response contains an output message with one or more content blocks. For a basic text response, the generated text is available at response["output"]["message"]["content"]. Production code should not assume that every response contains exactly one text block: the selected model and API features may produce different content structures.
Applications should also inspect response metadata and errors, preserve request identifiers where available, and handle throttling or temporary service failures. Keep the model ID configurable so that a deployment can be changed without rewriting application logic.
How pricing generally works
There is no single TII-wide public API price covering all Falcon usage. Open model weights may be downloadable without a TII per-token API charge, subject to the relevant model license and terms. Downloading a model is not the same as operating it: self-hosting still creates infrastructure, storage, networking, monitoring, and engineering costs.
Managed access is billed by the selected provider. On Amazon Bedrock, costs and availability depend on the model listing, usage or throughput arrangement, AWS Region, and applicable AWS or marketplace terms. SageMaker deployments can add endpoint compute, storage, and scaling costs. Check the current AWS listing and console before estimating production spend; the research does not establish a universal Falcon price or fixed per-token rate.
Capabilities developers can access
Falcon capabilities are model- and host-dependent rather than guaranteed platform-wide.
- Text generation and chat: supported by conversational Falcon models when the serving layer accepts messages or an equivalent prompt format.
- Multilingual and Arabic use: available in relevant Falcon releases, but language quality and supported languages vary by model.
- Image input and OCR: available for selected releases such as Falcon Perception and related vision-oriented models, not for the entire Falcon catalog.
- Audio and video analysis: described for selected Falcon 3 materials, but the exact input format and hosted availability must be confirmed.
- Code generation and reasoning: advertised by some newer Falcon variants, including smaller models designed for efficient deployment.
- Tool or function calling: available for some newer models or serving configurations, but not universal across Falcon.
File upload, web search, persistent assistants, agent runtimes, and structured JSON output should not be treated as default TII API features. They may be implemented by a host, application layer, or model-specific serving stack and must be verified for the chosen deployment.
Streaming and advanced requests
Amazon Bedrock provides the ConverseStream operation for streaming supported model responses. Streaming returns partial content events as generation proceeds, allowing an interface to display text progressively instead of waiting for the complete response.
stream = client.converse_stream(
modelId=model_id,
messages=[
{
"role": "user",
"content": [
{"text": "List three practical Falcon deployment options."}
],
}
],
inferenceConfig={"maxTokens": 250},
)
for event in stream.get("stream", []):
delta = event.get("contentBlockDelta", {}).get("delta", {})
if "text" in delta:
print(delta["text"], end="", flush=True)
print()Streaming support depends on the selected model and deployment. Treat the event structure as part of the Bedrock API contract and test it with the exact Falcon listing you plan to use.
Tool calling, files, and structured output
Some Falcon variants or serving configurations may support function or tool calling. However, the research does not establish a universal TII tool-calling schema. Verify the selected model's input and output contract before building an application around tool calls.
There is no verified universal TII file-upload API. Multimodal or document-processing models may accept particular content types, but file handling can also belong to the host or your own application. A common integration pattern is to validate files in your service, convert them into the format required by the selected model, and apply size and content restrictions before inference.
Structured outputs are likewise not confirmed as a platform-wide Falcon feature. If an application requires strict JSON, validate the generated response in your application and confirm whether the selected host and model offer a supported constrained-output mechanism. Do not assume that a normal text response is schema-valid merely because it resembles JSON.
SDKs, playgrounds, and developer tools
TII does not have a verified first-party SDK for a universal hosted Falcon API. For AWS-managed access, use the current AWS SDK for the chosen language, such as boto3 for Python or @aws-sdk/client-bedrock-runtime for JavaScript. These SDKs use AWS authentication and the Bedrock Runtime API.
The Amazon Bedrock console can be used to inspect available models and experiment with managed deployments, subject to account access and regional availability. TII also provides official model pages and repositories for self-hosting. A local deployment may expose an API supplied by the chosen inference framework, but that endpoint is created and operated by the deploying organization rather than by TII.
Important limits and production considerations
There is no universal TII rate limit, latency guarantee, or uptime commitment identified in the supplied research. Managed limits depend on the provider: Bedrock applies account, Region, model, and throughput quotas, while SageMaker behavior depends on endpoint capacity and scaling configuration. Self-hosted latency depends on hardware, model size, batching, queueing, prompt length, and serving configuration.
Before production deployment, verify all of the following for the exact model and host:
- license permissions, including restrictions on shared hosted inference or fine-tuning;
- model availability in the required Region;
- supported message, image, audio, video, or document formats;
- streaming, tool calling, and any structured-output behavior;
- quotas, concurrency, timeout behavior, and retry guidance;
- billing for inference, endpoint capacity, storage, and marketplace services;
- data retention, logging, regional processing, and training or improvement terms;
- safety evaluation, monitoring, authentication, and abuse controls.
Self-hosting offers more control over data location and retention, but the operator becomes responsible for GPU capacity, scaling, network security, access control, observability, patching, safety controls, and model updates. Managed hosting reduces infrastructure work but adds provider dependencies and provider-specific costs and policies.
When Falcon is a good or poor choice
Falcon is a good choice when a team wants downloadable model weights, self-hosting options, control over deployment, or access to specific multilingual, Arabic, efficient, edge, or multimodal research models. It can also fit organizations that already operate AWS infrastructure and want selected Falcon releases through Bedrock or SageMaker.
It is a poorer choice when the primary requirement is a single polished API with stable cross-model behavior, one pricing schedule, universal JSON enforcement, built-in web search, persistent assistants, or broad first-party support. Those characteristics are not established for the TII Falcon ecosystem as a whole.
The safest implementation approach is to treat the model, host, and serving API as one deployment contract. Select the exact Falcon release, verify its license and capabilities, use the host's current native SDK or API, and keep model identifiers and provider-specific settings configurable.
