Developer platform

Aleph Alpha PhariaAI API

Developer overview for Aleph Alpha, including API access, pricing, SDK support, endpoints, capabilities and platform policies.

API key Required
Primary API PhariaInference API
SDK support Current first-party SDKs are Python packages: pharia-inference-sdk, pharia-studio-sdk, and pharia-data-sdk. The legacy aleph-alpha-client remains available for compatibility but its repository was archived on September 22, 2026. Direct HTTP integration is
Rate limits No universal public rate-limit table was identified. Limits are likely deployment-, contract-, workspace-, or installation-specific and should be confirmed with Aleph Alpha or the system administrator.
Platform

API overview

Endpoints

API access

Base URL https://api.aleph-alpha.com
Primary API PhariaInference API
Pricing

API pricing

Pricing model Enterprise quotation; deployment- and usage-dependent

No universal public token-price table was found for the current PhariaAI platform. Pricing may depend on hosted access, deployment, infrastructure, models, usage, support, and platform components.

Developer experience

SDKs & usability

SDKs Current first-party SDKs are Python packages: pharia-inference-sdk, pharia-studio-sdk, and pharia-data-sdk. The legacy aleph-alpha-client remains available for compatibility but its repository was archived on September 22, 2026. Direct HTTP integration is
Ease of use Moderate for enterprise developers. The platform provides a Playground, versioned API documentation, Python SDKs, and bearer-token authentication, but setup is deployment-specific and the current architecture is broader than a simple public REST API.
Documentation Good and improving. Current documentation covers PhariaAI architecture, PhariaStudio, SDK installation, authentication, Playground usage, versioned OpenAPI references, structured output, and deployment administration. Some endpoint details are exposed thr
Latency No universal public latency guarantee was identified. Latency depends on model, hardware, batching, concurrency, network, and whether the service is hosted or self-managed.
Features

API capabilities

✓ Streaming
✓ Function calling
✓ Assistants API
✓ File uploads
✓ Fine-tuning
✓ Image input
✓ Structured outputs
✓ Playground
Feature notes

The current root developer platform is PhariaAI. PhariaInference provides versioned inference APIs, while PhariaStudio provides the development workspace and Playground and PhariaData provides data, document, file, repository, and search capabilities. Current documentation lists PhariaInference API v4.7.0 as the newest documented v4 release. Streaming is supported. Structured output with JSON Schema and structured function calling are documented for compatible deployments and models. Multimodal input is model-dependent. File and document handling is available through PhariaData and related platform services, but direct file-input syntax is deployment- and endpoint-dependent. Fine-tuning and model refinement are part of PhariaStudio capabilities, subject to installation and access. The legacy Intelligence Layer SDK is deprecated; Aleph Alpha recommends pharia-inference-sdk, pharia-studio-sdk, and pharia-data-sdk. The code examples use the compatibility-oriented completion client and a configurable host because hosted and self-managed PhariaAI installations can expose different service URLs and endpoint versions. Developers should validate the exact route and request schema in the target deployment's Swagger reference before production use.

Examples

API examples

curl --fail-with-body --silent --show-error "$PHARIA_INFERENCE_URL/complete" \
  -H "Authorization: Bearer ${AA_TOKEN}" \
  -H "Content-Type: application/json" \
  --data @- | jq -r '.completions[0].completion'
{
  "model": "pharia-1-llm-7b-control",
  "prompt": "Explain data sovereignty in one paragraph.",
  "maximum_tokens": 128,
  "temperature": 0.2
}

# Streaming example; verify the exact route and payload against the target PhariaInference deployment.
curl --fail-with-body --silent --show-error --no-buffer "$PHARIA_INFERENCE_URL/complete" \
  -H "Authorization: Bearer ${AA_TOKEN}" \
  -H "Content-Type: application/json" \
  --data @- | jq -r '.completion // .completions[0].completion // empty'
{
  "model": "pharia-1-llm-7b-control",
  "prompt": "List three benefits of sovereign AI.",
  "maximum_tokens": 128,
  "stream": true
}
# Install: pip install aleph-alpha-client
import os
from aleph_alpha_client import Client, CompletionRequest, Prompt

TOKEN = os.environ["AA_TOKEN"]
HOST = os.environ.get("PHARIA_INFERENCE_URL", "https://api.aleph-alpha.com")
MODEL = os.environ.get("AA_MODEL", "pharia-1-llm-7b-control")

client = Client(token=TOKEN, host=HOST)
request = CompletionRequest(
    prompt=Prompt.from_text("You are a concise technical assistant.\n\nExplain data sovereignty in one paragraph."),
    maximum_tokens=128,
    temperature=0.2,
)

try:
    response = client.complete(request, model=MODEL)
    print(response.completions[0].completion)
except Exception as exc:
    raise SystemExit(f"Aleph Alpha request failed: {exc}")

# Streaming pattern
stream = client.complete_with_streaming(
    CompletionRequest(
        prompt=Prompt.from_text("Give three practical benefits of sovereign AI."),
        maximum_tokens=128,
    ),
    model=MODEL,
)
for item in stream:
    print(item, end="", flush=True)
import process from "node:process";

const token = process.env.AA_TOKEN;
const baseUrl = process.env.PHARIA_INFERENCE_URL || "https://api.aleph-alpha.com";
const model = process.env.AA_MODEL || "pharia-1-llm-7b-control";

if (!token) {
  throw new Error("AA_TOKEN is required");
}

const response = await fetch(`${baseUrl}/complete`, {
  method: "POST",
  headers: {
    Authorization: `Bearer ${token}`,
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model,
    prompt: "You are a concise technical assistant. Explain data sovereignty in one paragraph.",
    maximum_tokens: 128,
    temperature: 0.2
  })
});

if (!response.ok) {
  const body = await response.text();
  throw new Error(`Aleph Alpha HTTP ${response.status}: ${body}`);
}

const data = await response.json();
console.log(data.completions?.[0]?.completion ?? data.completion ?? data);
<?php
$token = getenv('AA_TOKEN');
$baseUrl = getenv('PHARIA_INFERENCE_URL') ?: 'https://api.aleph-alpha.com';
$model = getenv('AA_MODEL') ?: 'pharia-1-llm-7b-control';

if (!$token) {
    throw new RuntimeException('AA_TOKEN is required');
}

$payload = json_encode([
    'model' => $model,
    'prompt' => 'You are a concise technical assistant. Explain data sovereignty in one paragraph.',
    'maximum_tokens' => 128,
    'temperature' => 0.2,
], JSON_THROW_ON_ERROR);

$ch = curl_init($baseUrl . '/complete');
curl_setopt_array($ch, [
    CURLOPT_POST => true,
    CURLOPT_POSTFIELDS => $payload,
    CURLOPT_HTTPHEADER => [
        'Authorization: Bearer ' . $token,
        'Content-Type: application/json',
        'Accept: application/json',
    ],
    CURLOPT_RETURNTRANSFER => true,
    CURLOPT_TIMEOUT => 120,
]);

$body = curl_exec($ch);
if ($body === false) {
    throw new RuntimeException('cURL error: ' . curl_error($ch));
}

$status = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);

$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
if ($status < 200 || $status >= 300) {
    throw new RuntimeException('Aleph Alpha HTTP ' . $status . ': ' . $body);
}

$result = $data['completions'][0]['completion'] ?? $data['completion'] ?? null;
if ($result === null) {
    throw new RuntimeException('No completion found in response: ' . $body);
}

echo $result . PHP_EOL;
Policies

Data & usage

Data training

Aleph Alpha positions PhariaAI around sovereign, customer-controlled enterprise and government deployments. The reviewed current developer documentation does not establish one universal public rule stating that all API submissions are or are not used for model training. Training-use terms should therefore be confirmed in the applicable contract, hosted-service terms, or self-managed deployment configuration. Self-managed deployments provide stronger operational control because inference and submitted data can remain within the organization's environment.

Data retention

No universal public retention period was identified in the reviewed current API documentation. Retention depends on the hosted service, deployment configuration, workspace, logging settings, and contract. Self-managed PhariaAI installations allow the organization to control infrastructure and operational data policies more directly.

Rate limits

No universal public rate-limit table was identified. Limits are likely deployment-, contract-, workspace-, or installation-specific and should be confirmed with Aleph Alpha or the system administrator.

Developer guide

Aleph Alpha PhariaAI API: Beginner’s Guide to the Developer Platform

Aleph Alpha’s current developer platform is PhariaAI, an enterprise-focused environment that combines PhariaInference for model access, PhariaStudio for application development and evaluation, and PhariaData for documents, files, repositories, and search. Developers can use token-authenticated APIs for inference, streaming, structured JSON output, function calling, multimodal workloads on supported models, and organization-specific deployments.
Aleph Alpha’s developer offering is now centered on PhariaAI rather than a single consumer chatbot endpoint. It is designed mainly for enterprises, government organizations, and regulated teams that need control over data, deployment, models, and application workflows. The platform can be used through hosted services or organization-specific sovereign and on-premises installations, so exact URLs, models, permissions, limits, and available features depend on the target deployment.
PhariaAI is Aleph Alpha’s current enterprise developer platform, combining PhariaInference for model access, PhariaStudio for application development, and PhariaData for documents and search. It supports token-authenticated inference, streaming, structured JSON output, function calling, multimodal workloads on supported models, and sovereign deployment patterns. Pricing, limits, available models, retention, and file behavior depend on the specific hosted or self-managed installation.

What the PhariaAI API is and when to use it

PhariaAI is Aleph Alpha’s current developer platform. Its main components have distinct roles: PhariaInference provides model inference APIs, PhariaStudio provides a development workspace and Playground, and PhariaData provides data, document, file, repository, connector, and search capabilities.

This structure is intended for teams building controlled AI applications rather than individuals looking for a self-service consumer chatbot. Typical projects include internal assistants, document analysis, retrieval-augmented generation, domain-specific workflows, government applications, and systems that must run on European infrastructure or inside an organization’s own environment.

Use PhariaAI when deployment control, sovereignty, explainability, domain adaptation, and enterprise integration matter. It is a less natural choice for casual experimentation, low-cost public applications, or developers seeking a globally standardized consumer API with one public pricing table and identical limits for every customer.

Getting access and obtaining credentials

Access is normally organization-based. Depending on the arrangement, Aleph Alpha or an administrator provides access to a hosted PhariaAI environment, a private installation, or an organization-specific service URL. Public documentation does not establish one universal onboarding process for every deployment.

API requests use a bearer token. In PhariaStudio, an authorized user can copy a bearer token from their profile. Hosted API access may use an Aleph Alpha account and profile, while self-managed installations can issue credentials through the organization’s own administration process.

Store the token in an environment variable rather than putting it directly in source code. The exact service URL is also deployment-specific, so use the inference URL supplied by the administrator or shown in the target installation’s documentation.

export AA_TOKEN="your-token"
export PHARIA_INFERENCE_URL="https://your-pharia-inference-host"
export AA_MODEL="pharia-1-llm-7b-control"

Choosing the API and model

Start with PhariaInference when your application needs text or multimodal model inference. Use PhariaStudio when you need a workspace for prompt development, debugging, evaluation, application construction, or model refinement. Use PhariaData when the application must manage documents, files, datasets, repositories, connectors, or search stores.

Model identifiers and supported operations are deployment-dependent. The publicly documented Pharia-1-LLM-7B-control family is one example, but a model that exists in one installation may not be enabled in another. Confirm the available model list, supported input types, and request schema in the target deployment’s API reference or Swagger documentation before writing production code.

Making a first inference request

The following compatibility-oriented HTTP example illustrates the documented completion pattern. Replace the host and model with values supported by your PhariaAI installation. The exact route and payload should be checked against the versioned PhariaInference reference for that deployment.

curl --fail-with-body --silent --show-error 
  "$PHARIA_INFERENCE_URL/complete" 
  -H "Authorization: Bearer ${AA_TOKEN}" 
  -H "Content-Type: application/json" 
  --data @- | jq -r '.completions[0].completion'
{
  "model": "pharia-1-llm-7b-control",
  "prompt": "Explain data sovereignty in one paragraph.",
  "maximum_tokens": 128,
  "temperature": 0.2
}

The request supplies a model, a prompt, and generation settings. maximum_tokens places an upper bound on the generated completion, while temperature influences variation. Available parameters and their names can vary with the API version and deployment.

Understanding the response

A completion response contains generated text in a completion object or in the first item of a completions collection, depending on the endpoint response format. Applications should handle the documented response schema rather than assuming that every deployment returns exactly the same wrapper.

For production code, check the HTTP status before parsing the body, handle authentication and validation errors, and log request identifiers or service errors where the deployment exposes them. Do not assume that a successful HTTP response means the generated content is suitable for a business decision; application-level validation is still required.

How pricing generally works

Aleph Alpha does not publish a universal self-service token-price table for the current PhariaAI developer platform in the reviewed materials. Pricing is generally provided through an enterprise quotation.

The commercial arrangement may depend on hosted access, the deployment model, infrastructure, selected models, usage, support, and optional components such as development or data services. Hosted and private deployments can therefore have substantially different cost structures.

Before implementation, ask for the pricing basis, included environments, model availability, usage measurement, support terms, data-processing terms, rate limits, and costs associated with private or on-premises operation. Do not estimate production cost from a generic per-token assumption unless it is present in your contract.

Capabilities available to developers

Streaming responses

Streaming is supported for compatible inference workflows. Instead of waiting for the complete answer, an application can process output as it becomes available. This is useful for interactive interfaces and long responses, but the event or chunk format must be confirmed for the selected API version.

curl --fail-with-body --silent --show-error --no-buffer 
  "$PHARIA_INFERENCE_URL/complete" 
  -H "Authorization: Bearer ${AA_TOKEN}" 
  -H "Content-Type: application/json" 
  --data @- | jq -r '.completion // .completions[0].completion // empty'
{
  "model": "pharia-1-llm-7b-control",
  "prompt": "List three benefits of sovereign AI.",
  "maximum_tokens": 128,
  "stream": true
}

Tool and function calling

Structured function calling is documented for compatible models and configurations. The model returns information describing a requested function and its arguments; your application then validates the arguments, executes the external operation, and sends the result back into the conversation.

Function calling does not grant the model direct access to your database or services. Your application remains responsible for authorization, input validation, side effects, timeouts, and error handling. Check the supported function-calling schema for the model and PhariaInference version you are using.

Structured JSON output

PhariaAI supports structured output using JSON Schema in supported chat-completion workflows. This is useful when the response must be consumed by another program, such as an extraction pipeline, classification service, or workflow engine.

Schema-constrained output reduces parsing ambiguity, but your application should still validate the returned JSON, handle missing or invalid values, and plan for model or deployment-specific restrictions.

Multimodal input and files

Multimodal input is available for supported models and deployments. File and document functionality is also available through PhariaData and related platform services. These services cover files, documents, datasets, repositories, connectors, and search stores.

There is no single universal file-input syntax or file-type list established for every PhariaAI installation in the supplied documentation. Confirm supported formats, upload routes, size limits, indexing behavior, and access permissions with the target deployment before designing a document workflow.

Customization and refinement

PhariaStudio supports organization-specific application development, evaluation, retrieval-augmented generation workflows, and model fine-tuning or refinement processes where enabled. Availability depends on the installation, permissions, selected model, and commercial arrangement.

SDKs, Playground, and developer tools

Aleph Alpha’s current first-party SDK direction is Python-oriented and separated by platform component. The relevant packages are the pharia-inference-sdk, pharia-studio-sdk, and pharia-data-sdk. Direct HTTP integration is an option for other languages when the required endpoint is exposed.

The older Intelligence Layer SDK is deprecated and should not be the basis for a new integration. The legacy client may remain relevant to existing compatibility work, but new projects should begin with the current Pharia SDKs and versioned PhariaAI documentation.

PhariaStudio includes a Playground where authorized users can select an available model, test prompts, change settings, inspect results, and export prompt code. It is an enterprise development workspace, not an anonymous public playground, and its features depend on the organization’s installation and permissions.

A practical advanced workflow

  1. Use PhariaStudio to test a prompt and identify an available model.
  2. Use PhariaData to organize the documents, files, or search stores required by the application.
  3. Call PhariaInference with a bearer token from a protected server-side application.
  4. Use structured output when downstream code needs a predictable JSON shape.
  5. Use function calling only for explicitly approved operations, and validate every returned argument before execution.
  6. Measure latency, concurrency, response quality, and failure behavior in the actual hosted or self-managed environment.

This separation helps keep model inference, application development, and enterprise data operations distinct while still allowing them to work together in one platform.

Limits and production considerations

No universal public rate-limit table was identified for the current PhariaAI platform. Limits may depend on the deployment, contract, workspace, model, or installation. Confirm requests per minute, concurrency, token limits, quotas, and service commitments before launch.

There is also no universal public latency guarantee. Response time depends on the selected model, hardware, batching, concurrency, network conditions, and whether the service is hosted or self-managed. Benchmark the real deployment with representative prompts rather than relying on a provider-wide figure.

The reviewed documentation does not establish one universal retention period or one universal rule for whether all API submissions are used for model training. Confirm logging, retention, training use, deletion, and data-processing terms in the applicable contract or deployment configuration. Self-managed deployments can provide more direct control over infrastructure and operational data policies.

Finally, validate endpoint versions and model identifiers before production use. PhariaAI is a broader platform than a single fixed public endpoint, and capabilities can vary between installations.

Advantages and limitations

Advantages:

  • Designed for European, enterprise, government, and regulated use cases.
  • Supports hosted, private, sovereign, and on-premises deployment patterns.
  • Combines inference, application development, evaluation, data, documents, and search services.
  • Provides documented streaming, structured output, and function-calling capabilities for compatible deployments.
  • Offers Python-oriented SDKs and a browser-based development Playground.

Limitations:

  • Access and setup are more organization-specific than a typical self-service API.
  • Pricing, rate limits, retention, latency, available models, and file behavior are not defined by one universal public specification.
  • Many capabilities depend on deployment configuration, permissions, and model support.
  • There is no broadly marketed consumer subscription ecosystem for casual experimentation.
  • Developers using JavaScript, TypeScript, PHP, or Java may need to integrate through HTTP because equivalent current first-party SDK coverage was not established.

Who should use PhariaAI?

PhariaAI is a good choice for organizations that need sovereign or private AI deployment, European data and infrastructure options, domain-specific models, controlled enterprise workflows, and integrated document or search capabilities. It is particularly relevant when procurement, compliance, access control, and operational ownership are as important as model output.

It is a poorer fit for a hobby project that requires instant anonymous access, transparent public token pricing, globally uniform limits, or a large consumer application ecosystem. In those cases, the organization-specific onboarding and deployment model may add more complexity than the project requires.

Sources 13