What the PhariaAI API is and when to use it
PhariaAI is Aleph Alpha’s current developer platform. Its main components have distinct roles: PhariaInference provides model inference APIs, PhariaStudio provides a development workspace and Playground, and PhariaData provides data, document, file, repository, connector, and search capabilities.
This structure is intended for teams building controlled AI applications rather than individuals looking for a self-service consumer chatbot. Typical projects include internal assistants, document analysis, retrieval-augmented generation, domain-specific workflows, government applications, and systems that must run on European infrastructure or inside an organization’s own environment.
Use PhariaAI when deployment control, sovereignty, explainability, domain adaptation, and enterprise integration matter. It is a less natural choice for casual experimentation, low-cost public applications, or developers seeking a globally standardized consumer API with one public pricing table and identical limits for every customer.
Getting access and obtaining credentials
Access is normally organization-based. Depending on the arrangement, Aleph Alpha or an administrator provides access to a hosted PhariaAI environment, a private installation, or an organization-specific service URL. Public documentation does not establish one universal onboarding process for every deployment.
API requests use a bearer token. In PhariaStudio, an authorized user can copy a bearer token from their profile. Hosted API access may use an Aleph Alpha account and profile, while self-managed installations can issue credentials through the organization’s own administration process.
Store the token in an environment variable rather than putting it directly in source code. The exact service URL is also deployment-specific, so use the inference URL supplied by the administrator or shown in the target installation’s documentation.
export AA_TOKEN="your-token"
export PHARIA_INFERENCE_URL="https://your-pharia-inference-host"
export AA_MODEL="pharia-1-llm-7b-control"Choosing the API and model
Start with PhariaInference when your application needs text or multimodal model inference. Use PhariaStudio when you need a workspace for prompt development, debugging, evaluation, application construction, or model refinement. Use PhariaData when the application must manage documents, files, datasets, repositories, connectors, or search stores.
Model identifiers and supported operations are deployment-dependent. The publicly documented Pharia-1-LLM-7B-control family is one example, but a model that exists in one installation may not be enabled in another. Confirm the available model list, supported input types, and request schema in the target deployment’s API reference or Swagger documentation before writing production code.
Making a first inference request
The following compatibility-oriented HTTP example illustrates the documented completion pattern. Replace the host and model with values supported by your PhariaAI installation. The exact route and payload should be checked against the versioned PhariaInference reference for that deployment.
curl --fail-with-body --silent --show-error
"$PHARIA_INFERENCE_URL/complete"
-H "Authorization: Bearer ${AA_TOKEN}"
-H "Content-Type: application/json"
--data @- | jq -r '.completions[0].completion'
{
"model": "pharia-1-llm-7b-control",
"prompt": "Explain data sovereignty in one paragraph.",
"maximum_tokens": 128,
"temperature": 0.2
}The request supplies a model, a prompt, and generation settings. maximum_tokens places an upper bound on the generated completion, while temperature influences variation. Available parameters and their names can vary with the API version and deployment.
Understanding the response
A completion response contains generated text in a completion object or in the first item of a completions collection, depending on the endpoint response format. Applications should handle the documented response schema rather than assuming that every deployment returns exactly the same wrapper.
For production code, check the HTTP status before parsing the body, handle authentication and validation errors, and log request identifiers or service errors where the deployment exposes them. Do not assume that a successful HTTP response means the generated content is suitable for a business decision; application-level validation is still required.
How pricing generally works
Aleph Alpha does not publish a universal self-service token-price table for the current PhariaAI developer platform in the reviewed materials. Pricing is generally provided through an enterprise quotation.
The commercial arrangement may depend on hosted access, the deployment model, infrastructure, selected models, usage, support, and optional components such as development or data services. Hosted and private deployments can therefore have substantially different cost structures.
Before implementation, ask for the pricing basis, included environments, model availability, usage measurement, support terms, data-processing terms, rate limits, and costs associated with private or on-premises operation. Do not estimate production cost from a generic per-token assumption unless it is present in your contract.
Capabilities available to developers
Streaming responses
Streaming is supported for compatible inference workflows. Instead of waiting for the complete answer, an application can process output as it becomes available. This is useful for interactive interfaces and long responses, but the event or chunk format must be confirmed for the selected API version.
curl --fail-with-body --silent --show-error --no-buffer
"$PHARIA_INFERENCE_URL/complete"
-H "Authorization: Bearer ${AA_TOKEN}"
-H "Content-Type: application/json"
--data @- | jq -r '.completion // .completions[0].completion // empty'
{
"model": "pharia-1-llm-7b-control",
"prompt": "List three benefits of sovereign AI.",
"maximum_tokens": 128,
"stream": true
}Tool and function calling
Structured function calling is documented for compatible models and configurations. The model returns information describing a requested function and its arguments; your application then validates the arguments, executes the external operation, and sends the result back into the conversation.
Function calling does not grant the model direct access to your database or services. Your application remains responsible for authorization, input validation, side effects, timeouts, and error handling. Check the supported function-calling schema for the model and PhariaInference version you are using.
Structured JSON output
PhariaAI supports structured output using JSON Schema in supported chat-completion workflows. This is useful when the response must be consumed by another program, such as an extraction pipeline, classification service, or workflow engine.
Schema-constrained output reduces parsing ambiguity, but your application should still validate the returned JSON, handle missing or invalid values, and plan for model or deployment-specific restrictions.
Multimodal input and files
Multimodal input is available for supported models and deployments. File and document functionality is also available through PhariaData and related platform services. These services cover files, documents, datasets, repositories, connectors, and search stores.
There is no single universal file-input syntax or file-type list established for every PhariaAI installation in the supplied documentation. Confirm supported formats, upload routes, size limits, indexing behavior, and access permissions with the target deployment before designing a document workflow.
Customization and refinement
PhariaStudio supports organization-specific application development, evaluation, retrieval-augmented generation workflows, and model fine-tuning or refinement processes where enabled. Availability depends on the installation, permissions, selected model, and commercial arrangement.
SDKs, Playground, and developer tools
Aleph Alpha’s current first-party SDK direction is Python-oriented and separated by platform component. The relevant packages are the pharia-inference-sdk, pharia-studio-sdk, and pharia-data-sdk. Direct HTTP integration is an option for other languages when the required endpoint is exposed.
The older Intelligence Layer SDK is deprecated and should not be the basis for a new integration. The legacy client may remain relevant to existing compatibility work, but new projects should begin with the current Pharia SDKs and versioned PhariaAI documentation.
PhariaStudio includes a Playground where authorized users can select an available model, test prompts, change settings, inspect results, and export prompt code. It is an enterprise development workspace, not an anonymous public playground, and its features depend on the organization’s installation and permissions.
A practical advanced workflow
- Use PhariaStudio to test a prompt and identify an available model.
- Use PhariaData to organize the documents, files, or search stores required by the application.
- Call PhariaInference with a bearer token from a protected server-side application.
- Use structured output when downstream code needs a predictable JSON shape.
- Use function calling only for explicitly approved operations, and validate every returned argument before execution.
- Measure latency, concurrency, response quality, and failure behavior in the actual hosted or self-managed environment.
This separation helps keep model inference, application development, and enterprise data operations distinct while still allowing them to work together in one platform.
Limits and production considerations
No universal public rate-limit table was identified for the current PhariaAI platform. Limits may depend on the deployment, contract, workspace, model, or installation. Confirm requests per minute, concurrency, token limits, quotas, and service commitments before launch.
There is also no universal public latency guarantee. Response time depends on the selected model, hardware, batching, concurrency, network conditions, and whether the service is hosted or self-managed. Benchmark the real deployment with representative prompts rather than relying on a provider-wide figure.
The reviewed documentation does not establish one universal retention period or one universal rule for whether all API submissions are used for model training. Confirm logging, retention, training use, deletion, and data-processing terms in the applicable contract or deployment configuration. Self-managed deployments can provide more direct control over infrastructure and operational data policies.
Finally, validate endpoint versions and model identifiers before production use. PhariaAI is a broader platform than a single fixed public endpoint, and capabilities can vary between installations.
Advantages and limitations
Advantages:
- Designed for European, enterprise, government, and regulated use cases.
- Supports hosted, private, sovereign, and on-premises deployment patterns.
- Combines inference, application development, evaluation, data, documents, and search services.
- Provides documented streaming, structured output, and function-calling capabilities for compatible deployments.
- Offers Python-oriented SDKs and a browser-based development Playground.
Limitations:
- Access and setup are more organization-specific than a typical self-service API.
- Pricing, rate limits, retention, latency, available models, and file behavior are not defined by one universal public specification.
- Many capabilities depend on deployment configuration, permissions, and model support.
- There is no broadly marketed consumer subscription ecosystem for casual experimentation.
- Developers using JavaScript, TypeScript, PHP, or Java may need to integrate through HTTP because equivalent current first-party SDK coverage was not established.
Who should use PhariaAI?
PhariaAI is a good choice for organizations that need sovereign or private AI deployment, European data and infrastructure options, domain-specific models, controlled enterprise workflows, and integrated document or search capabilities. It is particularly relevant when procurement, compliance, access control, and operational ownership are as important as model output.
It is a poorer fit for a hobby project that requires instant anonymous access, transparent public token pricing, globally uniform limits, or a large consumer application ecosystem. In those cases, the organization-specific onboarding and deployment model may add more complexity than the project requires.
