Google DeepMind
Gemini API
Free tier plus prepaid and pay-as-you-go paid tiers; usage is primarily metered by model, input tokens, output tokens, cached tokens, storage duration, service tier, batch mode, and tool usage.
Free access is available for eligible models and limits. Paid access provides higher limits, advanced models, context caching, batch processing, and production features. Current prices vary by model and are listed on the official pricing page.
SDKs
Official Google GenAI SDKs for Python, JavaScript/TypeScript, Go, and Java. REST is also supported. Google recommends migrating from legacy Gemini SDKs to the Google GenAI SDK.
Ease of use
High. Google provides AI Studio, API-key creation, quickstarts, REST examples, and official production-ready SDKs. The recommended Interactions API offers a unified model and agent interface.
Documentation
Strong and extensive, with current quickstarts, API reference pages, migration guidance, SDK documentation, capability guides, pricing, rate limits, files, agents, and privacy documentation.
Streaming
Function calling
Assistants & agents
File uploads
Web search
+3
Reka AI
Reka API
Usage-based pay-as-you-go, primarily charged per input and output token; media and Vision actions may use separate units.
Published Reka Chat rates include Edge at $0.10/$0.10, Flash at $0.80/$2.00, and Core at $2.00/$6.00 per million input/output tokens. Image, video, and audio charges may apply.
SDKs
OpenAI Python and JavaScript SDKs are officially documented through the OpenAI-compatible interface. Official examples also cover Go, Java, and REST. A dedicated Reka SDK was not verified.
Ease of use
High for basic integrations because the Chat API is OpenAI-compatible and uses familiar chat-completions patterns. Complexity increases when combining Chat, Research, Vision, and Gateway product surfaces.
Documentation
Good and actively maintained API reference with quickstarts, model information, streaming, multimodal, tool, Research, and Vision documentation. Some product surfaces use different authentication conventions and base URLs.
Streaming
Function calling
File uploads
Fine-tuning
Web search
+3
LG AI Research
EXAONE API through FriendliAI
Usage-based hosted inference; Dedicated Endpoints are available as an alternative capacity-based deployment model.
No publicly verified EXAONE-specific price was found. Pricing is determined by FriendliAI account, Model API, or Dedicated Endpoint terms.
SDKs
OpenAI-compatible clients can be used. Verified practical options include the official OpenAI Python SDK, the official OpenAI JavaScript SDK, raw HTTP, and any compatible client that allows a custom base URL.
Ease of use
High for developers familiar with OpenAI-compatible APIs; use the Friendli endpoint, token, model identifier, and standard chat-completion schema.
Documentation
Moderate. FriendliAI documents the OpenAI-compatible platform and LG AI Research publishes model documentation, but there is no single LG-owned API reference covering the complete commercial integration.
Streaming
Function calling
Structured outputs
Playground
NAVER AI
CLOVA Studio API
Usage-based billing based on model, purpose, and token usage
Current pricing is published in the NAVER Cloud Platform portal under AI Services > CLOVA Studio; official documentation does not provide one universal public price table.
SDKs
No provider-specific official SDK was identified in the consulted documentation. Official OpenAI Python and JavaScript SDKs can be used through the OpenAI-compatible endpoint; other languages can use REST.
Ease of use
Moderate. REST APIs are straightforward, and OpenAI-compatible endpoints simplify migration, but account subscription, key issuance, Korean-region availability, model-specific restrictions, and service-app review add setup requirements.
Documentation
Good and detailed. Current documentation covers native APIs, OpenAI compatibility, authentication, models, streaming, tools, structured outputs, tuning, limits, errors and quickstarts.
Streaming
Function calling
File uploads
Fine-tuning
Image input
+2
Aleph Alpha
Aleph Alpha PhariaAI API
Enterprise quotation; deployment- and usage-dependent
No universal public token-price table was found for the current PhariaAI platform. Pricing may depend on hosted access, deployment, infrastructure, models, usage, support, and platform components.
SDKs
Current first-party SDKs are Python packages: pharia-inference-sdk, pharia-studio-sdk, and pharia-data-sdk. The legacy aleph-alpha-client remains available for compatibility but its repository was archived on September 22, 2026. Direct HTTP integration is
Ease of use
Moderate for enterprise developers. The platform provides a Playground, versioned API documentation, Python SDKs, and bearer-token authentication, but setup is deployment-specific and the current architecture is broader than a simple public REST API.
Documentation
Good and improving. Current documentation covers PhariaAI architecture, PhariaStudio, SDK installation, authentication, Playground usage, versioned OpenAPI references, structured output, and deployment administration. Some endpoint details are exposed thr
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+3
Xiaomi HyperAI
Xiaomi MiMo API Open Platform
Pay-as-you-go token billing, fixed Token Plan subscriptions, and discounted Batch API billing
Overseas real-time pricing includes MiMo V2.6 Pro at $0.435 per million input tokens and $0.87 per million output tokens, and MiMo V2.6 Flash at $0.14 per million input tokens and $0.28 per million output tokens. Cached-input pricing is lower. Batch infer
SDKs
Official examples use the OpenAI Python SDK and Anthropic Python SDK through compatible base URLs. JavaScript and PHP developers can use the OpenAI-compatible HTTP API; no Xiaomi-specific SDK was verified.
Ease of use
Moderate to easy because the platform supports OpenAI and Anthropic-compatible protocols and works with common SDKs. Account, credential type, regional endpoint, model capability, and billing differences require attention.
Documentation
Good and improving. Official documentation includes quickstarts, API references, model catalogs, examples, structured output, multimodal input, web search, batch processing, rate limits, error codes, and migration updates.
Streaming
Function calling
Assistants & agents
File uploads
Web search
+3
Technology Innovation Institute (TII)
TII Falcon Developer Platform
Open-weight and self-hosted deployment; managed inference through partner cloud platforms using their billing and authentication
No universal TII-hosted API price was verified. Falcon model access may be free under the applicable license, while Amazon Bedrock and SageMaker usage is billed by AWS.
SDKs
No verified first-party TII SDK. Managed integrations can use AWS SDKs such as boto3 and AWS SDK for JavaScript, or compatible HTTP clients.
Ease of use
Moderate. Self-hosting requires model-serving and GPU expertise; managed AWS deployment is easier but requires AWS configuration and model-specific setup.
Documentation
Moderate. TII provides model pages, licenses, research, and release information, while detailed API operations are documented by the selected hosting provider.
Streaming
Function calling
Fine-tuning
Image input
Playground
SenseTime
SenseNova API
Usage- and quota-based API pricing; selected models may have public-beta quotas
SenseNova U1.5 Lite currently has a documented free public-beta quota of 1,500 requests every five hours. General commercial model pricing was not reliably verified from the accessible official pages.
SDKs
Raw HTTPS and OpenAI-compatible clients are supported. An official current Python, JavaScript, or PHP SDK package for the Token Plan API was not reliably verified.
Ease of use
Moderate; REST and OpenAI-compatible patterns are familiar, but endpoint generations and model-specific documentation require careful verification.
Documentation
Moderate; official endpoint documentation is available, but much of the canonical material is Chinese-language and the public English coverage is incomplete.
Streaming
Function calling
Assistants & agents
Image input
Playground
Tencent AI
Tencent Cloud TokenHub API
Model-specific usage pricing with pay-as-you-go billing and Token Plan enterprise quota options
Pricing varies by model and usage dimension, including input tokens, output tokens, cached tokens, reasoning tokens, image usage and search calls. Current prices must be checked in the TokenHub model catalog.
SDKs
OpenAI Python and JavaScript SDKs work through the compatible API by changing base_url. Tencent Cloud signed SDKs support Python, Java, PHP, Go, Node.js, .NET, C++, Ruby and other languages for control-plane APIs.
Ease of use
High for developers familiar with OpenAI-compatible APIs; configure a bearer API key, custom base URL and model ID. Regional account and model activation requirements still apply.
Documentation
Good and actively updated, with API usage guides, protocol references, model listings, migration documentation, error codes, multimodal examples and enterprise operations guidance.
Streaming
Function calling
File uploads
Web search
Image input
+2
StepFun
StepFun Open Platform API
Usage-based pay-as-you-go pricing, with model-specific token, character, duration, per-image, search-call, and file-storage billing units; Step Plan subscription access is separate.
Current international prices include Step 5 Preview at $1.00/M cache-miss input tokens, $0.05/M cache-hit input tokens, and $2.70/M output tokens; Step 3.7 Flash at $0.20/$0.04/$1.15; Step 3.5 Flash at $0.10/$0.02/$0.30. Internet search is $0.006/call, im
SDKs
OpenAI-compatible SDK usage is officially documented for Python and JavaScript. Raw HTTP is supported for any language. No separate first-party StepFun Python, JavaScript, or PHP SDK requirement was verified.
Ease of use
High for developers familiar with OpenAI-compatible APIs. Standard HTTP endpoints and OpenAI Python and JavaScript SDK configuration are documented. Region-specific accounts and model-specific feature differences require attention.
Documentation
Good and improving. Current English documentation includes quickstarts, API references, model guides, pricing, rate limits, privacy, file APIs, Responses, realtime, tools, migration notices, and runnable Python, JavaScript, and cURL examples.
Streaming
Function calling
Assistants & agents
File uploads
Web search
+3
ByteDance Seed
BytePlus ModelArk
Pay-as-you-go usage billing with token-based online and batch inference, per-image pricing for image models, model-specific video pricing, and hourly or monthly model-unit pricing.
Published Seed examples include seed-2-0-lite-260428 at USD 0.25 per million input tokens and USD 2.00 per million output tokens for prompts up to 128K tokens; seed-2-0-mini-260428 is listed at USD 0.10 per million input tokens and USD 0.40 per million ou
SDKs
OpenAI-compatible Python and JavaScript integrations are documented. BytePlus/ModelArk SDK examples are available for Python, Go, Java, and other supported environments. Community SDKs should not be treated as official unless identified in the documentati
Ease of use
Good for developers familiar with OpenAI-compatible APIs. Main setup requirements are BytePlus account activation, model activation, regional base URL selection, API-key management, and exact versioned model IDs.
Documentation
Good and detailed, with current model pages, API references, pricing, authentication, playground, streaming, files, agents, reasoning, and multimodal examples. Documentation is model-version and region specific.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+4
AI21 Labs
AI21 Studio API
Usage-based token pricing with enterprise and private-deployment agreements
Hosted access is usage-based and account/model dependent; current commercial terms should be verified in the AI21 account or pricing documentation. Enterprise private deployments use negotiated pricing.
SDKs
Official Python and TypeScript SDKs. Raw HTTP is also supported. The SDKs include chat completions, streaming, Maestro, library files, and conversational RAG functionality.
Ease of use
Moderate to easy. The REST API follows a familiar chat-completions pattern, and official Python and TypeScript SDKs simplify authentication, requests, streaming, Maestro, and file operations.
Documentation
Good and actively maintained. Current documentation covers AI21 Studio, Jamba, Maestro, SDKs, authentication, rate limits, files, retrieval, function calling, fine-tuning, deployment, and API references.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+3
IBM watsonx
IBM watsonx.ai API
Plan-based and usage-based pricing
Free playground and trial allocations are available. Paid Essentials is pay-as-you-go; Standard has a published starting instance fee of USD 1,110 per month, with additional model, token, capacity, hosting, tuning and feature charges varying by region and
SDKs
Official Python package ibm-watsonx-ai and official Node.js package @ibm-cloud/watsonx-ai. REST HTTP access is available for any language, including PHP.
Ease of use
Moderate. Prompt Lab and official SDKs simplify onboarding, but projects, regions, IAM authentication, versioned requests and model-specific capabilities add enterprise setup complexity.
Documentation
Good and broad, with official REST API, Python SDK, Node.js SDK, tutorials, Prompt Lab code generation and service-specific documentation. Exact behavior can vary by model, plan and deployment type.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+3
Cerebras
Cerebras Inference API
Free tier, pay-per-token developer access, and custom enterprise pricing
Current public model metadata lists model-specific input/output token prices; GPT OSS 120B is listed at approximately $0.35 per million input tokens and $0.75 per million output tokens. Free access has lower limits; enterprise pricing is custom.
SDKs
Official Python package: cerebras_cloud_sdk. Official TypeScript/Node.js package: @cerebras/cerebras_cloud_sdk. OpenAI-compatible Python and Node.js clients can also be configured with the Cerebras base URL. Other languages can use HTTP.
Ease of use
High for developers familiar with OpenAI-compatible APIs. The platform provides a conventional REST interface, bearer authentication, official Python and TypeScript SDKs, and OpenAI client compatibility.
Documentation
Good and technically focused. Documentation covers quickstart, API reference, models, streaming, tool use, structured outputs, rate limits, projects, service tiers, files, and batch processing. Some newer features are preview-only.
Streaming
Function calling
File uploads
Fine-tuning
Structured outputs
+1
Baidu
Baidu Qianfan API
Usage-based pricing by model and service; token-based inference, token plans, quota plans, and separate charges for selected tools, media, batch, or training services.
Prices vary by model and may use input/output token rates, image or video units, web-search calls, token plans, or TPM/RPM quota plans. The current model catalog exposes model-specific pricing.
SDKs
Official Qianfan Python SDK is available. Baidu Cloud SDK resources cover multiple languages. The current v2 documentation also supports the OpenAI Python client for compatible inference requests. Direct HTTPS integration is available for JavaScript and P
Ease of use
Moderate. The v2 API is straightforward for OpenAI-compatible chat requests, but Baidu-specific authentication, model availability, regional/account requirements, Agents, and fine-tuning add platform complexity.
Documentation
Good and extensive, with current Chinese-language API reference pages, model catalog documentation, SDK material, Agent documentation, and endpoint examples. English coverage is more limited.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+4
Databricks Data + AI Platform
Databricks Data and AI Developer Platform
Usage-based Databricks Units and cloud charges; model access may be pay-per-token, priority pay-per-token, provisioned throughput, or compute-based serving
Pricing varies by cloud, region, workspace tier, model, and serving mode. Foundation Model APIs can charge separately for input and output tokens; provisioned throughput and custom serving use hourly or compute-based charges.
SDKs
Official Databricks SDK for Python, Databricks SDKs and tooling for platform operations, Databricks OpenAI client integration, MLflow Deployments SDK, Databricks CLI, and OpenAI-compatible standard clients. Raw REST is available for any language.
Ease of use
Moderate. OpenAI-compatible interfaces make model calls familiar, but workspace configuration, Unity Catalog, endpoint permissions, regions, quotas, and Databricks authentication add operational complexity.
Documentation
Strong and extensive, with current product guides, REST references, SDK documentation, model availability tables, lifecycle policies, and cloud-specific pages. Exact behavior can vary by cloud and model.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+3
Allen Institute for Artificial Intelligence (Ai2)
OlmoEarth API
Request-based platform access; no public standard API pricing identified
OlmoEarth access is available by account request. No public per-token, per-request, or subscription API price was identified in the reviewed official documentation.
SDKs
No provider-specific official SDK was identified. The interactive API browser can generate client code for common languages and libraries; standard HTTP clients are suitable.
Ease of use
Moderate for developers familiar with REST and GeoJSON; API keys, interactive documentation, and generated client examples are provided, but access is account-based and the API is specialized for Earth-observation workflows.
Documentation
Good and current for authentication, area management, platform concepts, and API discovery. The interactive API browser provides endpoint schemas and generated examples, while some operational details such as quotas and pricing are not publicly specified.
Fine-tuning
Playground
NVIDIA AI
NVIDIA API Catalog and NIM
Hosted API Catalog endpoints are free for NVIDIA Developer Program prototyping within applicable limits; production self-hosted NIM generally requires NVIDIA AI Enterprise licensing priced by GPU.
NVIDIA Developer Program access is free for prototyping. NVIDIA AI Enterprise production licensing starts at $4,500 per GPU per year or approximately $1 per GPU per hour in cloud environments. Hosted endpoint pricing and limits can vary by endpoint.
SDKs
Official Python quickstarts and broad raw HTTP support. OpenAI-compatible Python and JavaScript clients can be configured with NVIDIA’s base URL. NVIDIA also provides first-party SDKs and toolkits across the broader CUDA, NeMo, NIM, Triton, and AI Enter
Ease of use
Easy for prototyping through build.nvidia.com, API-key generation, browser previews, generated code, and OpenAI-compatible HTTP patterns. Self-hosting requires NVIDIA GPU infrastructure, containers, model selection, and operational deployment work.
Documentation
Strong and extensive, with centralized API documentation, model-specific references, NIM deployment guides, support matrices, quickstarts, release notes, and NeMo ecosystem documentation. Details can vary by NIM version and model.
Streaming
Function calling
Assistants & agents
Fine-tuning
Image input
+2
Yandex AI
Yandex AI Studio API
Usage-based cloud billing
Prices vary by model and resource consumption, including input/output tokens, image or audio processing, batch jobs, dedicated instances, agents, and fine-tuning. Consult the current official price list for exact rates.
SDKs
Official Yandex Cloud SDKs support Node.js, Go, Python, Java, and .NET. A dedicated Yandex AI Studio SDK is available, and compatible Python and JavaScript OpenAI clients can be used with the OpenAI-compatible endpoint. No PHP-specific SDK is required for
Ease of use
Moderate. The OpenAI-compatible endpoint lowers migration effort, while native Yandex Cloud APIs require familiarity with folders, service accounts, IAM permissions, model URIs, and cloud billing.
Documentation
Good and actively maintained. Documentation covers REST, gRPC, OpenAI-compatible access, Responses API, SDKs, agents, files, fine-tuning, quotas, pricing, and tutorials. Capability details can vary by model and feature status.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+4
Z.ai
Z.ai API
Usage-based token pricing with prepaid resource packages, usage bundles, and enterprise or volume arrangements; separate subscription quotas apply to coding plans.
Public pricing is model- and account-dependent. Z.ai advertises usage bundles and prepaid API resources, while some business pricing is handled through sales. Coding plans are separate and currently advertise individual plans from $18 per month, subject t
SDKs
Official Python SDK via zai-sdk and official Java SDK are documented. OpenAI-compatible integration is also documented. JavaScript/TypeScript and PHP can use standard HTTP requests; a first-party JavaScript or PHP SDK was not verified in the current offic
Ease of use
Good. The API is OpenAI-compatible, uses conventional bearer-key authentication and chat-completions requests, and provides official Python and Java SDK documentation. Advanced capabilities require model-specific configuration.
Documentation
Good and broad. Current documentation covers quick starts, HTTP calls, Python and Java SDKs, API reference pages, model guides, streaming, tools, MCP, structured output, web search, agents, and specialized media APIs. Some quotas and account-specific acce
Streaming
Function calling
Assistants & agents
File uploads
Web search
+3
MiniMax
MiniMax Open Platform API
Usage-based pay-as-you-go with token, character, image, audio, video-second, request, and task-based pricing; separate Token Plan and credit options are also available.
MiniMax-M3 standard pricing is listed at $0.30 per million input tokens and $1.20 per million output tokens for requests up to 512k input tokens; larger-context and priority pricing is higher. Other modalities use separate unit prices.
SDKs
Official documentation supports the Anthropic SDK as the recommended language integration and the OpenAI SDK for Python and Node.js. Raw HTTP is available for PHP and other languages. Official MiniMax MCP implementations are available in Python and JavaSc
Ease of use
Good. Standard HTTP APIs are available, and language models can be accessed with the official OpenAI SDK or the recommended Anthropic SDK. Media APIs use native HTTP, multipart, asynchronous-task, or WebSocket patterns.
Documentation
Good and broad, with current API-overview, model, pricing, rate-limit, SDK, file, media, and server-tool documentation. Some capabilities are documented across separate endpoint pages and availability can vary by model or region.
Streaming
Function calling
File uploads
Web search
Image input
+1
Moonshot AI
Kimi API Platform
Pay-as-you-go token billing with separate pricing for model inference, web search, and batch processing
Input and output tokens are billed per usage; K3 supports cache-aware pricing, batch inference has separate lower-cost pricing, and file upload/extraction APIs are temporarily free.
SDKs
Official OpenAI Python and Node.js SDKs are supported through the OpenAI-compatible endpoints; official Anthropic SDKs are supported through the Anthropic-compatible Messages endpoint; raw HTTP works in any language.
Ease of use
Easy for developers familiar with OpenAI-compatible APIs; use an API key, change the base URL, select a current Kimi model, and send standard HTTP or SDK requests.
Documentation
Good and currently maintained, with quickstarts, model guides, API references, examples, migration notices, pricing documentation, tool workflows, file workflows, and troubleshooting guidance.
Streaming
Function calling
Assistants & agents
File uploads
Web search
+3
Microsoft Copilot
Microsoft Foundry API
Usage-based Azure consumption pricing with model-specific token, media, batch, deployment, and provisioned-throughput charges
Pricing varies by model, input/output tokens, region, deployment type, batch usage, and Provisioned Throughput Units. Consult the live Azure pricing calculator and Azure OpenAI pricing page for exact rates.
SDKs
Official OpenAI SDK support for Python and JavaScript/TypeScript, plus Microsoft Foundry SDKs for Python, JavaScript/TypeScript, C#, and Java. Microsoft also provides Azure identity libraries and language-specific Azure clients.
Ease of use
Moderate to high. The OpenAI-compatible Responses API is straightforward, while Azure resources, deployments, RBAC, quota, regions, and networking add enterprise configuration.
Documentation
Strong and extensive, with official quickstarts, REST references, SDK guidance, quota documentation, model lifecycle information, and privacy documentation. The breadth of Azure options requires careful endpoint and version selection.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+4
Amazon
Amazon Bedrock API
Usage-based AWS pricing determined by model, modality, tokens or other model-specific units, Region, service tier and optional capacity commitments.
Model-specific on-demand pricing is published by AWS. Bedrock also supports Standard, Flex, Priority, Batch and Reserved or Provisioned Throughput options for supported models. There is no single universal API subscription price.
SDKs
Official AWS SDK support includes Python Boto3, JavaScript or TypeScript AWS SDK v3, Java, Go, .NET, C++, Ruby, PHP and other AWS SDK languages.
Ease of use
Moderate. Converse provides a unified interface for supported models, while AWS credentials, IAM policies, Regions, model access and model-specific request formats add operational complexity.
Documentation
Extensive and technically detailed, with separate user guides, API references, quotas, model-availability pages, SDK examples and service-specific documentation.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+4
Cohere
Cohere API
Usage-based pricing by model and endpoint; token-based for generative models, search-unit pricing for reranking, token-based pricing for embeddings, and separate deployment pricing for managed dedicated infrastructure.
Trial API keys are free but rate-limited. Production usage is metered. Current model pricing varies; Command A+ is listed as free until applicable rate limits are reached, while other models and endpoints have published per-token or per-unit prices.
SDKs
Official SDKs for Python, TypeScript, Java and Go. REST/HTTP access is available for other languages, including PHP.
Ease of use
Good. The REST API is conventional, the v2 Chat API uses a clear messages structure, and official SDKs provide typed clients and helpers. Advanced RAG and tool workflows require more application-side orchestration.
Documentation
Good to very good. Cohere provides current guides, API references, SDK repositories, model pages, migration guidance, rate-limit documentation, cookbooks and deprecation notices.
Streaming
Function calling
Assistants & agents
File uploads
Image input
+2
Mistral AI
Mistral API
Usage-based pricing by model and service; token-based for most text models, with separate pricing units for OCR, audio, tools, batch, priority, cached input, and fine-tuning products.
Text models are generally priced per million input and output tokens. Batch processing is listed at 50% of standard pricing, cached input can receive discounts of up to 90%, and regional inference adds 10%. Current prices vary by model and service.
SDKs
Official Python SDK package mistralai and official TypeScript/JavaScript SDK package @mistralai/mistralai. Third-party SDKs are available for other languages.
Ease of use
High for standard chat requests because the REST API follows a familiar chat-completions structure and official Python and TypeScript SDKs are available. Advanced agents, libraries, regional routing, and tool workflows require additional concepts.
Documentation
Good and broad, with current quickstarts, API reference pages, SDK documentation, cookbooks, model lifecycle information, regional-inference guidance, and explicit deprecation notices. Some legacy fine-tuning material remains marked deprecated.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+4
DeepSeek
DeepSeek API
Usage-based token pricing with separate cached-input, uncached-input, and output rates; peak and off-peak pricing applies.
deepseek-flash: $0.003–$0.006 cached input, $0.15–$0.30 uncached input, and $0.60–$1.20 output per 1M tokens. deepseek-v4-pro: $0.022–$0.044 cached input, $0.66–$1.32 uncached input, and $1.98–$3.96 output per 1M tokens.
SDKs
Official examples use the OpenAI Python and JavaScript SDKs configured with DeepSeek's base URL. Anthropic-compatible integrations are also documented. PHP can use raw HTTP; no official DeepSeek PHP SDK was verified.
Ease of use
High for developers familiar with OpenAI-compatible APIs. The platform uses standard HTTP, Bearer authentication, JSON requests, and compatible Python and JavaScript SDK configuration.
Documentation
Good and actively updated, with API reference pages, quick starts, feature guides, pricing, error codes, files, vision, Responses API, and agent integration documentation.
Streaming
Function calling
Assistants & agents
File uploads
Image input
+1
Qwen
Qwen API
Primarily pay-as-you-go token billing; eligible services also support batch pricing, prepaid or subscription plans, and dedicated deployment billing.
Prices vary by model, region, deployment scope, context length, caching, and thinking mode. Current examples include Qwen3.8-Max at $2 per 1 million input tokens and $6 per 1 million output tokens in the Singapore international scope. Batch inference for
SDKs
Official DashScope SDK plus compatibility with official OpenAI SDKs for Python and JavaScript/TypeScript through a custom base URL. Raw HTTP is also supported.
Ease of use
Good. OpenAI-compatible endpoints allow existing OpenAI integrations to be adapted by changing the API key, base URL, and model name. Regional workspace setup and model availability require additional configuration.
Documentation
Good and actively maintained. Documentation covers quickstarts, OpenAI-compatible APIs, DashScope, model-specific references, pricing, regions, tools, structured outputs, files, fine-tuning, privacy, and migration topics.
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+4
xAI
xAI API
Usage-based pricing billed by tokens, tool invocations, images, video duration, audio duration or characters depending on the service; prepaid credits and enterprise invoicing are available.
The documented Grok 4.7 price is $2 per 1M short-context input tokens and $6 per 1M output tokens. Long-context, cached-input, reasoning, media, voice and server-side tool prices vary by model and operation.
SDKs
Official Python SDK xai-sdk with native gRPC support; OpenAI Python and JavaScript SDK compatibility through the xAI base URL; JavaScript support through @ai-sdk/xai and Vercel AI SDK; gRPC and raw HTTP are also supported.
Ease of use
High for developers familiar with OpenAI-compatible APIs. API keys are created in the Console, the REST API follows familiar schemas, and official or compatible SDK options are available.
Documentation
Strong and actively maintained, with quickstarts, model catalog, REST and gRPC references, capability guides, pricing, rate limits, security FAQs and code examples. Some APIs and integrations are marked early access or may change.
Streaming
Function calling
Assistants & agents
File uploads
Web search
+3
Meta AI
Meta Model API
Pay-as-you-go usage pricing
Muse Spark Standard: $1.25/M input tokens, $0.15/M cached input tokens, $4.25/M output tokens. Contributor: $0.10/M input, $0.002/M cached input, $0.20/M output. Muse Image: $0.01/image. Voice Transcribe: $0.18/hour. SAM 3.1: $2.50/1,000 images or $0.20/1
SDKs
Official Python package llama-api-client and official TypeScript package llama-api-client; Meta also documents compatibility with the OpenAI SDK and Anthropic SDK.
Ease of use
High for developers familiar with OpenAI-compatible or Anthropic-compatible APIs; set a base URL and bearer key, then use the Responses, Chat Completions, or Messages format.
Documentation
Good and improving; current documentation includes quickstarts, API references, capability guides, pricing, rate limits, SDK guidance, cookbooks, and help-center articles.
Streaming
Function calling
Assistants & agents
File uploads
Web search
+3
Claude
Claude API
Usage-based pay-as-you-go token pricing, with separate pricing for some tools, caching, and service tiers
Pricing varies by model and is billed per input and output token. Batch processing is generally 50% lower than standard pricing; prompt caching and web search have separate pricing rules.
SDKs
Official SDKs for Python, TypeScript, C#, Go, Java, PHP, and Ruby; official ant CLI; direct REST and cURL support.
Ease of use
High for standard Messages API integrations; official SDKs, Workbench, typed responses, streaming helpers, retries, and extensive examples reduce setup effort.
Documentation
Strong and detailed official documentation with API reference pages, language-specific SDK guides, feature guides, migration notes, and operational documentation.
Streaming
Function calling
Assistants & agents
File uploads
Web search
+3
OpenAI
OpenAI API
Usage-based pay-as-you-go pricing, primarily per million input, cached-input, cache-write, and output tokens, with separate charges for some tools and specialized services.
Current pricing varies by model and processing mode. As listed on the current pricing page, flagship examples include GPT-6 Astra at $10 per 1M short-context input tokens and $50 per 1M output tokens under Standard processing; GPT-6 Sol at $2 input and $1
SDKs
Official SDKs and client libraries are available for JavaScript/TypeScript, Python, .NET, Java, Go, and Ruby. The Agents SDK is officially available for TypeScript and Python. PHP clients are community-maintained, so raw HTTPS is the portable PHP approach
Ease of use
High for common use cases. The Responses API has a unified request format, official SDKs, automatic environment-variable authentication, direct HTTPS access, Playground testing, and helper properties such as output_text. Advanced multi-step tool workflows
Documentation
High. Current official documentation includes a quickstart, API reference, guides for Responses, tools, structured outputs, files, rate limits, production, privacy, latency, SDKs, models, and migrations. Deprecated APIs are identified, although the rapidl
Streaming
Function calling
Assistants & agents
File uploads
Fine-tuning
+4