Ollama

Local AI model runtime

Ollama lets users download, run, customize, and integrate open AI models on macOS, Windows, Linux, servers, and Docker environments. It provides a desktop application, command-line interface, local REST API, OpenAI-compatible API access, model library, and optional cloud-hosted model access.

Company Ollama Inc.
Free plan Yes
Paid plans from $20/month
Ease of use Moderate

What you can do with Ollama

Key features
✓
Run open models locally

Download and execute supported language, vision, embedding, and other open models on local hardware.

✓
Chat with models

Use the desktop application or command line to interact with downloaded models.

✓
Analyze files and images

Drag PDFs and code files into the desktop app and send images to compatible vision models.

✓
Connect models to applications

Use the REST API, OpenAI-compatible API, Python library, or JavaScript library in software projects.

✓
Support coding agents

Launch and connect compatible developer tools and coding agents such as Claude Code, Codex, OpenCode, and VS Code.

✓
Generate structured and tool-enabled responses

Use structured outputs and tool calling with models and API clients that support them.

✓
Use cloud-hosted models

Run larger models through Ollama's cloud service when local hardware is insufficient.

How Ollama works

Install Ollama on a supported computer, select or download a model from the library, and interact with it through the desktop app, terminal, or local API. Developers can then connect the running model to applications, coding agents, or other tools using Ollama's supported interfaces.

INPUTS
Text promptsImagesPDF filesCode filesAPI requests
OUTPUTS
TextCodeStructured JSONEmbeddingsData analysis

Who Ollama is for

BEST FOR

Developers and technical users who want to run open models locally, keep sensitive prompts on their own hardware, experiment with different model families, build AI applications, or connect models to coding agents.

LESS SUITED FOR

Users seeking a polished, fully managed consumer chatbot with no installation, no hardware requirements, no command-line usage, and no responsibility for choosing or operating models.

Strengths & limitations

+ Strengths

  • Local execution can keep prompts and model interactions on the user's device
  • broad support for open model families
  • simple command-line workflow
  • desktop chat application for macOS and Windows
  • local REST and OpenAI-compatible APIs
  • useful developer integrations
  • support for vision, embeddings, structured outputs, and tool calling
  • optional cloud access for larger models
  • permissive MIT license for the core open-source repository.

– Limitations

  • Local performance depends heavily on available RAM, GPU, storage, and model size
  • setup and model management are more technical than hosted chatbot services
  • model quality and capabilities vary substantially between downloaded models
  • local usage requires users to supply and maintain hardware
  • cloud usage introduces token billing and account requirements
  • Ollama is not a full agent builder, workflow automation platform, or general-purpose collaborative workspace.

Pricing & access

FREE ACCESS Free plan available

Local model usage is available without a subscription. Ollama's cloud service also has a Free plan with starter usage credits, access to starter models, and pay-as-you-go usage credits for additional access.

PAID ACCESS $20/month

Local models can be run without subscription charges, subject to the user's hardware and storage. Cloud pricing includes Free, Pro at $20/month, Max at $100/month, Team at $500/month, and custom Enterprise pricing. Cloud model usage is measured per million tokens. Annual billing is available for Pro at $200/year.

FREE TRIAL No free trial listed
USAGE LIMITS Plan limits apply

Cloud usage is limited by included monthly usage credits, model-specific token rates, and concurrency limits. The current pricing page lists one concurrent request for Free, three for Pro, and ten for Max and Team. Local model execution is not metered by Ollama, but it depends on available system resources.

Platforms & access

– Web app
– Mobile app
✓ Desktop app
– Browser extension
✓ API
– Embeddable

macOS; Windows 10 or later; Linux; Docker; local servers; REST API; OpenAI-compatible API

Product format: standalone

Product specs

Standard features
✓ Web access
✓ File upload
– Memory
– Custom agents
– Scheduled automation
– Knowledge base
– Website ingestion
– Code execution
– Computer actions
✓ Integrations
– Webhooks
– MCP support
– Bring your own key
✓ Model selection
✓ Collaboration
– Shared workspace
✓ Admin controls
– Role permissions
✓ Analytics
– Templates
– No-code
– Project workspace
– Brand tools
– Performance scoring

Ollama supports model downloading and management, local and cloud inference, multimodal models, document and image input in the desktop app, embeddings, structured outputs, tool calling, OpenAI-compatible APIs, Python and JavaScript libraries, Docker deployments, and integrations with coding agents. Users remain responsible for model licenses and hardware capacity.

Categories & capabilities

Browse similar tools

Integrations & models

INTEGRATIONS Connected workflows

Ollama provides integrations and launch workflows for coding agents and developer tools including Claude Code, Codex, OpenCode, VS Code, OpenClaw, Pi, n8n, and other compatible applications. It also supports Python and JavaScript libraries, REST access, OpenAI-compatible APIs, Docker, LangChain, and LlamaIndex integrations.

MODELS Models used

Model choice varies. The official model library includes publicly documented models from providers such as Meta, Google, DeepSeek, Qwen, Mistral, Microsoft, NVIDIA, OpenAI, IBM, and others.

Privacy & data Data handling, AI training, retention and security
↓
Data handling

Local execution keeps prompts, responses, model interactions, and locally processed content on the user's device according to Ollama's privacy policy. Cloud-hosted requests are processed transiently to provide the service. Ollama states that it does not sell personal information and collects account, device, download, diagnostic, and limited usage metadata.

AI training

Ollama states that it does not use user inputs or outputs to train AI models. The pricing and privacy materials state that cloud prompts and responses are not logged or trained on.

Data retention

Locally processed prompts and responses are not collected by Ollama. For cloud-hosted models, prompt and response content is processed transiently and is not stored beyond the time required to fulfill the request according to the privacy and pricing materials. Account information, billing records, support communications, and operational metadata may be retained under the privacy policy.

Security

The local runtime can keep model interactions on the user's own device. Cloud models are hosted primarily in the United States, with Europe and Singapore used for additional capacity. Ollama states that cloud hosting partners are required to provide no-logging, no-training, and zero-data-retention policies. Enterprise includes custom security questionnaires.

About Ollama

Ollama gives developers, researchers, and technical users a practical way to run open AI models on their own computers. It combines a model library, desktop chat app, command-line workflow, local REST API, and OpenAI-compatible endpoints, while also offering cloud models when local hardware is not sufficient.

What is Ollama?

Ollama is a local AI model runtime rather than a conventional hosted chatbot. After installing it on a supported computer or server, users can download models, run them through a desktop application or terminal, and connect them to other software through an API. The core appeal is control: users choose the model, provide the hardware, and can keep locally processed prompts and responses on their own device.

Ollama also provides optional cloud-hosted models. This gives users a way to work with larger models when their computer does not have enough memory or processing capacity, but cloud use introduces accounts, usage credits, token-based billing, and different data-handling considerations.

What people use Ollama for

Running open models locally

Ollama's main workflow is downloading models from its library and running them on a Mac, Windows PC, Linux system, Docker deployment, or local server. The available model families come from a range of providers, including Meta, Qwen, DeepSeek, and Mistral AI. Capabilities and quality depend on the selected model rather than Ollama alone, so model choice, licensing, size, and hardware all matter.

Building applications around models

Developers can use Ollama's local REST API, OpenAI-compatible API, Python library, or JavaScript library to add model responses to their own applications. Common uses include private assistants, retrieval-augmented generation systems, document summarization, code tools, image-aware applications, and experiments with embeddings. Structured outputs and tool calling are also available for supported models and client workflows.

Supporting coding workflows

Ollama can serve as the model layer for coding tools and developer workflows. Its integrations and launch workflows include tools such as Claude Code, Codex, OpenCode, VS Code, OpenClaw, Pi, and n8n. This makes it useful for developers who want to test or operate models locally rather than sending every coding request to a managed AI service.

Working with files and images

The desktop application supports chat and file-based interactions, including PDFs and code files. Images can be sent to models that support vision inputs. These features are useful for document review, visual question answering, and code or file analysis, but they should not be confused with a dedicated document-management or research platform: the available behavior depends heavily on the selected model and the local setup.

How Ollama fits into an AI workflow

A typical workflow starts with installing Ollama, selecting a model, and downloading it from the model library. Users can then chat in the desktop application, issue commands from the terminal, or send requests to the local API. Developers often place Ollama behind an application interface, coding tool, or framework such as LangChain or LlamaIndex.

This makes Ollama different from tools such as LM Studio, which also focuses on running models locally, and from general-purpose hosted assistants. Ollama is primarily infrastructure and a model access layer: its value is in model management, local inference, APIs, and integration rather than in providing a single polished assistant with a fixed set of consumer features.

Platforms and technical requirements

Ollama is available for macOS, Windows 10 or later, and Linux. It can also be deployed with Docker or used on local servers. The desktop application is available on macOS and Windows, while Linux is primarily distributed through installation scripts and command-line or server packages.

Local performance depends on the computer. RAM, GPU capability, storage, and model size can substantially affect whether a model runs comfortably and how quickly it responds. Users are responsible for downloading models, maintaining storage, and checking the relevant model licenses. This hardware responsibility is a central difference from fully managed AI services.

Pricing and cloud access

Running models locally does not require an Ollama subscription, although it does require suitable hardware and storage. Ollama also offers a cloud Free plan with starter usage credits. Paid cloud access starts with Pro at $20 per month; the research lists Max at $100 per month, Team at $500 per month, and custom Enterprise pricing. Cloud usage is measured using published per-million-token rates, and monthly plans include usage credits.

Cloud limits vary by plan and model. The current pricing information lists one concurrent request for Free, three for Pro, and ten for Max and Team. Pro annual billing is listed at $200 per year. Because cloud prices, credits, and model availability can change, users with predictable or high-volume workloads should check the current pricing page rather than treating a monthly plan as unlimited access.

Privacy and data considerations

For local inference, Ollama states that prompts, responses, model interactions, and locally processed content remain on the user's device. This can make the local runtime a practical option for organizations that do not want sensitive material sent to a third-party inference service, provided the device and surrounding application environment are managed appropriately.

Cloud requests are handled differently. Ollama states that cloud prompts and responses are processed transiently, are not used to train AI models, and are not retained beyond the time needed to provide the service. Account, billing, diagnostic, download, and operational metadata may still be collected or retained under the privacy policy. Cloud hosting is primarily in the United States, with Europe and Singapore used for additional capacity.

Strengths and limitations

  • Strong fit for local control: Ollama makes it relatively straightforward to run open models on personal hardware or a private server.
  • Useful developer interfaces: The REST API, OpenAI-compatible endpoints, Python and JavaScript libraries, Docker support, and integrations make it practical to embed models in software.
  • Broad model choice: Users can experiment with model families from multiple providers instead of being locked to one hosted model.
  • Flexible deployment: The same general runtime can support desktop experimentation, terminal use, APIs, coding tools, and cloud access.
  • Hardware dependence: Response speed and model size are constrained by available RAM, GPU resources, storage, and system configuration.
  • More technical than a managed chatbot: Users must install the runtime, choose models, manage downloads, and understand the limitations and licenses of those models.
  • Variable output quality: Ollama provides the runtime, not a guarantee of consistent model quality. Results differ substantially across model families, sizes, and configurations.
  • Limited collaboration and automation: It is not primarily a shared workspace, full agent builder, or general workflow automation platform.

Who should use Ollama?

Ollama is a good fit for developers, AI engineers, researchers, privacy-conscious users, and technical teams that want to test open models, build applications, power coding workflows, or keep routine inference on their own infrastructure. It is especially useful when API compatibility and local control matter more than having a fully managed user experience.

It is a weaker fit for someone who simply wants an immediately available consumer chatbot, managed storage for collaborative work, or reliable high-end performance without maintaining hardware. In those cases, a hosted assistant or specialized AI application may involve less setup. Ollama is best understood as a flexible local model platform with a convenient interface, not as a finished replacement for every category of AI software.

Ollama is a local AI runtime and desktop application for downloading and running open models, connecting them to software and coding tools, and optionally using cloud-hosted models when local hardware is insufficient.

Answers to Frequently Asked Questions

Can Ollama run AI models locally without a subscription?
Yes. Running models locally with Ollama does not require an Ollama subscription, but users need suitable hardware, storage, and a model that their computer can handle. Ollama also offers optional cloud-hosted models through paid and free plans.
What is Ollama used for?
Ollama is used to download and run AI models locally on macOS, Windows, Linux, Docker deployments, and local servers. It also provides APIs and libraries for building private assistants, coding tools, document summarizers, retrieval-augmented generation systems, and other AI applications.
Does Ollama send local prompts and responses to the cloud?
For local inference, Ollama states that prompts, responses, model interactions, and locally processed content remain on the user's device. Cloud requests are handled separately and may involve account, billing, diagnostic, download, and operational metadata under Ollama's privacy policy.
How much does Ollama cloud access cost?
Ollama offers a cloud Free plan with starter usage credits. Paid access starts with Pro at $20 per month, while the listed Max and Team plans cost $100 and $500 per month respectively; Enterprise pricing is custom. Cloud usage is measured through token-based rates and plan credits, so users should check the current pricing page for the latest limits and model availability.
What hardware does Ollama require?
Ollama runs on macOS, Windows 10 or later, and Linux, and it can also be deployed with Docker or on local servers. Performance depends on available RAM, GPU capability, storage, model size, and system configuration. Larger models generally require more memory and processing capacity.