Ollama is a tool for downloading and running open-weight large language models on your own computer or server. It wraps model fetching, quantized inference, and API serving into a single command (e.g. ollama run qwen3.6), so models like Qwen, DeepSeek, GLM, Gemma, and Kimi can run locally as easily as installing an app, and any application can talk to them through a local OpenAI-compatible endpoint. Running models on your own hardware is free and requires no account; Ollama Inc. also operates a per-token cloud model service, which is entirely optional. The project is open source under the MIT license and, as of 2026-10-02, has roughly 182k stars on GitHub, making it one of the most widely used local LLM runtimes.

At a Glance
- URL: https://ollama.com
- Type: Local LLM runtime / open-source developer tool (MIT license)
- Cost: Free and unlimited for local models; cloud models billed per token (Free / Pro / Max / Team / Enterprise tiers)
- Account: Not required for local use; an account and API key are only needed for cloud models
- Interface: Official desktop app (macOS/Windows) + CLI + local REST API; website and docs in English
- Platforms: macOS 14+, Windows 10 22H2+, Linux, Docker
Background
The ollama/ollama repository was created in June 2023 and is maintained by Ollama Inc. under the MIT license. As of 2026-10-02 it has about 182k stars and 18.1k forks; the latest release is v0.35.0, published on 2026-09-28, with commits landing the same day we checked, indicating active maintenance.

According to a July 2026 announcement on the official blog, Ollama was then serving 8.9 million developers and had raised $88M from investors including Benchmark, Theory Ventures, 8VC, and Y Combinator. Its inference backend was originally built on llama.cpp; since mid-2026, Apple Silicon builds also ship an MLX engine, which Ollama says made Gemma 4 up to about 90% faster with coding agents via multi-token prediction (official blog, June 2026).
Core Capabilities
One command to run a model. After installation, ollama run <model> downloads the model and drops you into a chat. ollama pull, ollama ps, and ollama rm handle downloading, inspecting what's loaded (including GPU/CPU split), and deleting models.
Local OpenAI-compatible API. Ollama exposes a REST API at http://localhost:11434, implementing OpenAI's Chat Completions and Responses APIs plus an Anthropic Messages compatibility layer. Existing OpenAI SDK clients only need a new base_url to switch to local models. Official Python and JavaScript SDKs are also available.
Customization via Modelfile. A Modelfile lets you build on an existing model with your own system prompt, sampling parameters, and chat template, then package it with ollama create. GGUF weights can be imported the same way.
Coding agents and desktop apps. Commands like ollama launch claude, ollama launch codex, and ollama launch opencode wire coding agents such as Claude Code, Codex CLI, OpenCode, and OpenClaw directly to Ollama models. The macOS/Windows desktop app can also connect Claude Desktop and ChatGPT Desktop (Codex mode) to Ollama. The README maintains a large community integrations list covering Open WebUI, Continue, LangChain, LlamaIndex, Dify, and many others.
Installation and Platform Requirements
- macOS: Requires macOS Sonoma (14) or newer; Apple M-series chips get CPU+GPU acceleration, while Intel machines are CPU-only.
- Windows: Requires Windows 10 22H2 or newer (Home/Pro); NVIDIA GPUs need driver 551.61 or newer. Install via
irm https://ollama.com/install.ps1 | iexor the setup package. - Linux:
curl -fsSL https://ollama.com/install.sh | shinstalls Ollama as a systemd service; there is no official GUI. - Docker: The official
ollama/ollamaimage is published on Docker Hub.
Plan for disk usage: the documentation notes that model files can run from tens to hundreds of gigabytes, and the storage location may need to be moved if the home directory is short on space.
Model Library and Cloud Service
The model library is sorted by popularity and recency. Entries carry capability tags such as vision, tools, thinking, embedding, decision, and cloud, along with parameter size, context length, and cumulative pull counts. As of 2026-10-02, popular entries include qwen3.6 (6.9M pulls), qwen3.8 (3M pulls), glm-5.3-flash, and deepseek-v4.1-flash.

Since 2025, Ollama also hosts cloud models: entries with a :cloud suffix (e.g. glm-5.3:cloud) run on Ollama's own infrastructure, so machines without enough local capacity can still call frontier-scale open-weight models — 753B parameters with a 1M-token context in GLM-5.3's case. Cloud calls go through https://ollama.com/v1 (or are relayed by a signed-in local server) and require an account with an API key.

Pricing (per the pricing page, as of 2026-10-02):
- Local models: always free and unlimited ("Running models on your own hardware is always unlimited").
- Free $0: Includes starter usage for a smaller set of models; purchased credits unlock all cloud models, billed per token.
- Pro 200/yr): $60 of monthly usage credits and access to larger Pro models.
- Max $100/mo: $300 of monthly credits and 10 concurrent requests.
- Team $500/mo (early access): Unlimited users sharing $1,000 of monthly credits, with centralized billing.
- Enterprise: Custom pricing with model access controls and budgets.
Privacy and Data Handling
Per the official FAQ: prompts and data from local runs never leave your machine. For cloud models, Ollama processes prompts and responses to provide the service but states it does not store or log that content, never trains on it, collects only basic account info and limited usage metadata, and does not sell data. Cloud models are hosted in the US, Europe, and Singapore. Cloud features can be disabled entirely for local-only operation. These commitments are Ollama's own statements; review the privacy policy before sending sensitive workloads to the cloud.
Good Fits
- Developers who want a self-contained local LLM environment exposing an OpenAI-compatible endpoint to existing apps.
- Individuals or teams with privacy requirements running open models on-prem or offline.
- Driving coding agents (Claude Code, Codex, OpenCode, etc.) with open-weight models to control token costs.
- Packaging customized models with fixed system prompts and parameters via Modelfile for redistribution.
- Users with modest local hardware who still want access to frontier open-weight models via metered
:cloudvariants.
Limitations
- Local inference demands substantial VRAM/RAM and disk space; model files range from tens to hundreds of GB, and low-memory machines are limited to small models.
- The default context window is 4096 tokens; long-context use requires manually raising
OLLAMA_CONTEXT_LENGTHornum_ctx, which increases memory pressure. - Intel Macs run on CPU only and are slow; Linux has no official GUI, so expect the command line.
- Cloud models are metered per token, require an account, and route content through Ollama's servers (with no-store/no-train claims made by Ollama itself).
- Output quality depends entirely on the chosen open model; Ollama runs and distributes models but does not vouch for their capabilities.
References
- Ollama website and pricing page (checked 2026-10-02)
- GitHub repository ollama/ollama (star and release data verified via the GitHub API on 2026-10-02)
- Official docs: Quickstart, OpenAI compatibility, FAQ, macOS, Windows
- Official blog: Ollama: all aboard open models, Ollama's transparent pricing







