Ollama
Run open models locally with one command.
The simplest way to download and run open-weight models on macOS, Windows and Linux, with an OpenAI-compatible API.
- Vendor
- Ollama
- Category
- Inference & model hosting
- Pricing
- Open source
- Open source
- Yes
- License
- MIT
- Platforms
- macOS, Windows, Linux
- Launched
- 2023
- Website
- ollama.com
Features
- One-line model pulls
- Local API
- GPU acceleration
Best for
- Local & on-device
Works with
- gpt-oss-20b — OpenAI. Compact open-weight reasoning model that runs on consumer hardware.
- Gemma 4 31B — Google. Open-weight dense multimodal model from Google DeepMind.
- Qwen3.6 35B A3B — Alibaba (Qwen). Tiny-active-parameter open MoE — a local-inference favorite.
More inference & model hosting
- OpenRouter — OpenRouter. One API for hundreds of models.
- vLLM — vLLM project. High-throughput LLM serving engine.
- llama.cpp — ggml. LLM inference in C/C++ on any hardware.
- Amazon Bedrock — AWS. Foundation models on AWS.
- Vertex AI — Google Cloud. Google Cloud’s AI platform.
- LM Studio — Element Labs. Desktop app to discover and run local LLMs.
- Azure AI Foundry — Microsoft. Build and run AI apps on Azure.
- Groq — Groq. Ultra-fast inference on custom LPUs.