llama.cpp
LLM inference in C/C++ on any hardware.
Portable inference engine and GGUF format powering most local-LLM apps, from laptops to phones.
- Vendor
- ggml
- Category
- Inference & model hosting
- Pricing
- Open source
- Open source
- Yes
- License
- MIT
- Platforms
- CPU, Apple Silicon, CUDA, Vulkan
- Launched
- 2023
- Website
- github.com/ggml-org/llama.cpp
Features
- GGUF quantization
- Runs on CPU
- Server mode
Best for
- Local & on-device
More inference & model hosting
- Ollama — Ollama. Run open models locally with one command.
- OpenRouter — OpenRouter. One API for hundreds of models.
- vLLM — vLLM project. High-throughput LLM serving engine.
- Amazon Bedrock — AWS. Foundation models on AWS.
- Vertex AI — Google Cloud. Google Cloud’s AI platform.
- LM Studio — Element Labs. Desktop app to discover and run local LLMs.
- Azure AI Foundry — Microsoft. Build and run AI apps on Azure.
- Groq — Groq. Ultra-fast inference on custom LPUs.