llama.cpp

LLM inference in C/C++ on any hardware.

Portable inference engine and GGUF format powering most local-LLM apps, from laptops to phones.

Vendor
ggml
Category
Inference & model hosting
Pricing
Open source
Open source
Yes
License
MIT
Platforms
CPU, Apple Silicon, CUDA, Vulkan
Launched
2023
Website
github.com/ggml-org/llama.cpp

Features

Best for

More inference & model hosting

Links