SGLang
High-throughput serving engine for LLMs and VLMs.
Fast open-source serving with RadixAttention prefix caching, speculative decoding, structured-output acceleration and large-scale expert parallelism for MoE models.
- Vendor
- LMSYS / SGLang community
- Category
- Inference & model hosting
- Pricing
- Open source
- Open source
- Yes
- License
- Apache 2.0
- Platforms
- Linux, NVIDIA, AMD, TPU
- Launched
- 2024
- Website
- docs.sglang.ai
Features
- RadixAttention caching
- Expert parallelism
- Fast JSON decoding
Best for
- High volume / low cost
- Local & on-device
- Real-time / low latency
More inference & model hosting
- Ollama — Ollama. Run open models locally with one command.
- OpenRouter — OpenRouter. One API for hundreds of models.
- vLLM — vLLM project. High-throughput LLM serving engine.
- llama.cpp — ggml. LLM inference in C/C++ on any hardware.
- Amazon Bedrock — AWS. Foundation models on AWS.
- Vertex AI — Google Cloud. Google Cloud’s AI platform.
- LM Studio — Element Labs. Desktop app to discover and run local LLMs.
- Azure AI Foundry — Microsoft. Build and run AI apps on Azure.