SGLang

High-throughput serving engine for LLMs and VLMs.

Fast open-source serving with RadixAttention prefix caching, speculative decoding, structured-output acceleration and large-scale expert parallelism for MoE models.

Vendor
LMSYS / SGLang community
Category
Inference & model hosting
Pricing
Open source
Open source
Yes
License
Apache 2.0
Platforms
Linux, NVIDIA, AMD, TPU
Launched
2024
Website
docs.sglang.ai

Features

Best for

More inference & model hosting

Links