vLLM

High-throughput LLM serving engine.

Open-source inference engine with PagedAttention and continuous batching; the de-facto standard for self-hosted serving.

Vendor
vLLM project
Category
Inference & model hosting
Pricing
Open source
Open source
Yes
License
Apache 2.0
Platforms
Linux, GPU, TPU
Launched
2023
Website
vllm.ai

Features

Best for

More inference & model hosting

Links