vLLM
VerifiedHigh-throughput open-source LLM inference and serving engine.
vLLM is an Apache-2.0 serving engine known for high throughput via PagedAttention and continuous batching, with an OpenAI-compatible server. It is a de facto standard for self-hosting open models on GPUs.
License
Open source (Apache-2.0)
Deployment
Self-hosted
Pricing
Free and open-source; you provide the GPUs.
SDKs / languages
Python, Any (OpenAI API)
Founded
2023
Strengths
- High throughput and efficiency
- OpenAI-compatible endpoint
- Broad model + hardware support
Limitations
- You operate the GPUs
- Tuning needed for best performance
Alternatives
Baseten
Deploy and scale model inference in production.
Fireworks AI
Fast managed inference for open models and fine-tunes.
Groq
Ultra-low-latency inference on custom LPU hardware.
LM Studio
Desktop app to run local LLMs with a GUI and local server.
Modal
Serverless GPU compute for AI workloads.
Ollama
Run open LLMs locally with one command.
Facts verified 19 July 2026. Neutral summary, not an endorsement; verify current details with the vendor.