TensorRT-LLM
Partial infoNVIDIA's optimized inference library for LLMs on GPUs.
TensorRT-LLM is NVIDIA's Apache-2.0 library for compiling and running LLMs with maximum performance on NVIDIA GPUs, often paired with Triton Inference Server. It targets peak throughput and latency on NVIDIA hardware.
License
Open source (Apache-2.0)
Deployment
Self-hosted
Pricing
Free and open-source; requires NVIDIA GPUs.
SDKs / languages
Python, C++
Strengths
- Peak performance on NVIDIA GPUs
- Advanced quantization (FP8/INT4)
- Pairs with Triton
Limitations
- NVIDIA-only
- Build/compile step is complex
Alternatives
Baseten
Deploy and scale model inference in production.
Fireworks AI
Fast managed inference for open models and fine-tunes.
Groq
Ultra-low-latency inference on custom LPU hardware.
LM Studio
Desktop app to run local LLMs with a GUI and local server.
Modal
Serverless GPU compute for AI workloads.
Ollama
Run open LLMs locally with one command.
Facts verified 19 July 2026. Neutral summary, not an endorsement; verify current details with the vendor.