Modal
Partial infoServerless GPU compute for AI workloads.
Modal is a serverless platform for running Python code and AI workloads on GPUs, billed per-second, commonly used to self-serve LLM inference, batch jobs, and fine-tuning without managing infrastructure. You bring the serving code (e.g. vLLM).
License
Proprietary
Deployment
Managed
Pricing
Per-second GPU/CPU compute billing.
SDKs / languages
Python
Founded
2021
Strengths
- Serverless GPUs, per-second billing
- Great for bursty/batch workloads
- Full control over serving code
Limitations
- You build the serving layer
- Python-centric
Alternatives
Baseten
Deploy and scale model inference in production.
Fireworks AI
Fast managed inference for open models and fine-tunes.
Groq
Ultra-low-latency inference on custom LPU hardware.
LM Studio
Desktop app to run local LLMs with a GUI and local server.
Ollama
Run open LLMs locally with one command.
Replicate
Run and deploy any model via a simple API.
Facts verified 19 July 2026. Neutral summary, not an endorsement; verify current details with the vendor.