vLLM
High-throughput, memory-efficient LLM inference and serving engine.
Official website / GitHub ↗ Visit
What is this technology for?
Its role in production environments and why it is taught in our curriculums.
vLLM is a high-throughput, memory-efficient inference engine for large language models, with PagedAttention and continuous batching — the reference for production LLM serving.
What you will learn
The hands-on skills you will gain on this technology in our curriculums.
- Serve LLMs with vLLM on OpenShift AI
- Configure quantization and batching for throughput
- Monitor and scale inference endpoints
- Tune memory efficiency with PagedAttention
Latest News & Ecosystem Updates
Recent innovations, major releases, and key industry milestones in the ecosystem.
vLLM v0.9: Native Speculative Decoding & Advanced Reasoning Model Support
50% reduction in time-per-output-token for complex reasoning models and broad support for modern AI hardware.
Standardization of vLLM as the Default Serving Engine on OpenShift AI
Streamlined production deployments featuring automatic autoscaling triggered by request queue depth.
Featured in the curriculum
Find this technology in the following modules of our Red Hat certified curriculums.
AS300 — AI Platform Engineer
View curriculum →
Red Hat OpenShift AI (RHOAI)
AI267 · Module 7 — AS300
GenAI Fundamentals + Granite + RHEL AI
AI296 · Module 8 — AS300