KServe
Standardized, serverless model serving on Kubernetes with autoscaling.
Official website / GitHub ↗ Visit
What is this technology for?
Its role in production environments and why it is taught in our curriculums.
KServe standardizes model serving on Kubernetes: it deploys, scales and manages inference endpoints for ML and LLM models with serverless capabilities.
What you will learn
The hands-on skills you will gain on this technology in our curriculums.
- Deploy models as KServe InferenceServices
- Configure autoscaling and canary rollouts
- Serve models with vLLM, OpenVINO and Triton runtimes
- Secure and monitor inference endpoints
Latest News & Ecosystem Updates
Recent innovations, major releases, and key industry milestones in the ecosystem.
KServe v0.15: v2 Data Plane GA & Intelligent Multi-Model Routing
Declarative serverless model serving with scale-to-zero, GPU fractional sharing, and real-time drift detection.
Automated Inference Pipeline Delivery with Knative & Istio
End-to-end mTLS encryption and fine-grained canary traffic shaping for production AI microservices.
Featured in the curriculum
Find this technology in the following modules of our Red Hat certified curriculums.
AS300 — AI Platform Engineer
View curriculum →
Red Hat OpenShift AI (RHOAI)
AI267 · Module 7 — AS300