The Real Cost of GPU Inference on Kubernetes
Where the money leaks when you serve models on K8s, and the practical fixes that plug each hole.
The uncomfortable math
GPUs are usually the largest line item in an AI infrastructure budget, and Kubernetes makes it remarkably easy to waste them.
Consider a simple, illustrative example. Suppose you rent a GPU at $4 per hour and your inference server can sustain 1,000 output tokens per second when it's busy.
Scenario | Tokens per hour | Effective cost per 1M tokens |
|---|---|---|
GPU fully utilized | 3.6M | ~$1.11 |
GPU 30% utilized | ~1.1M | ~$3.70 |
GPU 10% utilized | ~0.36M | ~$11.10 |
(Numbers are illustrative; plug in your own hardware price and measured throughput.)
Same hardware, same model, and a 10x difference in unit cost, purely from utilization. That gap is where the savings live. The rest of this article walks through the six most common leaks and how to fix each.
Rule of thumb: an idle GPU costs exactly the same as a busy one. Your job is to make sure the expensive silicon spends its time doing useful work.
Leak 1: Whole-GPU allocation for small models
By default, Kubernetes treats GPUs as indivisible integers through the nvidia.com/gpu extended resource:
A pod that asks for one GPU gets the entire GPU. If your 7B-parameter model needs 16 GB of memory and modest compute on an 80 GB accelerator, you're paying for 64 GB and most of the compute you'll never touch.
You have three main options for sharing:
Option A: MIG (Multi-Instance GPU)
On supported NVIDIA data-center GPUs (A100, H100, and newer), MIG partitions one physical GPU into up to seven isolated instances, each with dedicated memory, cache, and compute.
✅ Hardware-level isolation and predictable performance
✅ Great for serving several small or medium models side by side
⚠️ Fixed partition profiles, so reconfiguring may require draining the node
⚠️ Only available on supported hardware
Pods then request a specific slice, for example:
Option B: Time-slicing
The NVIDIA device plugin can advertise one physical GPU as several schedulable replicas. Pods take turns on the hardware.
✅ Works on almost any NVIDIA GPU
✅ Simple to configure, ideal for dev and test clusters
⚠️ No memory isolation. One pod can exhaust GPU memory and crash its neighbors
⚠️ No performance guarantees, so latency can vary under contention
Option C: Consolidate at the serving layer
Instead of one pod per model, run a single inference server that hosts multiple models or multiple LoRA adapters on a shared base model. Many modern serving stacks support loading adapters dynamically, which can replace a dozen near-identical deployments with one.
Which to choose?
Need | Best fit |
|---|---|
Production, predictable latency, multiple small models | MIG |
Dev/test, bursty and low-stakes workloads | Time-slicing |
Many fine-tuned variants of one base model | Multi-adapter serving |
One large model that fills the GPU anyway | Whole GPU, and that's fine |
Also consider GPU-aware scheduling: use node labels, taints, and tolerations so that GPU nodes run only GPU workloads, and CPU-only pods don't accidentally pin an expensive node in place.
Leak 2: Autoscaling on the wrong signal
The Horizontal Pod Autoscaler defaults to CPU utilization. For an inference server, that metric is nearly meaningless: the CPU can sit at 15% while the GPU is saturated and requests pile up in a queue.
Scale on signals that reflect real load:
Queue depth (requests waiting)
Requests in flight per replica
Time-to-first-token (TTFT) latency
GPU utilization (useful, but a lagging and coarse indicator)
Popular inference servers such as vLLM expose Prometheus metrics for queue length and running requests. You can feed them to the HPA through KEDA or a Prometheus adapter. Here's an example using KEDA:
(Metric names vary by server and version, so check what yours exports.)
Tune the dynamics, not just the trigger
Scale up fast, down slowly. Each scale-down that's followed by a scale-up costs you a cold start (see next section).
Use a stabilization window to avoid flapping.
Set sensible max replicas so a traffic spike or a runaway client can't scale you into a five-figure bill overnight.
Add rate limits and queue limits at the gateway, so overload becomes a fast
429rather than an ever-growing queue.
Leak 3: Cold starts
Scaling up isn't instant. A new replica has to:
Get scheduled (possibly waiting for a new GPU node to be provisioned)
Pull a container image, which is often 10 to 20 GB with CUDA and framework layers
Download model weights, often 10 to 150+ GB depending on the model
Load the weights into GPU memory
Warm up (compile kernels, allocate the KV cache)
That can easily add up to several minutes, during which users are waiting, or queues are growing.
Ways to shorten it:
Keep weights out of the image. Store them on a fast shared volume or object storage, and stream or cache them on the node. This keeps images small and lets you update models without rebuilding.
Cache on local NVMe. A node-local cache means the second pod on a node loads in seconds rather than minutes.
Pre-pull images onto GPU nodes with a DaemonSet so the pull isn't on the critical path.
Use faster loading formats (for example, safetensors) and loaders designed for fast weight streaming.
Keep a warm pool. Maintain a minimum replica count, or a small buffer of spare capacity, for anything user-facing.
Add readiness probes that reflect reality. Don't route traffic until the model is loaded and warmed up, or the first users will hit errors.
Scale-to-zero: great for internal tools, batch endpoints, and low-traffic dev environments. Usually a bad idea for customer-facing endpoints, where the first request after idle would wait minutes.
Leak 4: Ignoring batching and the serving engine
Here is the most under-appreciated fact in inference cost: the same GPU can deliver wildly different throughput depending on the serving software.
Naive serving handles requests one at a time, or waits for a full batch to finish before starting the next. Since generation lengths vary, short requests sit idle behind long ones, and the GPU is underused.
Modern inference engines fix this with several techniques:
Technique | What it does |
|---|---|
Continuous batching | Adds new requests into a running batch as others finish, instead of waiting for the whole batch |
Paged KV-cache management | Stores attention key/value cache in flexible blocks, cutting memory waste and allowing more concurrent requests |
Prefix caching | Reuses computation for shared prompt prefixes such as system prompts and long documents |
Chunked prefill | Interleaves processing of long prompts with ongoing generation to keep latency steady |
Speculative decoding | Uses a small draft model to propose tokens the large model verifies in bulk, speeding up generation |
Moving from a basic serving setup to an engine with continuous batching and paged KV-cache management can multiply throughput on identical hardware. Often the cheapest GPU is the one you don't have to buy.
Practical tips:
Benchmark with realistic traffic. Use your real prompt and output length distributions, not a single fixed length. Test at the concurrency you expect.
Tune
max concurrent sequencesand KV-cache memory fraction. Too low wastes GPU, too high causes out-of-memory errors or preemption.Put shared prefixes first. If every request starts with the same 2,000-token system prompt, prefix caching turns that into nearly free work.
Separate workloads by shape. Long-context summarization and short chat behave very differently, and mixing them on one deployment can hurt both.
Leak 5: Running models bigger than they need to be
The cheapest token is one generated by a model that fits your problem.
Quantization (FP8, INT8, 4-bit methods such as AWQ or GPTQ) shrinks memory footprint and often increases throughput. It can let a model fit on fewer or smaller GPUs. Always re-run your quality evals afterward, because the impact varies by model and task.
Right-size the model. Many production tasks (classification, extraction, routing, summarization) work well on a small model that's fine-tuned or well-prompted. Reserve the large model for hard cases.
Route intelligently. A router that sends easy requests to a small model and escalates hard ones to a large one can cut costs substantially while preserving quality.
Trim tokens. Shorter prompts, concise output instructions, and sensible
max_tokenslimits directly reduce compute.Cache responses for repeated or near-identical queries where correctness allows.
Each of these changes should go through your evaluation pipeline. Cost savings that quietly degrade quality aren't savings.
Leak 6: Paying on-demand for everything
Not all GPU work has the same requirements, so it shouldn't all be bought the same way.
Workload | Interruption-tolerant? | Good fit |
|---|---|---|
Real-time serving, customer-facing | No | On-demand or committed/reserved capacity |
Baseline steady traffic | No | Reserved instances or savings plans |
Bursty extra capacity | Sometimes | Mix of on-demand and spot with fast fallback |
Batch inference, embeddings | Yes | Spot / preemptible |
Offline evals | Yes | Spot / preemptible |
Fine-tuning and training | Yes, with checkpoints | Spot / preemptible |
For interruptible workloads:
Checkpoint frequently so a preemption loses minutes of work, not hours.
Handle termination notices gracefully and drain in-flight work.
Diversify across instance types and zones to improve spot availability.
Use queue-based job runners so interrupted work is simply retried.
For steady baseline load, commitments (reserved instances, savings plans) typically offer meaningful discounts compared to on-demand pricing. Size them to your floor of usage, not your peak.
Measure first: the metrics that matter
You can't optimize what you can't see. Build a dashboard around these:
Efficiency
GPU utilization and GPU memory utilization, per pod and per node
Fraction of GPU-hours spent idle
Batch size and KV-cache utilization
Performance
Time-to-first-token (p50, p95, p99)
Inter-token latency
Throughput (tokens per second per GPU)
Queue depth and rejected requests
Cost
Cost per million tokens (input and output separately)
Cost per request, and cost per successful task
Cost per team, per model, and per environment (use labels and cost-allocation tooling)
That cost-per-token metric is the most important one on the list. It turns infrastructure decisions (MIG profile, quantization, spot vs. on-demand) into numbers that non-infrastructure people can reason about, and makes tradeoffs concrete: "This change raises p99 latency by 80 ms but cuts unit cost by 35%. Is that worth it?"
Tooling to consider: DCGM Exporter for GPU metrics, Prometheus and Grafana for dashboards, your serving engine's built-in metrics, and a Kubernetes cost tool such as OpenCost or Kubecost for allocation.
Putting it together
A realistic optimization roadmap, ordered by effort versus payoff:
Instrument. Get utilization and cost-per-token visible. Without this, everything else is guesswork.
Fix the serving engine. Adopt continuous batching, paged KV-cache, and prefix caching if you haven't already. This is often the biggest single win.
Fix autoscaling. Scale on queue depth or TTFT, not CPU, and tune scale-down behavior.
Share GPUs where models are small: MIG for production, time-slicing for dev.
Shrink cold starts with local caching, pre-pulled images, and a warm pool.
Right-size and quantize. Validate with evals.
Optimize purchasing. Commit for your baseline, use spot for everything interruptible.
Do them roughly in that order. Steps 1 through 3 are usually cheap to implement and capture a large share of the savings.
A cost-reduction checklist
GPU utilization and cost per million tokens are on a dashboard
GPU nodes are tainted so only GPU workloads run on them
Inference uses an engine with continuous batching and paged KV-cache
Prefix caching is enabled where prompts share long prefixes
Autoscaling uses queue depth or latency, not CPU
Scale-down has a cooldown; max replicas are capped
Model weights are cached locally and kept out of the container image
Customer-facing endpoints keep a warm minimum
Small models share GPUs via MIG or multi-model/adapter serving
Quantization and smaller models have been evaluated for suitable workloads
Batch, eval, and training jobs run on spot with checkpointing
Baseline capacity is covered by commitments
Closing thought
GPU cost isn't a mystery, and it isn't only a procurement problem. It's an engineering problem with well-understood levers: utilization, batching, right-sizing, and purchasing strategy. Teams that treat cost per token as a first-class metric, right next to latency and reliability, routinely cut their inference bills dramatically without users noticing anything except faster responses.