The Real Cost of GPU Inference on Kubernetes

Where the money leaks when you serve models on K8s, and the practical fixes that plug each hole.

No headings found on page

Tech

11

Min Read

Share:

The uncomfortable math

GPUs are usually the largest line item in an AI infrastructure budget, and Kubernetes makes it remarkably easy to waste them.

Consider a simple, illustrative example. Suppose you rent a GPU at $4 per hour and your inference server can sustain 1,000 output tokens per second when it's busy.

Scenario

Tokens per hour

Effective cost per 1M tokens

GPU fully utilized

3.6M

~$1.11

GPU 30% utilized

~1.1M

~$3.70

GPU 10% utilized

~0.36M

~$11.10

(Numbers are illustrative; plug in your own hardware price and measured throughput.)

Same hardware, same model, and a 10x difference in unit cost, purely from utilization. That gap is where the savings live. The rest of this article walks through the six most common leaks and how to fix each.

Rule of thumb: an idle GPU costs exactly the same as a busy one. Your job is to make sure the expensive silicon spends its time doing useful work.

Leak 1: Whole-GPU allocation for small models

By default, Kubernetes treats GPUs as indivisible integers through the nvidia.com/gpu extended resource:

resources:
  limits:
    nvidia.com/gpu: 1
resources:
  limits:
    nvidia.com/gpu: 1
resources:
  limits:
    nvidia.com/gpu: 1

A pod that asks for one GPU gets the entire GPU. If your 7B-parameter model needs 16 GB of memory and modest compute on an 80 GB accelerator, you're paying for 64 GB and most of the compute you'll never touch.

You have three main options for sharing:

Option A: MIG (Multi-Instance GPU)

On supported NVIDIA data-center GPUs (A100, H100, and newer), MIG partitions one physical GPU into up to seven isolated instances, each with dedicated memory, cache, and compute.

  • ✅ Hardware-level isolation and predictable performance

  • ✅ Great for serving several small or medium models side by side

  • ⚠️ Fixed partition profiles, so reconfiguring may require draining the node

  • ⚠️ Only available on supported hardware

Pods then request a specific slice, for example:

resources:
  limits:
    nvidia.com/mig-1g.10gb: 1
resources:
  limits:
    nvidia.com/mig-1g.10gb: 1
resources:
  limits:
    nvidia.com/mig-1g.10gb: 1
Option B: Time-slicing

The NVIDIA device plugin can advertise one physical GPU as several schedulable replicas. Pods take turns on the hardware.

  • ✅ Works on almost any NVIDIA GPU

  • ✅ Simple to configure, ideal for dev and test clusters

  • ⚠️ No memory isolation. One pod can exhaust GPU memory and crash its neighbors

  • ⚠️ No performance guarantees, so latency can vary under contention

Option C: Consolidate at the serving layer

Instead of one pod per model, run a single inference server that hosts multiple models or multiple LoRA adapters on a shared base model. Many modern serving stacks support loading adapters dynamically, which can replace a dozen near-identical deployments with one.

Which to choose?

Need

Best fit

Production, predictable latency, multiple small models

MIG

Dev/test, bursty and low-stakes workloads

Time-slicing

Many fine-tuned variants of one base model

Multi-adapter serving

One large model that fills the GPU anyway

Whole GPU, and that's fine

Also consider GPU-aware scheduling: use node labels, taints, and tolerations so that GPU nodes run only GPU workloads, and CPU-only pods don't accidentally pin an expensive node in place.

Leak 2: Autoscaling on the wrong signal

The Horizontal Pod Autoscaler defaults to CPU utilization. For an inference server, that metric is nearly meaningless: the CPU can sit at 15% while the GPU is saturated and requests pile up in a queue.

Scale on signals that reflect real load:

  • Queue depth (requests waiting)

  • Requests in flight per replica

  • Time-to-first-token (TTFT) latency

  • GPU utilization (useful, but a lagging and coarse indicator)

Popular inference servers such as vLLM expose Prometheus metrics for queue length and running requests. You can feed them to the HPA through KEDA or a Prometheus adapter. Here's an example using KEDA:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: llm-server
spec:
  scaleTargetRef:
    name: llm-server
  minReplicaCount: 1
  maxReplicaCount: 8
  cooldownPeriod: 600          # seconds before scaling down
  triggers:
    - type: prometheus
      metadata:
        serverAddress: http://prometheus.monitoring:9090
        query: sum(vllm:num_requests_waiting)
        threshold: "5"
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: llm-server
spec:
  scaleTargetRef:
    name: llm-server
  minReplicaCount: 1
  maxReplicaCount: 8
  cooldownPeriod: 600          # seconds before scaling down
  triggers:
    - type: prometheus
      metadata:
        serverAddress: http://prometheus.monitoring:9090
        query: sum(vllm:num_requests_waiting)
        threshold: "5"
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: llm-server
spec:
  scaleTargetRef:
    name: llm-server
  minReplicaCount: 1
  maxReplicaCount: 8
  cooldownPeriod: 600          # seconds before scaling down
  triggers:
    - type: prometheus
      metadata:
        serverAddress: http://prometheus.monitoring:9090
        query: sum(vllm:num_requests_waiting)
        threshold: "5"

(Metric names vary by server and version, so check what yours exports.)

Tune the dynamics, not just the trigger
  • Scale up fast, down slowly. Each scale-down that's followed by a scale-up costs you a cold start (see next section).

  • Use a stabilization window to avoid flapping.

  • Set sensible max replicas so a traffic spike or a runaway client can't scale you into a five-figure bill overnight.

  • Add rate limits and queue limits at the gateway, so overload becomes a fast 429 rather than an ever-growing queue.

Leak 3: Cold starts

Scaling up isn't instant. A new replica has to:

  1. Get scheduled (possibly waiting for a new GPU node to be provisioned)

  2. Pull a container image, which is often 10 to 20 GB with CUDA and framework layers

  3. Download model weights, often 10 to 150+ GB depending on the model

  4. Load the weights into GPU memory

  5. Warm up (compile kernels, allocate the KV cache)

That can easily add up to several minutes, during which users are waiting, or queues are growing.

Ways to shorten it:

  • Keep weights out of the image. Store them on a fast shared volume or object storage, and stream or cache them on the node. This keeps images small and lets you update models without rebuilding.

  • Cache on local NVMe. A node-local cache means the second pod on a node loads in seconds rather than minutes.

  • Pre-pull images onto GPU nodes with a DaemonSet so the pull isn't on the critical path.

  • Use faster loading formats (for example, safetensors) and loaders designed for fast weight streaming.

  • Keep a warm pool. Maintain a minimum replica count, or a small buffer of spare capacity, for anything user-facing.

  • Add readiness probes that reflect reality. Don't route traffic until the model is loaded and warmed up, or the first users will hit errors.

Scale-to-zero: great for internal tools, batch endpoints, and low-traffic dev environments. Usually a bad idea for customer-facing endpoints, where the first request after idle would wait minutes.

Leak 4: Ignoring batching and the serving engine

Here is the most under-appreciated fact in inference cost: the same GPU can deliver wildly different throughput depending on the serving software.

Naive serving handles requests one at a time, or waits for a full batch to finish before starting the next. Since generation lengths vary, short requests sit idle behind long ones, and the GPU is underused.

Modern inference engines fix this with several techniques:

Technique

What it does

Continuous batching

Adds new requests into a running batch as others finish, instead of waiting for the whole batch

Paged KV-cache management

Stores attention key/value cache in flexible blocks, cutting memory waste and allowing more concurrent requests

Prefix caching

Reuses computation for shared prompt prefixes such as system prompts and long documents

Chunked prefill

Interleaves processing of long prompts with ongoing generation to keep latency steady

Speculative decoding

Uses a small draft model to propose tokens the large model verifies in bulk, speeding up generation

Moving from a basic serving setup to an engine with continuous batching and paged KV-cache management can multiply throughput on identical hardware. Often the cheapest GPU is the one you don't have to buy.

Practical tips:

  • Benchmark with realistic traffic. Use your real prompt and output length distributions, not a single fixed length. Test at the concurrency you expect.

  • Tune max concurrent sequences and KV-cache memory fraction. Too low wastes GPU, too high causes out-of-memory errors or preemption.

  • Put shared prefixes first. If every request starts with the same 2,000-token system prompt, prefix caching turns that into nearly free work.

  • Separate workloads by shape. Long-context summarization and short chat behave very differently, and mixing them on one deployment can hurt both.

Leak 5: Running models bigger than they need to be

The cheapest token is one generated by a model that fits your problem.

  • Quantization (FP8, INT8, 4-bit methods such as AWQ or GPTQ) shrinks memory footprint and often increases throughput. It can let a model fit on fewer or smaller GPUs. Always re-run your quality evals afterward, because the impact varies by model and task.

  • Right-size the model. Many production tasks (classification, extraction, routing, summarization) work well on a small model that's fine-tuned or well-prompted. Reserve the large model for hard cases.

  • Route intelligently. A router that sends easy requests to a small model and escalates hard ones to a large one can cut costs substantially while preserving quality.

  • Trim tokens. Shorter prompts, concise output instructions, and sensible max_tokens limits directly reduce compute.

  • Cache responses for repeated or near-identical queries where correctness allows.

Each of these changes should go through your evaluation pipeline. Cost savings that quietly degrade quality aren't savings.

Leak 6: Paying on-demand for everything

Not all GPU work has the same requirements, so it shouldn't all be bought the same way.

Workload

Interruption-tolerant?

Good fit

Real-time serving, customer-facing

No

On-demand or committed/reserved capacity

Baseline steady traffic

No

Reserved instances or savings plans

Bursty extra capacity

Sometimes

Mix of on-demand and spot with fast fallback

Batch inference, embeddings

Yes

Spot / preemptible

Offline evals

Yes

Spot / preemptible

Fine-tuning and training

Yes, with checkpoints

Spot / preemptible

For interruptible workloads:

  • Checkpoint frequently so a preemption loses minutes of work, not hours.

  • Handle termination notices gracefully and drain in-flight work.

  • Diversify across instance types and zones to improve spot availability.

  • Use queue-based job runners so interrupted work is simply retried.

For steady baseline load, commitments (reserved instances, savings plans) typically offer meaningful discounts compared to on-demand pricing. Size them to your floor of usage, not your peak.

Measure first: the metrics that matter

You can't optimize what you can't see. Build a dashboard around these:

Efficiency

  • GPU utilization and GPU memory utilization, per pod and per node

  • Fraction of GPU-hours spent idle

  • Batch size and KV-cache utilization

Performance

  • Time-to-first-token (p50, p95, p99)

  • Inter-token latency

  • Throughput (tokens per second per GPU)

  • Queue depth and rejected requests

Cost

  • Cost per million tokens (input and output separately)

  • Cost per request, and cost per successful task

  • Cost per team, per model, and per environment (use labels and cost-allocation tooling)

That cost-per-token metric is the most important one on the list. It turns infrastructure decisions (MIG profile, quantization, spot vs. on-demand) into numbers that non-infrastructure people can reason about, and makes tradeoffs concrete: "This change raises p99 latency by 80 ms but cuts unit cost by 35%. Is that worth it?"

Tooling to consider: DCGM Exporter for GPU metrics, Prometheus and Grafana for dashboards, your serving engine's built-in metrics, and a Kubernetes cost tool such as OpenCost or Kubecost for allocation.

Putting it together

A realistic optimization roadmap, ordered by effort versus payoff:

  1. Instrument. Get utilization and cost-per-token visible. Without this, everything else is guesswork.

  2. Fix the serving engine. Adopt continuous batching, paged KV-cache, and prefix caching if you haven't already. This is often the biggest single win.

  3. Fix autoscaling. Scale on queue depth or TTFT, not CPU, and tune scale-down behavior.

  4. Share GPUs where models are small: MIG for production, time-slicing for dev.

  5. Shrink cold starts with local caching, pre-pulled images, and a warm pool.

  6. Right-size and quantize. Validate with evals.

  7. Optimize purchasing. Commit for your baseline, use spot for everything interruptible.

Do them roughly in that order. Steps 1 through 3 are usually cheap to implement and capture a large share of the savings.

A cost-reduction checklist
  • GPU utilization and cost per million tokens are on a dashboard

  • GPU nodes are tainted so only GPU workloads run on them

  • Inference uses an engine with continuous batching and paged KV-cache

  • Prefix caching is enabled where prompts share long prefixes

  • Autoscaling uses queue depth or latency, not CPU

  • Scale-down has a cooldown; max replicas are capped

  • Model weights are cached locally and kept out of the container image

  • Customer-facing endpoints keep a warm minimum

  • Small models share GPUs via MIG or multi-model/adapter serving

  • Quantization and smaller models have been evaluated for suitable workloads

  • Batch, eval, and training jobs run on spot with checkpointing

  • Baseline capacity is covered by commitments

Closing thought

GPU cost isn't a mystery, and it isn't only a procurement problem. It's an engineering problem with well-understood levers: utilization, batching, right-sizing, and purchasing strategy. Teams that treat cost per token as a first-class metric, right next to latency and reliability, routinely cut their inference bills dramatically without users noticing anything except faster responses.

2026 Elias, All rights reserved

2026 Elias, All rights reserved

Create a free website with Framer, the website builder loved by startups, designers and agencies.