GPU Cost Allocation on Kubernetes: A 2026 Chargeback Playbook

Kubernetes bills GPUs by who held the card, never by who used it. Here is the instrumentation, label schema, and chargeback model that make GPU spend defensible.

By VVV Ops ·

Your CPU chargeback works. Your GPU chargeback is a rounding error wearing a suit. The gap is structural, not a tooling problem: GPU cost allocation on Kubernetes inherits a scheduling model built for whole devices, so the cluster faithfully records who held a card and has no idea who used one. A team that parked a notebook on an H100 at 4% occupancy for three weeks gets the same invoice line as the team that saturated it. We have watched that single gap turn a $30K monthly node group into an argument that nobody could settle with data. This is how to close it.

Why nvidia.com/gpu breaks the model you already have

Start with the constraint, because every bad GPU chargeback number traces back to it. The Kubernetes docs on scheduling GPUs are explicit: GPUs "are only supposed to be specified in the limits section", and if you set both, "these two values must be equal." There is no fractional request. nvidia.com/gpu: 1 means one entire physical device, bound to one container, for the life of the pod.

Now put that through a cost engine. The OpenCost specification defines workload cost as max(request, usage). For CPU that is a genuine signal, because requests are fractional and usage is measured in the same unit, so the two numbers can disagree and the disagreement is informative. For GPU the request is always the whole device. max(request, usage) collapses to "the whole device, every hour it was bound," and occupancy never enters the arithmetic.

The money involved makes this worth fixing rather than tolerating. AWS lists a p5.4xlarge in US East (N. Virginia) at an effective hourly rate of $5.191 per accelerator for one H100 under Capacity Blocks. The eight-GPU p5.48xlarge on the same page is $41.528 an hour. Over a 730-hour month that is $30,315.44, and your allocation report will assign all of it, to the decimal, without telling you whether a single tensor core did work.

This is the part our Kubernetes cost optimization framework does not reach. Right-sizing requests and killing idle workloads still applies to everything else on the node. It cannot touch a resource that has exactly one legal request value.

The sharing modes and what each one does to your cost model

You have four ways to hand a GPU to a workload, and they are not interchangeable from a billing standpoint.

| Mode | What the scheduler counts | Isolation | Honest unit of chargeback | |---|---|---|---| | Whole device | nvidia.com/gpu: 1 | Full | GPU-hours held | | Time-slicing | nvidia.com/gpu: 1 against an inflated replica count | None | Measured occupancy only | | MIG | nvidia.com/mig-1g.10gb: 1 | Hardware-partitioned | Partition-hours, at a fixed fraction of the card | | DRA | A ResourceClaim naming a DeviceClass | Depends on the driver | Claim duration per device |

Time-slicing is where teams lose the most credibility. NVIDIA's GPU Operator exposes it through a ConfigMap under sharing.timeSlicing.resources, where each entry carries a name and a replicas count, so one physical card can advertise four. The documentation is blunt about what you are buying: "Unlike Multi-Instance GPU (MIG), there is no memory or fault-isolation between replicas." It also warns that "a request for more than one time-sliced GPU does not guarantee that the pod receives access to a proportional amount of GPU compute power."

Read that warning as a billing rule. If four replicas do not deliver four equal quarters of compute, then dividing the card's price by four and sending each tenant a quarter of the bill is a number you cannot defend when a tenant disputes it. Do not bill time-sliced replicas as equal shares. Bill them by measured occupancy, which means you need the instrumentation in the next section before you enable time-slicing at all.

MIG is the opposite case. The GPU Operator applies a profile through the nvidia.com/mig.config node label under either the single or mixed strategy, and Kubernetes then sees named partition resources such as nvidia.com/mig-1g.10gb. Those partitions are enforced in hardware, so a static price per partition-hour is honest. Our recommendation for any card that supports MIG and any workload that fits inside a profile: use MIG. It is the only sharing mode where the simple allocation math is also the correct allocation math.

Dynamic Resource Allocation changes the shape of the problem. It has been stable since Kubernetes v1.35, and the resource.k8s.io API group models devices as ResourceSlice objects that drivers publish and ResourceClaim objects that pods bind. A claim is a durable, queryable record of which specific device a workload got and for how long. That is an attribution primitive the integer extended-resource count never had. Cost tooling has not caught up yet, so treat DRA as the direction to build toward rather than a reporting feature you can turn on this quarter.

Instrument before you allocate

Occupancy data comes from DCGM Exporter, and the metric you pick decides whether your numbers survive review. The default counter set enables DCGM_FI_DEV_GPU_UTIL, DCGM_FI_DEV_FB_USED, DCGM_FI_DEV_POWER_USAGE, DCGM_FI_PROF_GR_ENGINE_ACTIVE and DCGM_FI_PROF_PIPE_TENSOR_ACTIVE, and attaches namespace, pod, container and gpu labels to each sample.

Use DCGM_FI_PROF_GR_ENGINE_ACTIVE as your occupancy signal, not DCGM_FI_DEV_GPU_UTIL. The device utilization counter reports whether any kernel was resident during the sampling window, so a workload that launches one tiny kernel per millisecond reads close to 100% while the card is mostly idle. The profiling counter reports the fraction of time the graphics engine was actually active. On a card you are billing at five dollars an hour, that difference is the whole argument.

Two queries produce the numbers a chargeback report needs. Both assume a 30-second Prometheus scrape interval, so change the constant if yours differs.

# GPU-hours held per namespace over 30 days.
# Each sample represents one GPU held for one scrape interval (30s).
sum by (namespace) (
  count_over_time(DCGM_FI_DEV_GPU_UTIL{namespace!=""}[30d])
) * 30 / 3600

# Mean engine occupancy per namespace over the same window (0 to 1).
avg by (namespace) (
  avg_over_time(DCGM_FI_PROF_GR_ENGINE_ACTIVE{namespace!=""}[30d])
)

The first query is what you bill. The second is what you show alongside it. A namespace holding 5,840 GPU-hours at 0.06 mean occupancy is not a cost problem you solve with a discount, and putting both figures on the same row is what moves the conversation from procurement to engineering.

A label schema that survives an audit

Cost data is only as good as the dimensions you can group it by, and GPU workloads need more dimensions than web services do. Set them at the pod template, not with a tagging policy that runs after the fact.

# Illustrative. Replace acme with your own prefix.
metadata:
  labels:
    acme.dev/cost-center: "ml-platform"      # who pays
    acme.dev/workload-kind: "inference"      # inference | training | interactive
    acme.dev/model: "acme-embed-v3"          # which model, for per-model unit cost
    acme.dev/tenant: "internal"              # internal | customer-dedicated

Four labels, each answering a question a finance partner will ask. Separating inference from training matters more than it looks: training is a project cost with an end date, inference is unit cost that scales with revenue, and blending them into one GPU line makes both unreadable.

Map those labels onto FOCUS 1.4 columns when the data leaves the cluster. Resource ID carries the node or device identity, Tags carries your four labels, and Consumed Quantity against Pricing Unit is where GPU-hours belong. Landing on FOCUS means your Kubernetes numbers and your SaaS model-API invoices sit in one table, which is the point at which anyone can answer what a feature costs end to end. If you have not built the practice around this yet, our FinOps practice guide covers the operating model that consumes this data.

Choose an allocation model and defend it

| Model | Bill on | Use when | Fails when | |---|---|---|---| | Per-namespace GPU-hours | Device-hours bound | Teams own dedicated node pools | One namespace serves many products | | Per-MIG-partition | Partition-hours at a fixed fraction | Mixed small workloads on A100/H100 class cards | Workloads need more memory than the largest profile | | Occupancy-weighted | GPU-hours scaled by engine-active ratio | Time-slicing or shared research clusters | Idle cost has no owner, so someone must absorb it | | Per-token | Tokens served against total GPU-hours | Self-hosted inference behind one gateway | Batch and interactive traffic share the endpoint |

We pick per-namespace GPU-hours for training and per-token for inference, with MIG underneath wherever the hardware allows. Occupancy-weighted models are the tempting middle option and the one we advise against as a primary model, because scaling charges down by utilization leaves the idle remainder unallocated, and unallocated GPU cost quietly becomes nobody's problem. Report occupancy. Bill on hours held. Put the idle gap on the platform team's own line where it will get worked.

Per-token is the model worth building toward for anything customer-facing. Divide the full GPU-hour cost of a serving deployment, idle included, by tokens served in the same window. That gives a cost per million tokens you can hold against a vendor API quote and make an actual build-versus-buy call, rather than the vibes-based version most teams are running on.

What we would do in your first two weeks

  1. Deploy DCGM Exporter and confirm namespace and pod labels are populated on DCGM_FI_PROF_GR_ENGINE_ACTIVE. If they are empty, nothing downstream is worth building.
  2. Run both queries above over the last 30 days and put GPU-hours next to mean occupancy for every namespace. Expect the result to be uncomfortable.
  3. Apply the four labels to every GPU pod template. Block admission on the cost-center label once coverage passes 90%.
  4. Turn off time-slicing anywhere it is running in production. Move those workloads to MIG profiles or to their own smaller cards. An L4-backed g6.xlarge with 24 GB of GPU memory is the right answer for far more inference workloads than teams assume.
  5. Publish one report, to the teams and to finance at the same time. The first version being wrong is fine. The first version being private is not.

The State of FinOps 2026 data says 98% of FinOps practitioners now manage AI spend, up from 31% two years ago, and puts pre-deployment architecture costing among the capabilities practitioners want most. You cannot cost an architecture before you deploy it until you know what the last one actually cost. This work is the prerequisite.

For the wider reliability picture around these clusters, our Kubernetes production readiness checklist covers the operational side that GPU node pools tend to skip on the way to production.

When to Get Help

Bring someone in when the GPU line crosses roughly $50K a month, when two teams are disputing a shared cluster bill, or when you are deciding between self-hosted inference and a vendor API and cannot produce a defensible cost per million tokens for the self-hosted side. Those three situations have the same root cause and the same fix, and they get more expensive the longer the instrumentation gap stays open.

We do this work with engineering and finance in the same room, because a GPU allocation model that engineering cannot implement or finance will not accept is worse than no model. If that sounds like your quarter, get in touch.

Tags: gpu cost allocation on kubernetes, cloud cost optimization, finops implementation, infrastructure cost optimization, kubernetes gpu chargeback, cloud spend management