Most inference services use only part of a GPU, and a whole card is often more than a notebook, a batch-scoring job or a small model needs. Sharing one card between workloads sounds simple, but Kubernetes does not do it by default, and the different ways of doing it make very different promises. This article explains those promises, what a shared slice on Kube-DC guarantees and does not, and what we learned by running a real LLM on one.

Kubernetes hands out whole GPUs by default

The standard Kubernetes device-plugin model reports GPUs to the scheduler as whole numbers. A request for nvidia.com/gpu: 1 means the entire card, and every other pod waits. The scheduler has no notion of a memory quota, so sharing a card needs extra machinery, and there are three common approaches.

Approach Memory limit Isolation Works on a V100?
MIG (hardware partitioning) Enforced in hardware Memory and fault isolation at the hardware layer No: it needs Ampere or later (A100, H100, A30, H200), and the V100 is an older, Volta-generation card
Time-slicing None: replicas share the card’s memory No memory or fault isolation between replicas, according to NVIDIA’s documentation Yes
Software slicing (HAMi) Enforced in software: allocations beyond the slice are refused Software isolation by CUDA API interception; not hardware isolation, and compute is throttled on a best-effort basis Yes

HAMi, a Cloud Native Computing Foundation project, works by loading a library into every container that intercepts CUDA calls. The container sees only its slice of memory, an allocation that would exceed the slice returns an out-of-memory error, and kernel launches are throttled toward the requested compute share. HAMi’s own documentation is careful about the limits: applications that bypass the CUDA library, such as Docker-in-Docker or direct driver calls, are not covered, and the isolation is best-effort compared with MIG.

What a shared slice on Kube-DC guarantees

Kube-DC’s GPU packages use software slicing on NVIDIA V100 cards, and the platform documentation states the guarantees precisely. A slice is a fixed fraction of one GPU model; you request the product rather than tuning memory and compute freely. On a 32 GB V100, our packages offer 25% (8 GiB), 50% (16 GiB) and 100% (32 GiB).

  • The memory slice is enforced. Inside the container, nvidia-smi reports the size of the slice, not of the physical card, and the slice is a hard cap.
  • The compute share is cooperative. The platform steers each workload toward its share in steady state, but startup and some CUDA library paths can briefly exceed it. It is not guaranteed performance.
  • Quota is an entitlement, not a reservation. Holding quota does not reserve a physical slice. When every compatible GPU is busy, a valid workload waits in the Pending state until a slice frees up.
  • Slices are for containers. Shared capacity cannot be attached to a virtual machine.
A shared slice is a scheduling-and-accounting guarantee with runtime enforcement. It is not a security boundary. A workload that must not share silicon with another tenant needs a dedicated GPU.

What using one looks like

Since Kubernetes 1.34 the Dynamic Resource Allocation (DRA) APIs in resource.k8s.io/v1 are generally available, and Kube-DC uses them: a workload references a ResourceClaimTemplate that points at the DeviceClass for your product, and the platform validates the request when you apply it. Your project quota shows how many slices of each product you may hold at once:

kubectl get resourcequota -n <project-namespace> -o yaml | grep deviceclass

Once your pod is running, confirm the slice from inside the container. On the 8 GiB product the memory total is 8192 MiB even though the card has 32 GB:

nvidia-smi --query-gpu=name,memory.total --format=csv,noheader
# Tesla V100-PCIE-32GB, 8192 MiB

The complete, copy-and-edit manifest is in the Kube-DC documentation, and we deliberately do not reproduce it here, because it has to match the platform’s published contract exactly. In our own test, two attempts failed for that reason before the third succeeded:

  • A plain nvidia.com/gpu request on a raw pod was rejected by an admission policy, because shared GPU capacity is only offered through the claim-based pattern.
  • A hand-written DRA claim template was rejected by a second admission policy, because the labels, names and fixed capacity values must follow the documented shape.
  • The documented manifest was accepted, then the pod failed to start for an unrelated reason: the project’s pooled CPU quota was almost fully used by a Managed Cluster and a virtual machine. A GPU pod still needs CPU and memory from your package, so check that headroom before you deploy.

Running an LLM on a slice: what we learned

In our August 2026 end-to-end test we ran Ollama on an 8 GiB slice and served llama3.2:3b, then connected an AI agent to it, first from a virtual machine in the same project and then from a server outside the datacenter. The Ollama logs showed the card as a Tesla V100 with compute capability 7.0 and a total of 8.0 GiB, exactly the slice. Four details cost us time, and each one applies to anyone running a model on a slice.

1. Check that your software still supports the GPU generation

The V100 is a Volta-generation card, and current GPU software has started to move on. In our test the latest Ollama image skipped the GPU, because its CUDA build did not include compute capability 7.0. Pinning ollama/ollama:0.24.0 fixed it. This matches the upstream reports: an Ollama issue describes v0.30.0 failing on a Tesla V100 with the error “device kernel image is invalid” while v0.24.0 and earlier worked, and a maintainer later attributed one such failure to compressed CUDA kernels that need NVIDIA driver 550 or newer. Pin your image version and test upgrades deliberately.

2. The pod’s memory limit is separate from the GPU slice

Our first model load was killed with exit code 137 (OOMKilled). The cause was the container’s own memory limit of 512 MiB, which is far too small for Ollama’s host-side overhead, and had nothing to do with GPU memory. Size the container’s RAM request and limit independently of the fixed slice. We used 2 GiB and 4 GiB:

containers:
- name: ollama
  image: ollama/ollama:0.24.0
  env:
  - name: OLLAMA_CONTEXT_LENGTH
    value: "16384"
  resources:
    requests:
      memory: 2Gi
    limits:
      memory: 4Gi

3. The default context window can silently truncate an agent’s prompt

The agent we connected sends a system prompt and tool definitions of roughly 14,700 tokens. The Ollama server log showed the input being cut to its 4,096-token limit, which produced garbled, unreliable answers that looked like a model-quality problem. Raising the context length with the variable above fixed the cause. Remember that a longer context uses more of your slice’s memory, so size the model and the context together.

4. Models live in the pod unless you give them a volume

With the default temporary storage, every pod restart, including one caused by changing an environment variable, wiped the pulled models. Put the model directory on a persistent volume for anything beyond a demo. Also note that tenant roles on Kube-DC cannot use kubectl exec, so we pulled models with one-shot Jobs that call the Ollama HTTP API.

When a shared slice is the right tool

The Kube-DC documentation names the fit: model inference, notebooks, small training runs and batch scoring, where a job needs GPU acceleration but not a whole card. It is a poor fit when you need a hard isolation boundary, guaranteed throughput, or a model larger than the slice, and whole-card work belongs on dedicated hardware. Because slices queue when the card is busy, it also suits workloads that can wait for capacity more than latency-critical services with strict deadlines.

How to evaluate any shared-GPU offer

  • Is the memory limit enforced, or is it a suggestion? Time-slicing enforces none.
  • Is the compute share guaranteed, cooperative or unlimited, and what happens under contention?
  • Is quota a reservation, or can your workload wait in Pending when the card is busy?
  • Which GPU model is it, and does the software you plan to run still support that generation?
  • Can you get a whole card or a dedicated VM if you outgrow a slice, and what does isolation look like there?
  • Which CPU, memory and storage do you get alongside the GPU, and are they pooled with the rest of your workloads?

Sources

Related reading

Try it on Kube-DC

Our GPU packages add a slice of a shared NVIDIA V100 to a Kube-DC pool: 25%, 50% or 100% of the card, with the memory slice enforced. Alongside it you get dedicated vCPU, Ceph-backed NVMe storage, a dedicated public IPv4, S3-compatible object storage and your own tenant-administered Kubernetes clusters, with 24/7 email and ticket support and a 4-hour resolution target. We ran a real LLM and an AI agent on it end to end, and every step in this article is one you can repeat.

Explore Kube-DC Kubernetes