Persistent Model Cache
August 22, 2026 · View on GitHub
LLMKube includes a persistent model cache that avoids re-downloading models when InferenceServices are deleted and recreated.
Overview
Without persistent caching, models are downloaded via init container every time a pod starts. For large models (13B-70B), this means 26-40GB+ downloads taking 10-30+ minutes each time you recreate a deployment.
With persistent caching:
- A PVC is created automatically in each namespace where you deploy models
- Models are downloaded once to the namespace's PVC
- Subsequent pods mount the cache and skip download
- Delete/recreate cycles complete in seconds
Architecture
LLMKube uses per-namespace PVCs for model caching. This provides:
- Namespace isolation: Each namespace has its own cache
- No cross-namespace dependencies: Models work independently
- Simple RBAC: No need for cross-namespace access
┌─────────────────────────────────────────────────────────────┐
│ Namespace: production │
│ ┌─────────────────────────────────────────────────────┐ │
│ │ llmkube-model-cache PVC │ │
│ │ /models/<cache-key>/model.gguf │ │
│ └─────────────────────────────────────────────────────┘ │
│ ▲ ▲ │
│ │ (init container writes) │ (read-only) │
│ │ │ │
│ ┌────────┴────────┐ ┌────────────┴──────────────┐ │
│ │ First Pod │ │ Subsequent Pods │ │
│ │ - Downloads │ │ - Mount cache read-only │ │
│ │ - Caches model │ │ - Skip download │ │
│ └─────────────────┘ └───────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ Namespace: staging │
│ ┌─────────────────────────────────────────────────────┐ │
│ │ llmkube-model-cache PVC │ │
│ │ /models/<cache-key>/model.gguf │ │
│ └─────────────────────────────────────────────────────┘ │
│ ▲ │
│ │ │
│ ┌────────┴────────┐ │
│ │ Pods in staging │ │
│ └─────────────────┘ │
└─────────────────────────────────────────────────────────────┘
Deploying to Any Namespace
You can deploy models to any namespace using the CLI:
# Deploy to production namespace
llmkube deploy llama-3.1-8b --gpu -n production
# Deploy to staging namespace
llmkube deploy llama-3.1-8b --gpu -n staging
# Deploy to default namespace
llmkube deploy llama-3.1-8b --gpu
The controller will automatically:
- Create a
llmkube-model-cachePVC in the target namespace (if it doesn't exist) - Configure the pod's init-container to download the model to the PVC
- Mount the PVC read-only for the main container
Note: Each namespace has its own PVC, so the same model deployed to multiple namespaces will be downloaded once per namespace.
Cache Key
Models are cached using a SHA256 hash of the source URL (first 16 characters). This means:
- Models with the same source URL share cache entries
- Changing the source URL creates a new cache entry
- The cache key is stored in
Model.Status.CacheKey
Example:
Source: https://huggingface.co/TheBloke/Llama-2-7B-GGUF/resolve/main/llama-2-7b.Q4_K_M.gguf
Cache Key: a3b8c9d4e5f67890
Path: /models/a3b8c9d4e5f67890/model.gguf
Prefetch (Eager Download)
By default a Model with a remote source is only a declaration: nothing is
downloaded until the first InferenceService referencing it starts. Set
spec.prefetch: true to have the operator pull the artifact into the shared
cache immediately:
apiVersion: inference.llmkube.dev/v1alpha1
kind: Model
metadata:
name: llama-3b-prefetch
spec:
source: hf://bartowski/Llama-3.2-3B-Instruct-GGUF/Llama-3.2-3B-Instruct-Q4_K_M.gguf
format: gguf
prefetch: true
The controller runs an owner-referenced download Job that reuses the same
init-container downloader the serving path uses, writing into the
namespace's shared cache PVC (llmkube-model-cache, created on demand).
status.phase moves Downloading -> Ready; once Ready, the first
InferenceService starts from a cache hit instead of a cold download.
Limitations:
- Prefetch applies to remote sources (
https://,hf://). Local paths andpvc://sources ignore the field. - Prefetch targets the shared cache only.
perServicecache PVCs bind at serve time and cannot be pre-populated; pre-stage those fleets with apvc://source instead. - A failed prefetch Job sets
status.phase: Failed; check the Job's pod logs (<model-name>-prefetch). The Job self-cleans 24 hours after finishing and is garbage-collected with the Model.
See config/samples/model_prefetch.yaml for a complete example.
Configuration
Helm Values
modelCache:
# Enable persistent model cache (default: true)
enabled: true
# Provisioning mode: shared or perService.
# shared (default): a single cluster-wide PVC that all InferenceServices share.
# perService: operator-managed per-InferenceService PVC that binds on the node
# where the inference pod schedules. A separate user-managed override is
# available via spec.modelCache.claimName on individual InferenceServices.
mode: shared
# Storage size for model cache
size: 100Gi
# Storage class (leave empty for default)
storageClass: ""
# Access mode
# - ReadWriteOnce: Single-node clusters
# - ReadWriteMany: Multi-node clusters (requires NFS, EFS, etc.)
accessMode: ReadWriteOnce
# Mount path inside controller pod
mountPath: /models
# PVC annotations (e.g., for backup policies)
annotations: {}
Multi-Node Clusters
For multi-node clusters where pods may run on different nodes, you need a storage class that supports ReadWriteMany:
AWS EKS (EFS):
modelCache:
storageClass: efs-sc
accessMode: ReadWriteMany
GKE (Filestore):
modelCache:
storageClass: filestore-standard
accessMode: ReadWriteMany
Azure AKS (Azure Files):
modelCache:
storageClass: azurefile-premium
accessMode: ReadWriteMany
On-Premise (NFS):
modelCache:
storageClass: nfs-client
accessMode: ReadWriteMany
Strictly tainted GPU nodes
When a GPU node carries a hard NoSchedule taint, dynamic provisioning can fail for reasons unrelated to the inference pod's tolerations. The safest approach is to pre-provision a PVC bound to the tainted node and reference it through spec.modelCache.claimName.
The inference pod already receives the GPU toleration derived from the Model's hardware.gpu configuration. The model-downloader runs as an init container in that same pod, so a pre-existing PVC does not require a separate download Job.
The claim must already exist and be node-aligned by the cluster administrator's storage setup. claimName is user-owned: LLMKube mounts it through the same prep and download init containers as the built-in cache, but never creates, mutates, or deletes it. If the named PVC is missing, the InferenceService is marked Degraded rather than silently falling back to the shared cache.
Note: claimName is ignored for pvc:// model sources because those weights are already staged on the cluster and mounted read-only; no download occurs.
Some dynamic provisioners create a per-node helper pod to provision volumes. This helper pod is separate from the LLMKube inference pod and typically carries only the default not-ready/unreachable tolerations. When a GPU node has a hard NoSchedule taint, the helper pod may remain Pending with an untolerated-taint event — even though the inference pod itself has the correct GPU toleration.
Strict-taint choices:
- Pre-provision a static, node-aligned claim and use
claimName. - Use a provisioner whose helper configuration supports the taint.
- Apply a narrowly scoped cluster-level policy maintained by the cluster administrator.
Recognising a volume the provisioner never stamped
The failure above is worth calling out separately because it does not look like a storage failure. Some provisioners publish a PersistentVolume even when their helper never ran, carrying a node affinity that no node satisfies:
$ kubectl get pv <pv> -o jsonpath='{.spec.nodeAffinity}'
{"required":{"nodeSelectorTerms":[{"matchExpressions":[
{"key":"kubernetes.io/hostname","operator":"In","values":[""]}]}]}}
An empty value in that In term is the fingerprint: node label values are never
empty, so the volume can never be scheduled anywhere. The PVC still reports
Bound, which is why this reads as success.
The consuming pod then sits Pending forever, and the scheduler describes it in
one of two ways depending on whether the pod had already been assigned a node:
0/4 nodes are available: 1 node(s) didn't match PersistentVolume's node affinity
running PreBind plugin "VolumeBinding": binding volumes: pv "pvc-..."
node affinity doesn't match node "gpu-node-1": no matching NodeSelectorTerms
Both name node affinity, which sends you to the scheduler. The scheduler is not
at fault: it chose correctly, and the PVC's volume.kubernetes.io/selected-node
annotation agrees with it. Provisioning is what failed.
LLMKube reports this directly rather than passing the scheduler's wording through:
$ kubectl get inferenceservice <name> -o jsonpath='{.status.schedulingStatus}'
UnbindableModelCache
$ kubectl get inferenceservice <name> -o jsonpath='{.status.schedulingMessage}'
model cache volume was never provisioned: PersistentVolume "pvc-..." has a node
affinity that matches no node. Claim: <name>-model-cache. ...
To recover, delete the claim so its volume is released, then provision on a storage class whose helper can run on the tainted node. Which objects you have to remove depends on the cache mode:
| Cache mode | Deleting the InferenceService clears it |
|---|---|
perService | Yes, the claim is owner-ref'd and garbage-collected with it |
shared (default) | No, the shared claim intentionally outlives any one service |
claimName (bring your own) | No, LLMKube never deletes a user-owned claim |
In the latter two cases the unusable volume survives the InferenceService, so
delete the claim explicitly before retrying. Recreating the service against a
claim that is still bound to a broken volume reproduces the same Pending state.
Choosing a provisioner that tolerates taints
Whether this is fixable by configuration depends on the provisioner:
| Provisioner | Helper tolerations configurable |
|---|---|
rancher.io/local-path | Yes, via helperPod.yaml in its ConfigMap |
microk8s.io/hostpath | No, the helper template is built into the image |
For local-path, add tolerations to the helper template so it can run wherever a
volume is legitimately requested. This mirrors what a CSI node plugin does:
# local-path-config ConfigMap, helperPod.yaml key
apiVersion: v1
kind: Pod
metadata:
name: helper-pod
spec:
priorityClassName: system-node-critical
tolerations:
- operator: Exists
containers:
- name: helper-pod
image: busybox
Then point model caches at it:
modelCache:
storageClass: local-path
microk8s.io/hostpath exposes no equivalent knob, so on a tainted node the
remaining options are the pre-provisioned claimName route above or a different
storage class.
Per-InferenceService Cache PVC (Bring Your Own)
The cache backend above is an operator-global choice. To point a single
InferenceService at its own pre-existing, user-owned PVC — for example a
node-local volume for a large model pinned to one node, while everything else
rides the shared cache — set spec.modelCache.claimName:
apiVersion: inference.llmkube.dev/v1alpha1
kind: InferenceService
metadata:
name: llama-3.1-70b
spec:
modelRef: llama-3.1-70b
modelCache:
claimName: my-model-cache # pre-existing PVC in the same namespace
Behavior:
- The named PVC becomes the writable model cache for this workload only: the
same
model-cache-prepandmodel-downloaderinit containers run against it, weights land under the usual<cacheKey>/subdirectory, and the serving container mounts it read-only.RefreshPolicyand cache-key semantics are unchanged, so multiple models can safely share one claim. - The operator never creates or deletes the claim — you own it end-to-end
(unlike
perServicemode, where the operator provisions and garbage-collects<isvc>-model-cache). If the claim does not exist, the InferenceService is markedDegradedwith aModelCachePVCNotFoundevent instead of silently falling back to the shared cache. claimNametargets the download path, so it is ignored — with aModelCacheClaimIgnoredwarning event — for pre-stagedpvc://model sources (mounted read-only, no download), and whenever caching is inactive (chartmodelCache.enabled: false, a localfile://source, or a remote model not yet fingerprinted); in the inactive-caching cases the pod falls back to an ephemeral emptyDir and re-downloads on every restart.claimNameis mutable: changing it rolls the Deployment onto the new claim, but the old claim (and the weights on it) is not garbage-collected — clean it up yourself if it is no longer needed.- Node alignment is your responsibility: for an RWO or node-local claim, use
nodeSelectorso the pod lands where the PVC binds (aWaitForFirstConsumerlocal class binds on the first consumer; a pre-bound RWO PVC pins the pod). llmkube cache listdiscovers the shared cache and operator-managed per-service cache PVCs. A user-managedspec.modelCache.claimNamePVC is outside the operator's cache label/discovery contract and may not appear in the listing. Cache inspection may need a running pod or a transient inspector pod for PendingWaitForFirstConsumerclaims.
CLI Commands
List Cached Models
# List cached models in default namespace
llmkube cache list
# List from all namespaces
llmkube cache list -A
Output:
Model Cache Entries
═══════════════════════════════════════════════════════════════════════════════
CACHE KEY SIZE MODELS SOURCE
a3b8c9d4e5f67890 4.1 GiB llama-2-7b ...TheBloke/Llama-2-7B-GGUF/...
f1c314277254a2fd 7.2 GiB llama-3.1-8b ...meta-llama/Meta-Llama-3.1-8B/...
Total: 2 cache entries, 2 models
Clear Cache
# Clear cache for a specific model
llmkube cache clear --model llama-2-7b
# Clear all cache (with confirmation)
llmkube cache clear
# Force clear without confirmation
llmkube cache clear --force
Preload Models
Pre-download models before deploying them:
# Preload a catalog model
llmkube cache preload llama-3.1-8b
# Preload to a specific namespace
llmkube cache preload llama-3.1-8b -n production
This is useful for:
- Air-gapped environments (pre-populate cache on a connected machine)
- Reducing deployment time (model already cached)
- Bandwidth management (download during off-peak hours)
Air-Gapped Deployments
For air-gapped environments:
-
On a connected machine, preload models:
llmkube cache preload llama-3.1-8b llmkube cache preload mistral-7b -
Export the PVC. The weights live in the cache PVC of the namespace you preloaded into, not in the controller pod. The controller's own
/modelsis scratch space, and it never mounts a workload namespace's claim, so copying out ofllmkube-systemcannot reach them.Find the claim first. In
sharedmode (the default) it isllmkube-model-cache; inperServicemode it is<inferenceservice>-model-cache:kubectl -n <namespace> get pvcThen mount it in a throwaway pod and copy from there:
kubectl -n <namespace> run cache-export --restart=Never --image=busybox:1.36 \ --overrides='{"spec":{"containers":[{"name":"cache-export","image":"busybox:1.36","command":["sleep","3600"],"volumeMounts":[{"name":"cache","mountPath":"/models"}]}],"volumes":[{"name":"cache","persistentVolumeClaim":{"claimName":"llmkube-model-cache"}}]}}' kubectl -n <namespace> wait --for=condition=Ready pod/cache-export --timeout=120s kubectl cp <namespace>/cache-export:/models ./model-cache kubectl -n <namespace> delete pod cache-exportA ReadWriteOnce claim can only be mounted where it is already attached, so scale the InferenceService to zero first if the export pod will not schedule.
-
On the air-gapped cluster, import the cache by reversing the copy against a helper pod mounting the destination claim:
kubectl -n <namespace> run cache-import --restart=Never --image=busybox:1.36 \ --overrides='{"spec":{"containers":[{"name":"cache-import","image":"busybox:1.36","command":["sleep","3600"],"volumeMounts":[{"name":"cache","mountPath":"/models"}]}],"volumes":[{"name":"cache","persistentVolumeClaim":{"claimName":"llmkube-model-cache"}}]}}' kubectl -n <namespace> wait --for=condition=Ready pod/cache-import --timeout=120s kubectl cp ./model-cache/. <namespace>/cache-import:/models kubectl -n <namespace> delete pod cache-importThe destination claim must exist before the copy. Deploy the model once and let it fail to download, or create the PVC yourself, so the operator has a claim to bind.
-
Deploy models (they'll be found in cache):
llmkube deploy llama-3.1-8b --gpu
Troubleshooting
Model Not Using Cache
If models are still being downloaded via init container:
-
Check if the Model has a CacheKey:
kubectl get model llama-3.1-8b -n <namespace> -o jsonpath='{.status.cacheKey}' -
Verify the controller has cache enabled:
kubectl get deploy -n llmkube-system llmkube-controller-manager -o yaml | grep model-cache -
Check PVC exists in your namespace:
kubectl get pvc llmkube-model-cache -n <namespace> -
Check if model is cached in the PVC:
kubectl exec -n <namespace> <pod-name> -- ls -la /models/
Cache PVC Full
If the cache PVC runs out of space:
-
List cache entries:
llmkube cache list -n <namespace> -
Clear unused models:
llmkube cache clear --model <unused-model> -n <namespace> -
Or resize the PVC (if your storage class supports it):
kubectl patch pvc llmkube-model-cache -n <namespace> \ -p '{"spec":{"resources":{"requests":{"storage":"200Gi"}}}}'
Cache Corruption
If you suspect cache corruption:
-
Clear the specific cache entry by deleting the directory in the PVC:
# Find a pod in the namespace to exec into kubectl exec -n <namespace> <pod-name> -- rm -rf /models/<cache-key> -
Delete and recreate the InferenceService to trigger re-download:
kubectl delete inferenceservice llama-3.1-8b -n <namespace> kubectl apply -f inferenceservice.yamlOr delete and recreate the Model:
kubectl delete model llama-3.1-8b -n <namespace> kubectl apply -f model.yaml
Performance Considerations
- Storage Performance: Use SSD-backed storage for faster model loading
- Network: For ReadWriteMany, ensure low-latency network between nodes and storage
- Cache Size: Plan for 1.5-2x your total model sizes to allow for cache rotation
Disabling Cache
To disable persistent caching (not recommended):
# values.yaml
modelCache:
enabled: false
This will revert to the legacy behavior where each pod downloads the model via init container.
Security Considerations
Local and hostPath model sources (host-path allowlist)
Besides remote downloads, the model storage machinery can serve node-local sources: a Model whose spec.source is an absolute path (/srv/models/model.gguf) or a file:// URL is mounted into the inference pod via a hostPath volume. As of v0.9.0, these sources are gated by an operator-configured allowlist (GHSA-jw3m-8q7m-f35r).
The allowlist is empty by default, which disables local and hostPath sources entirely. With the default configuration, a Model with a local source is rejected: it goes to phase Failed with a SourceNotAllowed condition, nothing is fetched, and no hostPath volume is created. Remote sources (https://, pvc://, hf://) are unaffected.
To keep serving models from node-local disk, list the absolute directory roots you trust:
# values.yaml
modelSource:
allowedHostPathRoots:
- /srv/models
The equivalent controller flag is --allowed-host-path-roots (comma-separated). A source is permitted only when its cleaned path is within one of the configured roots; .. escapes are rejected. The check is lexical and does not resolve symlinks, so only allowlist roots whose contents you control.
Automatic shared cache on fsGroupPolicy: None backends (CephFS, NFS)
The automatic shared-model-cache workflow is a first-class supported path, including on fsGroupPolicy: None CSIs such as CephFS and NFS. On those backends Kubernetes does not apply the pod fsGroup to the volume, so the PVC root stays root:root and the non-root model downloader cannot write to it. The model-cache-prep init container fixes the ownership for you so the cache just works. You do not need to pre-stage models or hand-manage a PVC.
The prep is the compatibility implementation for these backends, and it runs as non-privileged root (uid 0) with ALL capabilities dropped and only CHOWN+FOWNER added (no privilege escalation, read-only rootfs, seccomp RuntimeDefault). Root here is a backend-driven requirement, not an avoidable misconfiguration: chowning a root-owned mount needs CAP_CHOWN, the prep shells out to chown, and a non-root process loses its capabilities across execve (Kubernetes has no ambient-capability field), so the exec'd chown would fail EPERM. Root retains its capabilities across exec. (Running the prep non-root broke the shared cache on these backends in 0.8.20; reverted in 0.8.21. A future change may move the chown into a small dedicated helper that performs the syscall in its own entrypoint, which would let the prep run non-root again without the exec capability loss.)
PSA restricted
Because the prep needs root plus CHOWN/FOWNER, it cannot satisfy Pod Security Admission restricted (which forbids running as root and adding any capability except NET_BIND_SERVICE). This is a property of the storage backend, not of the cache feature: it applies only in the narrow intersection of an fsGroupPolicy: None backend AND a namespace that enforces restricted. If you are in exactly that case, the options are:
- Use an
fsGroupPolicy: FileCSI so Kubernetes applies ownership itself and no prep is needed. - Use an
emptyDirmodel store (no shared persistent cache). - Relax the namespace policy to
baseline.
Shared-cache group-write multi-tenancy
In shared mode, the cache-prep init container sets chmod g+rwX /models, making the cache root writable by any pod in the same fsGroup. On a shared cache PVC, multiple InferenceServices share that fsGroup and can write to each other's cached model files.
Within one operator's trust domain this is fine. For a multi-tenant cache where distrusting tenants share the same namespace, this is a real consideration. Use per-service caches (modelCache.mode: perService) for isolation between distrusting tenants.