Generating Bundles
September 18, 2026 · View on GitHub
aicr bundle materializes a recipe into deployment-ready artifacts — one
folder per component, each with Helm values, checksums, and a README. This
guide covers the common bundling tasks: choosing a deployer, overriding values,
enabling or disabling components, pinning node scheduling, producing offline
bundles, and gating on component readiness.
This is a task-oriented how-to. For the complete flag list and exit codes, see
the aicr bundle section of the CLI Reference. For the
recipe → bundle → deploy → validate flow end to end, see the
End-to-End Tutorial.
Choose a deployer
The --deployer (-d) flag selects the output format. The bundle content is
the same validated configuration; only the serialization differs, so you can
re-render the same recipe for whatever pipeline you run:
| Deployer | Output |
|---|---|
helm (default) | Per-component Helm values + a deploy.sh that installs in dependency order. |
helmfile | A helmfile.yaml release graph. |
argocd | Argo CD Application manifests (app-of-apps), published from a Git repo (--repo). |
argocd-helm | A Helm chart app-of-apps; repoURL defaults to the push-target registry — plain helm install works with no --set repoURL needed. Override with --set repoURL=oci://mirror when mirroring. Bringing your own root Application? Set deployer.includeRootApp=false to render children-only — see Argo CD Deployer Options. |
flux | Flux HelmRelease manifests plus their source objects, and a plain Kustomize kustomization.yaml at the bundle root (not a Flux Kustomization CR). |
# GitOps with Argo CD, sourced from your config repo
aicr bundle --recipe recipe.yaml --deployer argocd \
--repo https://github.com/my-org/my-gitops-repo.git \
--output ./bundles
Bundle layout
This is the canonical description of what a bundle contains. The trees below
are frozen at v1 and gated by TestBundleLayoutMatchesManifest, so a path
shown here will not disappear or be renamed without a deliberate, reviewed
change. Automation may read these paths.
Every deployer writes bundle-info.yaml, checksums.txt, README.md and
recipe.yaml at the bundle root. Four of the five group components into
ordered NNN-<component> directories; Flux is the exception and uses a plain
<component> directory with shared sources/.
helm/ argocd/ flux/
001-cert-manager/ 001-cert-manager/ cert-manager/
values.yaml application.yaml helmrelease.yaml
cluster-values.yaml values.yaml nfd/
install.sh 002-nfd/ helmrelease.yaml
upstream.env application.yaml sources/
002-nfd/ values.yaml helmrepo-<host>.yaml
... app-of-apps.yaml gitrepo-<host>.yaml
deploy.sh bundle-info.yaml kustomization.yaml
bundle-info.yaml checksums.txt bundle-info.yaml
recipe.yaml README.md checksums.txt
checksums.txt recipe.yaml README.md
README.md recipe.yaml
helmfile shares Helm's per-component files but not its root: it writes
helmfile.yaml instead of deploy.sh. A recipe with dependencies also
produces one level-N.yaml per dependency depth, which is derived from the
recipe rather than fixed by the layout.
helmfile/
001-cert-manager/ same four files as helm
002-nfd/
helmfile.yaml
level-N.yaml one per dependency depth; absent when flat
recipe.yaml
bundle-info.yaml
checksums.txt
README.md
argocd-helm renders a Helm chart at the root — Chart.yaml, values.yaml,
values.schema.json — with one template per component.
Two kinds of name appear in these trees, and only one is a promise:
- Fixed names are contract.
deploy.sh,app-of-apps.yaml,kustomization.yaml,bundle-info.yaml,checksums.txt,values.yaml,helmrelease.yaml, and theNNN-<component>convention itself. - Derived names are not. Flux writes one
helmrepo-<host>.yamlper chart repository and helmfile onelevel-N.yamlper dependency depth, so both sets change with the recipe. Discover them by listing the directory rather than hardcoding a name.
Bundle info
Every bundle carries a bundle-info.yaml at its root, written unconditionally
by all five deployers — no flag turns it off. It answers three questions a
bundle cannot answer for itself: which deployer built it, which aicr binary
built it, and which Helm release landed in which directory.
layout.entrypoint names the file a consumer invokes or applies —
deploy.sh, helmfile.yaml, app-of-apps.yaml, Chart.yaml, or
kustomization.yaml — so automation reads one key instead of branching on
build.deployer.
layout.releases lists every Helm release the bundle installs, in deployment
order, and that ordering is normative: there is no ordinal field, so a
consumer reads sequence from list position rather than from a number. A
release injected alongside a component — a -pre folder when it declares
pre-install manifests, a -post folder when a component that also ships an
upstream chart declares post-install manifests, or a -readiness folder
under --readiness-hooks (none of the three tied to --vendor-charts) — has
no recipe component of its own, so it names its parent in component while
name carries its own suffixed name. One collision is deliberately left
undefined: a recipe that declares a component whose own name ends in -pre,
-post or -readiness alongside the matching base name makes component
deployer-dependent for that release, so do not rely on its value there. No
component in the shipped registry has such a name, and resolving this is a
reserved additive change — a later release may add an explicit field
distinguishing a primary release from an injected one.
bundle-info.yaml deliberately carries no component inventory:
recipe.yaml sits beside it at the bundle root and is already the source of
truth for what the recipe resolved to. build.settings.components in the
example below is not an inventory — it is the REST API's ?bundlers= filter
(there is no CLI flag for it), recorded because the filtered recipe.yaml
alongside it is written post-filter and would otherwise be indistinguishable
from an unfiltered bundle of a smaller recipe. It is absent unless that
filter was used; see Overrides that cannot take effect are
rejected for the same
bundlers= filter from the override side. bundle-info.yaml also carries no
timestamp — the record feeds checksums.txt, which is the subject of the
bundle attestation, so a wall-clock field would make every bundle
irreproducible.
Below is a fully populated example, generated from the shipped serializer and
extended with fields (repoURL, provenance, an injected -post/-readiness
pair) that a single bundle rarely exercises all at once. Key order is
alphabetical at every nesting level — serializer.MarshalYAMLDeterministic
sorts mapping keys before writing — so the top level reads apiVersion, build, kind, layout, metadata, not a hand-arranged order. Expect the same
alphabetical order in your own bundle.
apiVersion: aicr.run/v1
build:
deployer: argocd
recipe:
digest: sha256:d429490ec53e902e5a7f5a9b0221ab4a46f233ec70f808f1e62bbade2576e083
path: recipe.yaml
version: v0.22.0
settings:
appName: gpu-cluster
attested: true
checksums: true
components:
- cert-manager
- nfd
- network-operator
nodeScheduling:
accelerated:
selector:
nvidia.com/gpu.present: "true"
tolerations:
- effect: NoSchedule
key: nvidia.com/gpu
operator: Equal
value: present
system:
selector:
nodeGroup: system
tolerations:
- effect: NoSchedule
key: dedicated
operator: Equal
value: system
readinessHooks: true
repoURL: https://github.com/my-org/my-gitops-repo.git
serial: false
sharedStorageClass: shared-nfs
storageClass: local-nvme
targetRevision: main
vendorCharts: true
kind: BundleInfo
layout:
entrypoint: app-of-apps.yaml
provenance: provenance.yaml
releases:
- component: cert-manager
manifest: 001-cert-manager/application.yaml
name: cert-manager
namespace: cert-manager
path: 001-cert-manager
- component: nfd
manifest: 002-nfd/application.yaml
name: nfd
namespace: node-feature-discovery
path: 002-nfd
- component: network-operator
manifest: 003-network-operator/application.yaml
name: network-operator
namespace: network-operator
path: 003-network-operator
- component: network-operator
manifest: 004-network-operator-post/application.yaml
name: network-operator-post
namespace: network-operator
path: 004-network-operator-post
- component: network-operator
manifest: 005-network-operator-readiness/application.yaml
name: network-operator-readiness
namespace: network-operator
path: 005-network-operator-readiness
metadata:
version: v0.22.0
build.recipe.version and metadata.version can diverge: the former is the
aicr binary that resolved recipe.yaml (stamped by the recipe builder and
never restamped), the latter is the aicr binary that ran bundle. They read
the same on a one-step workflow and differ whenever recipe generation and
bundling run on different releases. build.settings includes only settings
whose effect is already observable elsewhere in the bundle's own files —
Fulcio/Rekor endpoints, certificate identity, and free-form --set overrides
are excluded by construction, since values.yaml already carries the effect
of the latter.
That test is applied per deployer, so repoURL, targetRevision and
appName appear only where the bundle shows them: all three under argocd,
repoURL and targetRevision under flux, appName alone under
argocd-helm (whose chart is URL-portable and takes the publish location at
helm install --set repoURL=... time, which is also why --repo warns that
it is ignored there), and none of the three under helm or helmfile. Where
a key does not apply it is omitted entirely rather than written empty.
Each deployer reports what it resolved, so the recorded value is the one the
bundle carries — including the fallback it applies when you pass no flag. A
flux bundle built without --repo records
repoURL: https://github.com/YOUR_ORG/YOUR_REPO.git, the same placeholder
written into sources/gitrepo-*.yaml, and targetRevision: main. That is
deliberate: the placeholder is what ships, and a bundle that needs its repo
URL replaced before it can be applied should say so rather than look
unconfigured.
Generated chart versions
Not every chart in a bundle comes from upstream. AICR generates one for each
Kustomize-derived component, each manifest-only component, each injected
-pre / -post / -readiness release, and — under --vendor-charts — as a
wrapper around every vendored upstream chart. A generated Chart.yaml has to
answer two different questions, so it answers them in separate fields:
| Field | Carries | Why |
|---|---|---|
version | the AICR version that produced the bundle | The chart's content is AICR's own, so AICR's version identifies the artifact. This is the aicr binary that ran bundle, which is not necessarily the one that generated the recipe. |
appVersion | the payload version | The upstream chart pin for a Helm component, the git ref for a Kustomize one. Free-form, so a ref like release-1.4 need not look like SemVer. |
aicr.run/component-version annotation | the payload version | The same value as appVersion, under a stable key to read from a live release. |
aicr.run/generated-by annotation | the AICR version | The same value as version. |
To read the payload version back out of a cluster, the rule is: use
aicr.run/component-version when it is present, otherwise use the release's
own chart version. A component installed straight from its upstream chart
carries neither annotation, and its release version already is the payload
version. The annotation's presence is exactly the signal that the chart
version describes the wrapper instead of what it wraps.
So for the gpu-operator-post release generated alongside gpu-operator, a
helm list reports chart gpu-operator-post-0.22.0 — the AICR version — while
its app version and aicr.run/component-version both read v26.7.0, the
gpu-operator pin those manifests accompany.
A component with no upstream pin — a manifest-only component, and the injected
releases belonging to one — ships only AICR-authored content, so appVersion
and the annotation report the AICR version too.
A build that is not release-stamped (aicr --version reports dev) reports
0.0.0-dev. Helm validates version: as SemVer 2 and refuses to load a chart
whose version is not, so dev cannot be written through verbatim.
The root chart argocd-helm renders is not one of these generated wrappers and
carries neither annotation. Its version: tracks the recipe's
metadata.version — the AICR build that generated the recipe, not the one
that ran bundle — so on a two-step workflow where the two binaries differ, it
will not match the wrappers alongside it. Only the 0.0.0-dev substitution
above is shared.
Verify a bundle you received with aicr verify — see
Artifact verification.
Override values
Use --set for scalar overrides, scoped per component as
component:path.to.field=value:
aicr bundle --recipe recipe.yaml \
--set gpuoperator:driver.version=570.86.16 \
--set gpuoperator:gds.enabled=true \
--output ./bundles
On recipes that carry an ADR-015 configuration profile
(metadata.selectedProfile, e.g. the AKS family's gpuStack), the
profile's owned paths are locked. The lock is enforced per surface:
aicr bundlestatic overrides (--set,--set-json,--set-file, or a config-file override, from any of those sources): a value identical to the selected one is accepted; a divergent value is rejected. The typed sources (--set-json,--set-file) are always rejected for the specialenabledpresence key — even when the value matches — because they would write a stray literalenabled:chart value instead of toggling the component.aicr mirror list --setoverrides: mirror exposes only the repeatable scalar--set(no--set-json/--set-file), and the same identical-accepted / divergent-rejected rule applies to it. Note thatmirror listdoes not apply a config file'sspec.bundle.deployment.setoverrides — pass image-affecting overrides to mirror via--setexplicitly.--dynamicexports: rejected whenever they merely intersect an owned path, regardless of the current value — install-time mutability of a locked path is itself the violation.- argocd-helm install-time values: any install-time key whose path equals, contains, or is contained by an owned path is rejected at Helm render time, even when the value is identical. Key presence alone trips the guard.
- Component presence (the synthetic
enabledowned path): not changeable by selecting a different profile value — profile fragments cannot assignenabled, so no--profilechoice adds or removes a component. The lock rejects removing (or bundle-subsetting away) an owned component; changing which components are present is a catalog/composition change.
Change owned value paths by regenerating with
aicr recipe --profile name=value instead; component presence is not
affected by reselection.
--set is scalar-only. For list or object values use --set-json (inline JSON)
or --set-file (value read from a file); both deep-merge objects and replace
lists/scalars, and take precedence over --set on the same path:
aicr bundle --recipe recipe.yaml \
--set-json agentgateway:allowedSourceRanges='["216.228.127.128/30"]' \
--output ./bundles
The agentgateway inference-gateway is private by default: with no
allowedSourceRanges override, the bundler scopes the LoadBalancer to private
RFC1918 ranges so it is not exposed to the public internet. Use the override
above to admit specific clients (e.g. a corporate VPN). See
Inference Gateway Network Exposure.
Enable or disable components
The special enabled key includes or excludes a component at bundle time
without editing the recipe:
# Skip the AWS EBS CSI driver for this bundle
aicr bundle --recipe recipe.yaml \
--set awsebscsidriver:enabled=false \
--output ./bundles
A recipe or overlay can also disable a component by default via
overrides.enabled: false (for example, a platform that ships its own
cert-manager). Such components are already excluded from the recipe's
deploymentOrder.
--set <component>:enabled=false disables a component the recipe leaves on.
A component the recipe disabled cannot be re-enabled at bundle time —
--set <component>:enabled=true on such a component is rejected with an error.
The recipe author disables a component because the target platform already
provides it, so re-enabling would install a conflicting second copy. To deploy
a component the recipe disables, edit the recipe/overlay instead.
Overrides that cannot take effect are rejected
An override whose component will not appear in the generated bundle is
rejected with an error rather than silently discarded. This covers a
component that is absent because the recipe disabled it, because
--set <component>:enabled=false removed it, because the bundlers=
filter excluded it, or because the component name is neither one the recipe declares nor a
registered valueOverrideKeys alias of one (usually a typo — registered
aliases such as gpuoperator for gpu-operator remain valid):
# Rejected: the second --set can never take effect
aicr bundle --recipe recipe.yaml \
--set nv-sentinel:enabled=false \
--set nv-sentinel:labeler.assumeDriverInstalled=true \
--output ./bundles
The two flags ask for contradictory things — remove the component, and
configure it — so the command fails instead of shipping a bundle with
one request quietly dropped. The rejection here is about the contradiction,
not about who owns the component: a profile-owned presence lock is a
separate rejection, described at the end of this section. Only a scalar --set <component>:enabled=false
is exempt on a declared component: it is the supported way to remove
one, and it is also accepted on a component the recipe already disables.
enabled=true on a component the bundlers= filter excludes is
rejected like any other ineffective override, and the enabled key is
never honored from --set-json/--set-file (present or absent — the
typed path would write a literal enabled: chart value instead of
toggling the component).
A component whose presence a configuration profile owns cannot be removed
at all — enabled=false on it is rejected regardless of the rule above.
NVSentinel is in that position on the AKS and GKE-COS families, whose
gpuStack profiles name it; see
NVSentinel on provider-installed-driver platforms.
The same rule applies to --set-json, --set-file, --dynamic, and
the REST API's equivalent parameters. For --dynamic no path is exempt,
enabled included: a dynamic path on an absent component exports
nothing (there is no cluster-values.yaml to defer it to), and a
dynamic path is never a removal idiom. Unknown component names were
already rejected by the bundlers= filter and by --dynamic
registry validation; this extends the same fail-closed rule to every
override source.
Pin node scheduling
Steer system components and GPU workloads onto the right nodes with selector and toleration flags (repeatable):
aicr bundle --recipe recipe.yaml \
--system-node-selector nodeGroup=system \
--system-node-toleration dedicated=system:NoSchedule \
--accelerated-node-selector nvidia.com/gpu.present=true \
--accelerated-node-toleration nvidia.com/gpu=present:NoSchedule \
--output ./bundles
Prepare DRA nodes when opting in to eviction coordination
DRA eviction coordination is opt-in. By default a bundle containing both
gpu-operator and nvidia-dra-driver-gpu adds no eviction node label, and the
DRA kubelet plugin runs on every accelerated node with no extra labeling. The
trade-off is that the plugin is not descheduled ahead of a GPU driver container
restart; aicr bundle warns about this where GPU Operator manages the driver,
and DRA Driver Upgrade Eviction
describes what can go wrong.
Generate with --dra-eviction-node-label key=value to opt in. The rest of this
section applies only then. The same applies to the corresponding -ocp
components.
Opting in also keeps dra-node-labeler in the bundle. It applies the
configured key=value to every node GFD labels nvidia.com/gpu.present=true
and never rewrites an existing value, so the node-pool labeling described below
is only needed when the labeler is not in the bundle: because you removed it
with --set dra-node-labeler:enabled=false, because a bundlers filter left
it out, or on OpenShift, where it is not yet wired (NVIDIA/aicr#2828). A
bundlers selection that names the labeler but omits the flag or one of its
prerequisites is rejected rather than rendered without it. Everything below that says "labeler disabled" applies to
those cases equally.
Choosing whether to opt in
The label is how GPU Operator's Driver Manager finds the plugin: it deschedules
the plugin by rewriting the label's value and restores it afterwards, so the
plugin's nodeSelector has to match the label for the mechanism to work at all.
That is what makes it a placement requirement, and why an unlabeled node ends up
with no plugin rather than with an uncoordinated one.
| Not opted in (default) | Opted in | |
|---|---|---|
| Node labeling | none needed | applied by dra-node-labeler from nvidia.com/gpu.present; every GPU node in the node pool definition only if the labeler is disabled |
| Plugin placement | every accelerated node | only nodes carrying the label |
| Driver restart | plugin is not descheduled first | plugin is descheduled and restored |
| If a node is missed | n/a | that node silently runs without DRA |
Opt in when GPU Operator manages the driver (driver.enabled=true). With the
bundled dra-node-labeler the label follows GFD's nvidia.com/gpu.present, so
nodes added later by autoscaling or replacement are labeled as soon as GFD sees
them; the remaining gap is a GPU node GFD has not labeled and that does not
already carry the configured key=value, which runs no kubelet plugin until
one of the two appears. If you disable the labeler, opt in only when you can
guarantee the label is set at provisioning time for every GPU node; otherwise
the default is the safer choice: a plugin that always runs, with a documented
risk at driver restarts, beats a plugin that silently does not run on some nodes.
There is nothing to opt in to where the driver is provider-installed
(driver.enabled=false — AKS azure-managed, GKE COS, OKE). GPU Operator
deploys no driver pod and therefore no Driver Manager, so no restart can occur
under the plugin and no warning is emitted.
Also note the mechanism is best-effort under k8s-driver-manager v0.12: the
configured label is paused in the same batch as other GPU operands and no wait
covers the standalone DRA kubelet plugin, so ordering against DRA claim holders
and completion of plugin teardown are not guaranteed. See
NVIDIA/k8s-driver-manager#250.
Opt-in requirement (labeler disabled only): if you pass
--set dra-node-labeler:enabled=false, label every GPU node that must run the DRA kubelet plugin before applying the bundle. Applying it first can reduce the DaemonSet to zero eligible nodes, interrupting ComputeDomain/IMEX and any whole-GPU resources advertised through DRA. With the labeler in the bundle there is nothing to pre-label; the plugin follows the labeler.
Set the label at node-pool provisioning time (labeler disabled)
This subsection applies only when dra-node-labeler has been disabled. Put the
label in the node pool definition — an EKS managed nodegroup labels
entry, a Karpenter NodePool spec.template.metadata.labels entry, or the
equivalent for your provisioner — alongside the nodeGroup=gpu-worker label
you already set there.
A one-off kubectl label node is a repair, not a configuration. It does not
survive node replacement or recycling, cluster autoscaling adding GPU nodes, or
a nodegroup scaled from zero. Any GPU node added afterwards arrives unlabeled
and silently runs without the DRA kubelet plugin, leaving the cluster
partially DRA-enabled. This is harder to detect than uniform failure,
because it is intermittent and node-dependent.
Use kubectl label only to repair nodes that already exist, and fix the node
pool definition in the same change so replacements inherit it. Substitute the
same key=value pair you passed to --dra-eviction-node-label — the examples
below use the documented default:
kubectl label node <node-name> nvidia.com/dra-kubelet-plugin=true
kubectl get nodes -l nvidia.com/dra-kubelet-plugin=true
The failure mode is silent
An unlabeled GPU node produces no error anywhere. helm install/helm upgrade reports success and the bundle's deploy.sh exits 0. What you get
instead is:
- no DRA kubelet plugin on any unlabeled node, and no
ResourceSlicesfrom it — so no ComputeDomain/IMEX capability there - the
nvidia-dra-driver-gpu-kubelet-pluginDaemonSet atDESIRED=0if no GPU node carries the label at all
Partial coverage is the shape node replacement and autoscaling produce: labeled nodes work normally while the rest silently lack DRA. A split cluster is harder to notice than uniform failure, because the DaemonSet looks healthy and only some workloads misbehave.
Because the absence is not self-announcing, confirm the selector matches the nodes you expect before applying the bundle, and check the DaemonSet afterwards.
This applies to existing clusters, not just fresh installs
The requirement is easy to read as a fresh-install prerequisite, but the
upgrade path is especially easy to miss. A cluster whose bundle was generated
before this selector existed has a working kubelet-plugin DaemonSet selecting
on nodeGroup=gpu-worker alone. Regenerating the bundle and running helm upgrade adds the second selector, and working functionality disappears —
still with no error. Revisit the node labels whenever you regenerate a bundle
for an existing deployment, not only when building a new cluster.
aicr bundle emits a non-blocking warning describing this requirement whenever
both components are enabled and an eviction label is configured. The
complementary opt-out warning fires only where GPU Operator manages the driver
(driver.enabled=true). See
Storage Class, where that warning is
described alongside the other cluster-state dependency reported the same way.
Upgrading from a build that applied the label implicitly
For a window on unreleased main, the eviction label was applied
automatically rather than requested. Bundles generated in that window — via the
CLI, the REST API, config, or a Go caller — received the
nvidia.com/dra-kubelet-plugin=true selector and the matching Driver Manager
environment entry without asking for them. No tagged release contains that
behavior, so only clusters built from main in that interval are affected.
At this head the same inputs produce the opposite result: omitting
--dra-eviction-node-label removes both the selector and the Driver Manager
entry. Regenerating and upgrading such a cluster therefore drops eviction
coordination silently — the reverse of the direction described above, and the
case the preceding subsection does not cover.
Decide deliberately, and only two answers are defensible:
- Keep coordination. Confirm the nodes still carry the label
(
kubectl get nodes -l nvidia.com/dra-kubelet-plugin=true), confirm the count matches the GPU nodes you expect, then pass--dra-eviction-node-label nvidia.com/dra-kubelet-plugin=trueexplicitly on every subsequent generation. Verify node labels before upgrading: passing the flag against unlabeled nodes takes the DaemonSet toDESIRED=0. - Accept the opt-out. Regenerate without the flag and accept the documented risks in DRA Driver Upgrade Eviction. The node labels become inert and can be removed at leisure.
The opt-out warning aicr bundle prints bounds the blast radius at generation
time, but it is not upgrade guidance: it fires on every unlabeled generation and
cannot know the cluster previously had coordination.
Custom label conventions and post-install checks
The flag both opts in and selects the convention: generate the bundle with
--dra-eviction-node-label key=value and apply that exact pair to the nodes.
AICR gives the full pair to the DRA node selector, but GPU Operator's Driver
Manager receives only the label key because its eviction contract matches and
temporarily removes the label by key.
After installation and after every GPU driver upgrade, monitor the kubelet plugin DaemonSet until all desired pods are ready. This also catches a Driver Manager rollout that did not restore the eviction label:
kubectl -n nvidia-dra-driver get daemonset \
nvidia-dra-driver-gpu-kubelet-plugin
The integration is not rendered when either component is absent. A dynamic
declaration intersecting kubeletPlugin.nodeSelector or driver.manager.env
is rejected because moving either path to install-time configuration would let
the two halves drift independently. See
DRA Driver Upgrade Eviction for
configuration details and NVIDIA's
GPU Operator DRA installation guide
for the upstream contract.
Produce an offline (vendored) bundle
--vendor-charts pulls upstream Helm chart bytes into the bundle at bundle
time, so the artifact needs no Helm chart registry egress at deploy time. Each
vendored chart is recorded in provenance.yaml with name, version, source URL,
and SHA256. Requires the helm binary on PATH.
aicr bundle --recipe recipe.yaml --vendor-charts --output ./bundles
Trade-off: vendoring freezes the chart version. A vendored bundle will keep installing a frozen chart even if upstream later yanks it for a CVE — you lose the fail-loud signal you get when pulling charts live. Container-image pulls may still require network access. For full air-gapped operation, also mirror images; see Air-Gap Mirror.
Recipe-side manifests of mixed components (AICR-authored manifests shipped
alongside a vendored upstream chart — for example the network-operator
NicClusterPolicy or the AKS nvidia-peermem-reloader DaemonSet) get the same
lifecycle under --vendor-charts as they do without it. The vendored primary
folder wraps only the upstream chart, and the manifests are emitted as a
separate <component>-post Helm release installed immediately after it:
002-network-operator/ # wrapper chart + charts/<chart>-<ver>.tgz
003-network-operator-post/ # recipe-side manifests, tracked release
Because they are ordinary members of that release rather than Helm hook
resources, they are patched in place by helm upgrade (three-way merge),
removed by helm uninstall, and applied normally by Argo CD under
syncPolicy.automated. Bundle-layer NNN- folder ordering sequences the two
releases, so any helm.sh/hook annotation a recipe manifest declares is
stripped when the -post chart is written — the same treatment the
non-vendored path has always applied.
Earlier releases injected these manifests into the vendored wrapper chart as
helm.sh/hook: post-installresources, which Helm never re-applied on upgrade, left behind on uninstall, and Argo CD silently skipped as a PostSync hook. Those live objects are not members of the new<component>-postrelease, so a redeploy of a rebundled layout fails with Helm ownership conflicts (exists and cannot be imported into the current release) until each resource is adopted or removed.
Adoption is the default path — it is non-destructive and required before
helm upgrade --install of the -post chart. For each previously
hook-injected resource (name and kind from the new
<NNN>-<component>-post/templates/ files, or from the old primary
templates/ if you still have that bundle):
# Include -n <ns> for namespaced kinds (same <ns> as release-namespace below).
# Omit -n for cluster-scoped kinds (CRD, ClusterRole, NicClusterPolicy, …).
kubectl label -n <ns> <kind>/<name> app.kubernetes.io/managed-by=Helm --overwrite
kubectl annotate -n <ns> <kind>/<name> \
meta.helm.sh/release-name=<component>-post \
meta.helm.sh/release-namespace=<ns> --overwrite
Prefer kubectl delete only for kinds where cascade is acceptable (for example
a ConfigMap or ClusterRole with no dependents). Do not kubectl delete -f
CRD or Namespace manifests from a migration cleanup — deleting a CRD
garbage-collects every CR of that type cluster-wide (Gateway API / Inference
Extension CRDs under agentgateway-crds are the concrete risk), and deleting a
Namespace removes everything inside it.
Gate on component readiness
--readiness-hooks emits a standalone readiness-gate chart for each component
that ships a readiness test, run as a post-component Job so the deployer blocks
on component-specific signals (e.g. GPU Operator ClusterPolicy state) that
Helm and Argo CD cannot assess natively. Supported with --deployer helm,
argocd, and argocd-helm; off by default.
aicr bundle --recipe recipe.yaml --readiness-hooks --output ./bundles
Deploy the bundle
For the default helm deployer, verify the bundle before installing it:
cd bundles && aicr verify . && chmod +x deploy.sh && ./deploy.sh
For GitOps deployers, commit/publish the manifests per your Argo CD or Flux workflow. Verify the bundle directory before it enters the repository or registry; a pull-based controller reconciles on its own schedule, so a pipeline step gates what gets published rather than what the cluster applies. For the per-deployer gates and which of them the cluster can enforce, see Gating Deployment on Verification.
After deploying, confirm the cluster matches the recipe with
aicr validate.