dpuflavor.md
September 10, 2026 ยท View on GitHub
[[TOC]]
Overview
DPUFlavor is a Kubernetes Custom Resource Definition (CRD) that defines configuration templates for DPU system-level settings. It serves as a blueprint that specifies how DPUs should be configured during provisioning, including kernel parameters, firmware settings, OVS configuration, network interfaces, and resource allocation.
API Version
- API Group:
provisioning.dpu.nvidia.com - API Version:
v1alpha1 - Kind:
DPUFlavor
Key Features
- Immutable Configuration: Once created, the DPUFlavor spec cannot be modified to ensure consistency across DPU deployments
- Comprehensive System Configuration: Covers all aspects of DPU system configuration from boot parameters to runtime settings
- Resource Management: Defines resource requirements and allocation policies
- Cluster deployment mode: Zero-trust vs host-trusted is configured on
DPFOperatorConfigand reflected onDPUstatus (deploymentMode), not onDPUFlavor - Template Reusability: Can be applied to multiple DPUs for consistent configuration
API Reference
DPUFlavorSpec
| Field | Type | Description |
|---|---|---|
grub | DPUFlavorGrub | All the parameters will be set in GRUB_CMDLINE_LINUX grub configuration |
sysctl | DPUFlavorSysctl | Kernel sysctl parameters which will be stored in /etc/sysctl.d/99-dpf.conf |
nvconfig | []NVConfig | The device configuration which will be applied by mlxconfig |
ovs | DPUFlavorOVS | Open vSwitch configuration applied by the DPU agent once per boot |
bfcfgParameters | []string | Parameters for the bf.cfg file. See BFCfg Parameters for important parameters |
configFiles | []ConfigFile | Custom configuration files. Users can use this configuration to overwrite files in the DPU file system or add content to existing files |
containerdConfig | ContainerdConfig | ContainerdConfig contains the configuration for containerd |
dpuResources | ResourceList | Minimum resources needed for BFB installation |
systemReservedResources | ResourceList | Resources reserved for system use |
hostNetworkInterfaceConfigs | []NetworkInterfaceConfig | Host-side network interface configuration |
dma | DPUFlavorDMA | Configures the DMA SF used by SNAP on BlueField-4 socket-direct systems |
serviceReadiness | ServiceReadiness | Configures the Service Readiness provisioning phase |
ServiceReadiness
Configures the Service Readiness provisioning phase.
| Field | Type | Description |
|---|---|---|
gate | string | The DPU.status.operationalConditions entry that must be True before the DPU leaves the Service Readiness phase. One of DPUServiceCriticalPodsReady or OperationalReady. Optional, with no default |
Unset is not the same as choosing a gate: the phase does not block, but a host hold requested through
DELAY_HOST_OS_INIT is still released on DPUServiceCriticalPodsReady. Setting gate makes the phase block on that
condition and points the host release at it too.
The hold itself is enabled by DELAY_HOST_OS_INIT=0x3 in NVConfig, not by this field, and requires Zero
Trust; the phase gate works in either deployment mode. See
Service Readiness.
DPUFlavorDMA
| Field | Type | Description |
|---|---|---|
enabled | *bool | Enables DMA SF creation by the dpu-agent. Defaults to false when unset, so the presence of the dma struct alone does not enable it. Only takes effect on BlueField-4 socket-direct systems. The created Scalable Function always uses sfnum 8000 (the SNAP discovery ABI value) |
DPUFlavorGrub
| Field | Type | Description |
|---|---|---|
kernelParameters | []string | Kernel boot parameters to be set in grub configuration |
DPUFlavorSysctl
| Field | Type | Description |
|---|---|---|
parameters | []string | Sysctl parameters to be applied |
NVConfig
| Field | Type | Description |
|---|---|---|
device | *string | Target device: "*" (all devices), "p0"/"P0" (port 0), or "p1"/"P1" (port 1). Case-insensitive. |
parameters | []string | Optional firmware parameters in KEY=VALUE format. When present, 1-64 entries, max 200 characters each; empty values such as FLAG= are allowed. |
force | *bool | Apply the parameters with mlxconfig --force and skip the mlxconfig q deferral filter. Defaults to false. |
Validation Constraints:
- Maximum of 3 nvconfig entries per DPUFlavor (one per device:
*,p0/P0,p1/P1) - Wildcard device (
"*") must be the sole entry when specified (no mixing with port-specific entries) - Device identifiers must be unique with case-insensitive matching (e.g.,
p0andP0are duplicates) - Parameters are optional. When present, provide 1-64
KEY=VALUEentries for each device with no whitespace; values may be empty
IB Mode to Ethernet Mode Configuration
Example for single port DPU:
nvconfig:
- device: '*'
parameters:
- LINK_TYPE_P1=ETH
Example for dual port DPU:
nvconfig:
- device: '*'
parameters:
- LINK_TYPE_P1=ETH
- LINK_TYPE_P2=ETH
Device-specific configuration (per-port):
nvconfig:
- device: 'p0'
parameters:
- LINK_TYPE_P1=ETH
- NUM_OF_VFS=8
- device: 'p1'
parameters:
- LINK_TYPE_P1=IB
Forcing parameters that firmware hides
Firmware hides some parameters from mlxconfig q until an enabling parameter is already active.
FORCE_ETH_PCI_SUBCLASS, for example, is invisible while ADVANCED_PCI_SETTINGS is off. A flavor
that sets both must set force: true:
nvconfig:
- device: '*'
force: true
parameters:
- ADVANCED_PCI_SETTINGS=1
- FORCE_ETH_PCI_SUBCLASS=1
Without force, such a flavor never converges: every pass resets the enabling parameter to its
default before validating the batch, so the hidden parameter is rejected every time. The failing run
commits that reset before it errors, leaving Next Boot at factory defaults and discarding the
flavor's other parameters too. With force, a single power cycle activates the whole batch.
To find out whether a parameter is hidden, ask mlxconfig rather than guessing. mlxconfig show_confs states the dependency explicitly, for example Configuration is available only when NV_PCI_CONF.ADVANCED_PCI_SETTINGS is TRUE. Which parameters are gated, and by which enabling
parameter, varies with firmware version and board.
Two costs to weigh before setting it:
forcerequires DOCA 3.5.0 or later. An older DOCA version accepts the flag but still rejects the hidden parameter, so the operation fails with an error such as-E- The Device doesn't support MAX_ACC_OUT_READ parameter. That message names the parameter rather than the cause, so check the DOCA version on the DPU when you see it on a flavor that setsforce.forceskips validation for the entire batch, not just the hidden parameter, so an invalid value is applied silently instead of being rejected. This is why it is opt-in per flavor.
force is honored only under spec.nvconfig. The field also appears under
spec.hostNetworkInterfaceConfigs[].nvconfig because both share a type, but nothing reads it
there.
DPUFlavorOVS
| Field | Type | Description |
|---|---|---|
rawConfigScript | string | Raw OVS configuration script. The DPU agent runs it once per boot; an agent restart in the same boot does not re-run it, a reboot or power cycle does |
ConfigFile
| Field | Type | Description |
|---|---|---|
path | string | File path on the DPU |
operation | DPUFlavorFileOp | File operation type (override or append) |
raw | string | File content |
permissions | string | File permissions (e.g., "0644") |
ContainerdConfig
| Field | Type | Description |
|---|---|---|
registryEndpoint | string | Container registry endpoint |
NetworkInterfaceConfig
| Field | Type | Description |
|---|---|---|
mtu | *int32 | MTU value (1280-9216) |
dhcp | *bool | Enable DHCP configuration |
portNumber | int32 | Port identifier (0 or 1) |
nvconfig | *NVConfig | Port-specific NVConfig settings |
Enumerations
DPUFlavorFileOp
override: Replace file content entirelyappend: Append to existing file content
BFCfg Parameters
The bfcfgParameters field accepts a list of KEY=VALUE strings that are written directly into the
bf.cfg file consumed by the BFB installer. Common parameters include:
| Parameter | Description |
|---|---|
UPDATE_ATF_UEFI | Update ATF/UEFI firmware during provisioning (yes/no) |
UPDATE_DPU_OS | Update the DPU operating system (yes/no) |
WITH_NIC_FW_UPDATE | Update NIC firmware during provisioning (yes/no) |
ubuntu_PASSWORD | Password hash for the ubuntu admin account on the DPU. If not set, DPF defaults to the well-known ubuntu:ubuntu credentials. It is strongly recommended to set this to a unique hashed password for every deployment |
Note
ubuntu_PASSWORD is filtered out of the generated bf.cfg and applied via cloud-init instead; all other parameters are written directly into the bf.cfg file.
To generate a SHA-512 password hash:
openssl passwd -6 'YourPassword'
Always use -6 (SHA-512). Do not use -1 (MD5), which is considered insecure.
For the full list of supported bf.cfg parameters, see the
BlueField BSP documentation.
Resource Management
DPU Resources
The dpuResources field specifies the minimum resources required for a BFB with this flavor to be installed on a DPU.
Using this field, the controller can understand if that flavor can be installed on a particular DPU. It should be set to the total amount of resources the system needs + the resources that should be made available for DPUServices to consume:
dpuResources:
cpu: 16
memory: 16Gi
nvidia.com/sf: 20
System Reserved Resources
The systemReservedResources field indicates resources consumed by the system (OS, OVS, DPF system etc) and are not made available for DPUServices to consume. DPUServices can consume the difference between dpuResources and systemReservedResources. This field must not be specified if dpuResources are not specified.:
systemReservedResources:
cpu: 4
memory: 4Gi
nvidia.com/sf: 4
The difference between dpuResources and systemReservedResources is available for DPUServices.
Examples
HBN-OVN DPUFlavor
apiVersion: provisioning.dpu.nvidia.com/v1alpha1
kind: DPUFlavor
metadata:
name: hbn-ovn
namespace: dpf-operator-system
spec:
bfcfgParameters:
- UPDATE_ATF_UEFI=yes
- UPDATE_DPU_OS=yes
- WITH_NIC_FW_UPDATE=yes
configFiles:
- operation: override
path: /etc/mellanox/mlnx-bf.conf
permissions: "0644"
raw: |
ALLOW_SHARED_RQ="no"
IPSEC_FULL_OFFLOAD="no"
ENABLE_ESWITCH_MULTIPORT="yes"
- operation: override
path: /etc/mellanox/mlnx-ovs.conf
permissions: "0644"
raw: |
CREATE_OVS_BRIDGES="no"
OVS_DOCA="yes"
- operation: override
path: /etc/mellanox/mlnx-sf.conf
permissions: "0644"
raw: ""
grub:
kernelParameters:
- console=hvc0
- console=ttyAMA0
- earlycon=pl011,0x13010000
- net.ifnames=0
- biosdevname=0
- iommu.passthrough=1
- cgroup_no_v1=net_prio,net_cls
- hugepagesz=2048kB
- hugepages=250
hostNetworkInterfaceConfigs:
- dhcp: true
mtu: 1500
portNumber: 0
nvconfig:
- device: '*'
parameters:
- PF_BAR2_ENABLE=0
- PER_PF_NUM_SF=1
- PF_TOTAL_SF=20
- PF_SF_BAR_SIZE=10
- NUM_PF_MSIX_VALID=0
- PF_NUM_PF_MSIX_VALID=1
- PF_NUM_PF_MSIX=228
- INTERNAL_CPU_MODEL=1
- INTERNAL_CPU_OFFLOAD_ENGINE=0
- SRIOV_EN=1
- NUM_OF_VFS=46
- LAG_RESOURCE_ALLOCATION=1
ovs:
rawConfigScript: |
_ovs-vsctl() {
ovs-vsctl --timeout 15 "$@"
}
# Remove default OVS configuration on the DPU and ensure no leftovers on the OVS kernel side
_ovs-vsctl --if-exists del-br ovsbr1
_ovs-vsctl --if-exists del-br ovsbr2
ovs-appctl --timeout 15 dpctl/del-dp system@ovs-system || true
_ovs-vsctl set Open_vSwitch . other_config:doca-init=true
_ovs-vsctl set Open_vSwitch . other_config:dpdk-max-memzones=50000
_ovs-vsctl set Open_vSwitch . other_config:hw-offload=true
_ovs-vsctl set Open_vSwitch . other_config:pmd-quiet-idle=true
_ovs-vsctl set Open_vSwitch . other_config:max-idle=20000
_ovs-vsctl set Open_vSwitch . other_config:max-revalidator=5000
_ovs-vsctl set Open_vSwitch . other_config:doca-congestion-threshold=60
_ovs-vsctl set Open_vSwitch . other_config:flow-limit=500000
_ovs-vsctl set Open_vSwitch . other_config:hw-offload-ct-unidir-udp-enabled=true
_ovs-vsctl remove Open_vSwitch . other_config default-datapath-type || true
if systemctl list-unit-files openvswitch-switch.service &>/dev/null; then
systemctl restart openvswitch-switch
elif systemctl list-unit-files openvswitch.service &>/dev/null; then
systemctl restart openvswitch
fi
_ovs-vsctl --may-exist add-br br-sfc
_ovs-vsctl set bridge br-sfc datapath_type=netdev
_ovs-vsctl set bridge br-sfc fail_mode=secure
_ovs-vsctl --if-exists del-br br-hbn
_ovs-vsctl --may-exist add-br br-hbn
_ovs-vsctl set bridge br-hbn datapath_type=netdev
_ovs-vsctl set bridge br-hbn fail_mode=secure
_ovs-vsctl --may-exist add-port br-sfc p0
_ovs-vsctl set Interface p0 type=dpdk
_ovs-vsctl set Interface p0 mtu_request=9216
_ovs-vsctl set Port p0 external_ids:dpf-type=physical
# Activate DOCA for OVNK
_ovs-vsctl set Open_vSwitch . external-ids:ovn-bridge-datapath-type=netdev
# setup ovnkube managed bridge, br-dpu (this corresponds to br-ex on ovnk docs)
_ovs-vsctl --may-exist add-br br-dpu
_ovs-vsctl br-set-external-id br-dpu bridge-id br-dpu
_ovs-vsctl br-set-external-id br-dpu bridge-uplink pbrdputobrovn
_ovs-vsctl set bridge br-dpu datapath_type=netdev
_ovs-vsctl --may-exist add-port br-dpu pf0hpf
_ovs-vsctl set Interface pf0hpf mtu_request=9216
_ovs-vsctl set Interface pf0hpf type=dpdk
# Create OVS bridge (br-ovn) in between the SC managed bridge and OVNK
_ovs-vsctl --may-exist add-br br-ovn
_ovs-vsctl set bridge br-ovn datapath_type=netdev
_ovs-vsctl --may-exist add-port br-ovn pbrovntobrdpu
_ovs-vsctl --may-exist add-port br-dpu pbrdputobrovn
# Patch br-ovn and br-dpu together
_ovs-vsctl set Interface pbrovntobrdpu type=patch options:peer=pbrdputobrovn
_ovs-vsctl set Interface pbrdputobrovn type=patch options:peer=pbrovntobrdpu
DPU Node Label Scripts
The DPU agent can run executable files on the DPU ARM and report their output as labels on the corresponding DPU cluster Node. This allows DPU-side hardware or software properties to be surfaced into Kubernetes scheduling decisions without any host-side tooling.
How it works
- On every provisioning run, the DPU agent scans the directory
/var/lib/dpf/dpuagent/node-label-scripts/on the DPU. - Each regular, executable file in that directory is run with a 30-second timeout. Shell scripts and compiled binaries are supported.
- Each non-empty line the file prints on stdout must have the form
<label-key-suffix>=<label-value>and produces one Node label. The key suffix is namespaced under thescripts.dpu.nvidia.com/prefix. A single file can therefore report any number of labels. - Labels from previous runs that are no longer emitted are automatically deleted from the Node.
Requirements:
- Every non-empty stdout line must contain a
=. Blank lines are ignored, and leading/trailing whitespace around the key and the value is trimmed. - The key suffix must be a valid Kubernetes label key suffix (alphanumeric,
-,_,.; max 63 characters). - The value must be a valid Kubernetes label value.
- The file name is not part of the label key and has no naming constraints.
- The file must be a regular file with at least one executable bit set (
chmod +x). - Directories in the scripts directory are silently ignored.
- If the same label key is emitted more than once, the last occurrence wins. Files run in sorted file name order, and within a file later lines win over earlier ones.
Deploying scripts via DPUFlavor
Use the configFiles field to place executable files into the default directory during provisioning:
apiVersion: provisioning.dpu.nvidia.com/v1alpha1
kind: DPUFlavor
metadata:
name: my-flavor
namespace: dpf-operator-system
spec:
configFiles:
- operation: override
path: /var/lib/dpf/dpuagent/node-label-scripts/network-info
permissions: "0700"
raw: |
#!/bin/bash
echo "test-label=some-data"
echo "link-speed=200G"
This example produces the Node labels scripts.dpu.nvidia.com/test-label=some-data and scripts.dpu.nvidia.com/link-speed=200G on the DPU cluster Node.
Note
Executables run during the dpuagent provisioning step on the DPU ARM, as root, with access to the DPU's hardware interfaces.
Best Practices
- Resource Planning: Always specify
dpuResourcesandsystemReservedResourcesto ensure proper resource allocation - Immutability: Plan your configuration carefully as DPUFlavor specs cannot be modified after creation
- Testing: Test DPUFlavor configurations in development environments before production deployment
- Documentation: Document custom configurations and their purposes for team understanding
- IB Mode Conversion: For DPUs initially in InfiniBand (IB) mode, always include
LINK_TYPE_P1=ETHin nvconfig parameters to convert to Ethernet mode. For dual port DPUs, also addLINK_TYPE_P2=ETH - NVConfig: Use wildcard (
device: '*') for uniform configuration across all devices. Use device-specific entries only when per-device configuration is required - DPU Credentials: Set
ubuntu_PASSWORDinbfcfgParametersto a strong hashed password. If omitted, the DPU is provisioned with defaultubuntu:ubuntucredentials.