Development
September 11, 2026 ยท View on GitHub
- Building
- Running
- Testing Prometheus exporter
- Testing device info exporter
- Building container image
- Testing container image
- Extracting sources from container image
- Testing in single-node cluster
- Driver stack updates
Building
See the build requirements in Level-Zero Go Bindings README.
Clone the repository:
git clone https://github.com/intel/xpumanager
Switch to it:
cd xpumanager/xpumd
And build the daemon:
make build
Running
sudo ./dist/xpumd --config config-example.yaml
Note
Extra privileges are required to get all the metrics, but some metrics
are available also without sudo, see Metrics.
Note
If not running this through normal user session, one may need to specify
intel_xpu_info exporter socket directory with e.g. XDG_RUNTIME_DIR=$PWD
environment variable (or full socket path with the
exporters.intel_xpu_info.endpoint config option).
Testing Prometheus exporter
In another terminal:
curl --no-progress-meter http://localhost:8080/metrics
Note
Example config restricts access to localhost for security reasons.
If this is run on another host, xpumd needs to be started with
--set exporters.prometheus.endpoint=0.0.0.0:8080 option.
Testing device info exporter
In another terminal, run the test client to receive device health information and changes:
sudo ./dist/xpuinfo-cli
Its output should look something like this:
info:
uuid: 8680457d-0800-0000-0002-000000000000
model: Intel(R) Graphics
pci: null
health:
- name: memory
status: 0
...
Building container image
docker build -t registry.local/xpumd:latest .
By default the image gets the Level-Zero GPU backend from the released
libze-intel-gpu1 package. It can also be built from the
compute-runtime sources, pinned to
the revision that the experimental extensions in the Go bindings are generated from:
docker build --build-arg BACKEND=src -t registry.local/xpumd:latest .
That is how the published images are built, as no released package implements the experimental extensions yet. It takes considerably longer, so the local default is the released package.
Either way, the package info and the license notice that the image carries describe the backend it actually ships.
Testing container image
Test the container with example config from the image:
docker run -it --rm --user 0 --cap-drop ALL --cap-add PERFMON \
--device /dev/dri --publish 8080:8080 registry.local/xpumd:latest \
--config /etc/xpumd/config-example.yaml
One can also modify the config, e.g. to drop the local gRPC health endpoint:
sed -i s/intel_xpu_info,// config-example.yaml
Map the modified config inside container, and ask daemon to use it:
docker run -it --rm --user 0 --cap-drop ALL --cap-add PERFMON \
--volume $PWD/config-example.yaml:/etc/xpumd/config.yaml:ro \
--device /dev/dri --publish 8080:8080 registry.local/xpumd:latest \
--config /etc/xpumd/config.yaml
Query Prometheus metrics:
curl --no-progress-meter http://localhost:8080/metrics
SELinux and Podman
If running on a distro with SELinux enabled and docker being
provided by podman i.e. container being run with normal user
privileges, add --security-opt label=disable Docker option,
otherwise contents of host volume mounts are inaccessible.
Testing container image with stub driver
The image also includes a stub Level-Zero loader library which can be used to test the daemon without a real GPU or driver.
Use LD_LIBRARY_PATH to load the stub driver and SYSMAN_STUB_CONFIG to
specify the config file:
docker run -it --rm --user 0 \
-e LD_LIBRARY_PATH=/usr/local/lib/xpumd/level-zero-stub \
-e SYSMAN_STUB_CONFIG=/etc/xpumd/level-zero-stub/example-config.yaml \
--publish 8080:8080 registry.local/xpumd:latest \
--config /etc/xpumd/config-example.yaml
To supply test-specific stub state, bind-mount a YAML file into the container
and point SYSMAN_STUB_CONFIG at that path:
docker run -it --rm --user 0 \
-e LD_LIBRARY_PATH=/usr/local/lib/xpumd/level-zero-stub \
-e SYSMAN_STUB_CONFIG=/etc/xpumd/level-zero-stub/example-config.yaml \
--volume $PWD/level-zero-go/level-zero-stub:/etc/xpumd/level-zero-stub:ro \
--publish 8080:8080 registry.local/xpumd:latest \
--config /etc/xpumd/config-example.yaml
Note
When using the stub loader, no /dev/dri device mapping or GPU-specific
Linux capabilities are required.
Extracting sources from container image
For (L)GPL compliance, the container image includes source packages in the
/sources directory for (L)GPL-licensed packages that were added on top of the
Ubuntu base image.
To extract these sources from the container image (locally built one is used as an example here):
- Run image in a temporary container:
docker create --name xpumd-temp registry.local/xpumd:latest
- Extract sources from it:
docker cp xpumd-temp:/sources ./sources
- And remove the container:
docker rm xpumd-temp
- List the source packages:
ls -lh ./sources
Testing in single-node cluster
After building the container image, load the image onto the cluster.
Load container image
Kind cluster
kind load docker-image registry.local/xpumd:latest
Containerd-based cluster
- Save image as a tarball (on build machine):
docker save registry.local/xpumd:latest -o xpumd-latest.tar
- Import tarball to container runtime (on cluster node):
sudo ctr -n k8s.io images import xpumd-latest.tar
(-n k8s.io option is needed for images to be visible to Kubernetes / crictl.)
Deploy with Helm
The most straightforward way to deploy the daemon for testing and debugging is to use privileged mode, which allows the daemon to access all the necessary devices without Kubernetes resource drivers:
helm install xpumd charts/xpumd --set image.repository=registry.local/xpumd --set image.pullPolicy=Never \
--set gpuAccess=none \
--set securityContextOverride.runAsUser=0 --set securityContextOverride.privileged=true
See Helm chart README for other deployment scenarios and detailed configuration options.
Integration tests
The project includes a suite of Kubernetes integration tests that verify the deployment and basic functionality using the stub Level-Zero driver.
Kind cluster
Run with:
make test-integration-kind
This runs everything locally: builds the container image, creates a disposable Kind cluster, deploys xpumd via the Helm chart (with stub driver enabled), and runs the tests.
Existing cluster
To run the integration tests against an existing cluster:
make test-integration-existing-cluster IMAGE_REPOSITORY=ghcr.io/intel/xpumanager/xpumd IMAGE_TAG=latest
Important
Ensure that the container image is present in/reachable by the cluster.
Driver stack updates
The build args of the BACKEND=src Dockerfile stage are copied from the
Level-Zero Go example Dockerfile by
make generate-dockerfile-args (part of make generate), so update them there.
XPUMD helper scripts and Docker files validate directly downloaded (driver) DEB and ZIP files against hard-coded checksums. If DEB / ZIP file versions are changed, build will error out after outputting the new checksums (so they can be updated after validation).