Apache Spark K8s Operator
July 17, 2026 · View on GitHub
Apache Spark™ K8s Operator is a subproject of Apache Spark and aims to extend K8s resource manager to manage Apache Spark applications via Operator Pattern.
Install Helm Chart
Apache Spark provides a Helm Chart.
- https://apache.github.io/spark-kubernetes-operator/
- https://artifacthub.io/packages/helm/spark-kubernetes-operator/spark-kubernetes-operator/
helm repo add spark https://apache.github.io/spark-kubernetes-operator
helm repo update
helm install spark spark/spark-kubernetes-operator
Building Spark K8s Operator
Spark K8s Operator is built using Gradle. To build, run:
./gradlew build -x test
Running Tests
./gradlew build
Updating Dependency Verification Metadata
When upgrading dependencies, regenerate gradle/verification-metadata.xml from scratch.
Since --write-verification-metadata is append-only and keeps stale entries of removed or
old dependencies, delete the existing file first so that the regenerated file contains
only the checksums of the current dependencies. Pass --refresh-dependencies so that
metadata files (BOM / parent POMs) already present in the local Gradle cache are
downloaded and recorded again — without it their checksums are silently dropped and
CI fails with a cold cache:
rm gradle/verification-metadata.xml
./gradlew --write-verification-metadata sha512 build --refresh-dependencies
Note that Gradle rewrites the file without the ASF license header, so restore the header after regeneration.
Build Docker Image
./gradlew buildDockerImage
Install Helm Chart from the source code
helm install spark -f build-tools/helm/spark-kubernetes-operator/values.yaml build-tools/helm/spark-kubernetes-operator/
Run Spark Pi App
$ kubectl apply -f examples/pi.yaml
$ kubectl get sparkapp
NAME CURRENT STATE AGE
pi ResourceReleased 4m10s
$ kubectl delete sparkapp/pi
Run Spark Cluster
$ kubectl apply -f examples/prod-cluster-with-three-workers.yaml
$ kubectl get sparkcluster
NAME CURRENT STATE AGE
prod RunningHealthy 10s
$ kubectl port-forward prod-master-0 6066 &
$ ./examples/submit-pi-to-prod.sh
{
"action" : "CreateSubmissionResponse",
"message" : "Driver successfully submitted as driver-20260110030233-0000",
"serverSparkVersion" : "4.2.0",
"submissionId" : "driver-20260110030233-0000",
"success" : true
}
$ curl http://localhost:6066/v1/submissions/status/driver-20260110030233-0000/
{
"action" : "SubmissionStatusResponse",
"driverState" : "FINISHED",
"serverSparkVersion" : "4.2.0",
"submissionId" : "driver-20260110030233-0000",
"success" : true,
"workerHostPort" : "10.1.1.172:44233",
"workerId" : "worker-20260110030145-10.1.1.172-44233"
}
$ kubectl delete sparkcluster prod
sparkcluster.spark.apache.org "prod" deleted
Run Spark Pi App on Apache YuniKorn scheduler
If you have not yet done so, follow YuniKorn docs to install the latest version:
helm repo add yunikorn https://apache.github.io/yunikorn-release
helm repo update
helm install yunikorn yunikorn/yunikorn --namespace yunikorn --version 1.8.0 --create-namespace --set embedAdmissionController=false
Submit a Spark app to YuniKorn enabled cluster:
$ kubectl apply -f examples/pi-on-yunikorn.yaml
$ kubectl describe pod pi-on-yunikorn-0-driver
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Scheduling 1s yunikorn default/pi-on-yunikorn-0-driver is queued and waiting for allocation
Normal Scheduled 1s yunikorn Successfully assigned default/pi-on-yunikorn-0-driver to node docker-desktop
Normal PodBindSuccessful 1s yunikorn Pod default/pi-on-yunikorn-0-driver is successfully bound to node docker-desktop
Normal Pulled 0s kubelet Container image "apache/spark:4.2.0-scala" already present on machine
Normal Created 0s kubelet Created container: spark-kubernetes-driver
Normal Started 0s kubelet Started container spark-kubernetes-driver
$ kubectl delete sparkapp pi-on-yunikorn
sparkapplication.spark.apache.org "pi-on-yunikorn" deleted from default namespace
Clean Up
Check the existing Spark applications and clusters. If exists, delete them.
$ kubectl get sparkapp
No resources found in default namespace.
$ kubectl get sparkcluster
No resources found in default namespace.
Remove HelmChart and CRDs.
helm uninstall spark
kubectl delete crd sparkapplications.spark.apache.org
kubectl delete crd sparkclusters.spark.apache.org
Contributing
Please review the Contribution to Spark guide for information on how to get started contributing to the project.