RAPIDS Accelerator for Apache Spark Testing
September 11, 2026 ยท View on GitHub
We have a stand-alone example that you can run in the integration tests.
The example is based off of the mortgage dataset you can download
here
and the code is in the com.nvidia.spark.rapids.tests.mortgage package.
Unit Tests
Unit tests implemented using the ScalaTest framework reside in the tests directory. This is unconventional and is done so we can run the tests on the close-to-final shaded single-shim version of the plugin. It also helps with how we collect code coverage.
The tests module depends on the aggregator module which shades external dependencies and
aggregates them along with internal submodules into an artifact supporting a single Spark version.
The minimum required Maven phase to run unit tests is package. Alternatively, you may run
mvn install and use mvn test for subsequent testing. However, to avoid dealing with stale jars
in the local Maven repo cache, we recommend to invoke mvn package -pl tests -am ... from the
spark-rapids root directory. Add -f scala2.13 if you want to run unit tests against
Apache Spark dependencies based on Scala 2.13.
To run targeted Scala tests use
-DwildcardSuites=<comma separated list of packages or fully-qualified test suites>
Or easier, use a combination of
-Dsuffixes=<comma separated list of suffix regexes> to restrict the test suites being run,
which corresponds to -q option in the
ScalaTest runner.
and
-Dtests=<comma separated list of keywords or test names>, to restrict tests run within test suites,
which corresponds to -z or -t options in the
ScalaTest runner.
For more information about using scalatest with Maven please refer to the scalatest documentation and the the source code.
Running Unit Tests Against Specific Apache Spark Versions
You can run the unit tests against different versions of Spark using the different profiles. The default version runs against Spark 3.3.0, to run against a specific version use a buildver property:
-Dbuildver=330(Spark 3.3.0)-Dbuildver=350(Spark 3.5.0)
etc
Please refer to the tests project POM to see the list of test profiles supported.
Apache Spark specific configurations can be passed in by setting the SPARK_CONF environment
variable.
Examples:
To run all tests against Apache Spark 3.3.0,
mvn package -pl tests -am -Dbuildver=330
To pass Apache Spark configs --conf spark.dynamicAllocation.enabled=false --conf spark.task.cpus=1
do something like.
SPARK_CONF="spark.dynamicAllocation.enabled=false,spark.task.cpus=1" mvn ...
To run all tests in ParquetWriterSuite in package com.nvidia.spark.rapids, issue
mvn package -pl tests -am -DwildcardSuites="com.nvidia.spark.rapids.ParquetWriterSuite"
To run all AnsiCastOpSuite and CastOpSuite tests dealing with decimals using Apache Spark 3.3.0 on Scala 2.13 artifacts, issue:
mvn package -f scala2.13 -pl tests -am -Dbuildver=330 -Dsuffixes='.*CastOpSuite' -Dtests=decimal
Parallel Unit Tests
Premerge runs the Scala unit tests in parallel, using
ParallelUnitTestRunner to run suites in separate worker JVMs. Add [serial ut] or [serial-ut]
to the PR title to run them serially when debugging a concurrency-only failure, reading a linear
ScalaTest log, or verifying a fix. Local runs are serial unless parallel execution is enabled
explicitly:
mvn package -pl tests -am -Drapids.parallelUnitTests=true -DparallelForkCount=4
- Parallel unit tests currently support at most four concurrent worker JVMs. Setting
parallelForkCounthigher than four does not increase concurrency. Further UT and IT parallelism tuning is tracked in #15344. -Dsuffixesand-Dtestsare not supported; the runner fails fast. Use-DwildcardSuites, which matches fully qualified suite-name prefixes.- Before starting workers, the runner detects free GPU memory and reserves 1 GiB for headroom.
It budgets 4 GiB per worker and uses the smallest of the resulting memory limit,
parallelForkCount, four workers, and the number of suite batches. For example, 9 GiB free permits two workers and 17 GiB permits four. One worker still runs all selected suites sequentially; less than 5 GiB free fails before any worker starts. - The GPU is shared. Each worker gets
rapids.test.gpu.allocFraction * 0.8 / workerCount, using the actual worker count. The default minimum pool fraction is also scaled by the startup free-to-total GPU memory ratio so memory already occupied by other processes does not inflate the minimum. The 4 GiB budget controls scheduling; suites that explicitly configure their own pools retain those settings. - Each suite has a watchdog controlled by
-DparallelSuiteTimeout, which defaults to 1800 seconds. On timeout, the runner captures ajstack, kills the worker, and fails the run. - The
RapidsDynamicPartitionPruningV1SuiteAEOffandRapidsDynamicPartitionPruningV1SuiteAEOnsuites are assigned to the same worker so they execute serially with each other because concurrent execution previously caused GPU broadcast contention. The list isDPP_SUITESinParallelUnitTestRunner.scala; the root cause remains tracked in #15401. - To debug a parallel failure, follow the
[wave-<run>-worker-<id>]log prefix for the failing suite. Re-run that suite alone with-DwildcardSuites=<fully.qualified.Suite>, both with and without-Drapids.parallelUnitTests=true, to determine whether concurrency caused the failure.
Integration Tests
Please refer to the integration-tests README