Working with S3

August 6, 2026 · View on GitHub

Object stores offered by CSPs such as AWS S3 are important for users of Gluten to store their data. This doc will discuss all details of configs, and use cases around using Gluten with object stores. In order to use an S3 endpoint as your data source, please ensure you are using the following S3 configs in your spark-defaults.conf. If you're experiencing any issues authenticating to S3 with additional auth mechanisms, please reach out to us using the 'Issues' tab.

Working with S3

Configuring S3 endpoint

S3 provides the endpoint based method to access the files, here's the example configuration. Users may need to modify some values based on real setup.

spark.hadoop.fs.s3a.impl                        org.apache.hadoop.fs.s3a.S3AFileSystem
spark.hadoop.fs.s3a.aws.credentials.provider    org.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider
spark.hadoop.fs.s3a.access.key                  XXXXXXXXX
spark.hadoop.fs.s3a.secret.key                  XXXXXXXXX
spark.hadoop.fs.s3a.endpoint                    https://s3.us-west-1.amazonaws.com
spark.hadoop.fs.s3a.connection.ssl.enabled      true
spark.hadoop.fs.s3a.path.style.access           false

Configuring S3 instance credentials

S3 also provides other methods for accessing, you can also use instance credentials by setting the following config

spark.hadoop.fs.s3a.use.instance.credentials true

Note that in this case, "spark.hadoop.fs.s3a.endpoint" won't take affect as Gluten will use the endpoint set during instance creation.

Configuring S3 IAM roles

You can also use iam role credentials by setting the following configurations. Instance credentials have higher priority than iam credentials.

spark.hadoop.fs.s3a.iam.role  xxxx
spark.hadoop.fs.s3a.iam.role.session.name xxxx

Note that spark.hadoop.fs.s3a.iam.role.session.name is optional.

Other authentatication methods are not supported yet

Log granularity of AWS C++ SDK in velox

You can change log granularity of AWS C++ SDK by setting the spark.gluten.velox.awsSdkLogLevel configuration. The Allowed values are: "OFF", "FATAL", "ERROR", "WARN", "INFO", "DEBUG", "TRACE".

Configuring Whether To Use Proxy From Env for S3 C++ Client

You can change whether to use proxy from env for S3 C++ client by setting the spark.gluten.velox.s3UseProxyFromEnv configuration. The Allowed values are: "false", "true".

Configuring S3 Payload Signing Policy

You can change the S3 payload signing policy by setting the spark.gluten.velox.s3PayloadSigningPolicy configuration. The Allowed values are: "Always", "RequestDependent", "Never".

  • When set to "Always", the payload checksum is included in the signature calculation.
  • When set to "RequestDependent", the payload checksum is included based on the value returned by "AmazonWebServiceRequest::SignBody()".

Configuring S3 Log Location

You can set the log location by setting the spark.gluten.velox.s3LogLocation configuration.

Configuring Async S3 Multipart Upload

You can enable asynchronous multipart part upload by setting spark.gluten.velox.s3UploadPartAsync to true. Use spark.gluten.velox.s3MaxConcurrentUploadNum to control the maximum number of in-flight part uploads per file, and spark.gluten.velox.s3UploadThreads to control the shared upload thread pool size. These settings apply to all buckets by default. To override a single bucket, use Velox's bucket-specific S3 keys through Gluten's static backend pass-through prefix, for example spark.gluten.velox.hive.s3.bucket.my-bucket.part-upload-async.

Known Limitation: Cross-Region S3 Access

If your S3 data and your configured endpoint are in different AWS regions — for example, your bucket is in us-west-2 but you have set the endpoint to us-east-1 (or left it unconfigured, causing requests to default to us-east-1) — you will encounter a runtime error.

Unlike the Java S3 SDK used by s3a:// (which often handles 301 cross-region redirects automatically), the AWS C++ SDK (aws-sdk-cpp) does not automatically follow 301 PermanentRedirect responses for S3 payload requests. It treats the redirect as a non-retriable error:

Error Type: SDK Error (100)
Error: PermanentRedirect — The bucket you are attempting to access must be addressed using the specified endpoint.

Resolution: Set the endpoint explicitly to the region where your bucket resides, for example:

spark.hadoop.fs.s3a.endpoint=s3.us-west-2.amazonaws.com

Local Caching support

Velox supports a local cache when reading data from S3 but not strictly tested and there are several limitations. Please refer Velox Local Cache part for more detailed configurations.

Configurations:

All configurations starts with spark.hadoop.fs.s3a.

✅ Supported ❌ Not Supported ⚠️ Partial Support 🔄 In Progress 🚫 Not applied or transparent to Gluten

Here is the list of hadoop s3 file system configurations:

NameDefault ValueGluten Honored
aws.credentials.provider(empty)⚠️
security.credential.provider.path(empty)
assumed.role.arn(empty)
assumed.role.session.name(empty)
assumed.role.policy(empty)
assumed.role.session.duration30m
assumed.role.sts.endpoint(empty)
assumed.role.sts.endpoint.region(empty)
assumed.role.credentials.providerorg.apache.hadoop.fs.s3a.SimpleAWSCredentialsProvider
delegation.token.binding(empty)
attempts.maximum5
socket.send.buffer8192
socket.recv.buffer8192
paging.maximum5000
multipart.size64M
multipart.threshold128M
multiobjectdelete.enabletrue
acl.default(empty)
multipart.purgefalse
multipart.purge.age86400
encryption.algorithm(empty)
encryption.key(empty)
signing-algorithm(empty)
block.size32M
buffer.dir{env.LOCAL_DIRS:-{hadoop.tmp.dir}}/s3a
fast.upload.bufferdisk
fast.upload.active.blocks4
readahead.range64K
user.agent.prefix(empty)
implorg.apache.hadoop.fs.s3a.S3AFileSystem
retry.limit7
retry.interval500ms
retry.throttle.limit20
retry.throttle.interval100ms
committer.namefile🚫
committer.magic.enabledtrue🚫
committer.threads8🚫
committer.staging.tmp.pathtmp/staging🚫
committer.staging.unique-filenamestrue🚫
committer.staging.conflict-modeappend🚫
committer.abort.pending.uploadstrue🚫
list.version2🚫
etag.checksum.enabledfalse
change.detection.sourceetag
change.detection.modeserver
change.detection.version.requiredtrue
ssl.channel.modedefault_jsse
downgrade.syncable.exceptionstrue
create.checksum.algorithm(empty)
audit.enabledtrue
vectored.read.min.seek.size128K
vectored.read.max.merged.size2M
vectored.active.ranged.reads4
experimental.input.fadviserandom
threads.max96
threads.keepalivetime60s
executor.capacity16
max.total.tasks16
connection.maximum25
connection.keepalivefalse
connection.acquisition.timeout60s
connection.establish.timeout30s
connection.idle.time60s
connection.request.timeout60s
connection.timeout200s
connection.ttl5m

Gluten new parameters:

NameDefault Value
access.key(none)
secret.key(none)
endpoint(none)
connection.ssl.enabledfalse
path.style.accessfalse
retry.limit(none)
retry.modelegacy
instance.credentialsfalse
iam.role(none)
iam.role.session.namegluten-session
endpoint.region(none)
aws.imds.enabledtrue

Gluten configures:

NameDefault Value
spark.gluten.velox.awsSdkLogLevelFATAL
spark.gluten.velox.s3UseProxyFromEnvfalse
spark.gluten.velox.s3PayloadSigningPolicyNever
spark.gluten.velox.s3LogLocation(none)
spark.gluten.velox.s3UploadPartAsyncfalse
spark.gluten.velox.s3MaxConcurrentUploadNum4
spark.gluten.velox.s3UploadThreads16