duckdb_hudi

June 19, 2026 ยท View on GitHub

A native, zero-JVM DuckDB extension written in Rust that allows you to scan Apache Hudi tables directly into DuckDB's high-performance vectorized storage engine.

By leveraging hudi-rs and duckdb-rs under the hood, this extension bypasses traditional Spark/JVM runtime dependencies entirely, making it ideal for fast, local analytical queries or embedded data lakehouse pipelines.

Features

  • Zero JVM/Spark Overhead: Pure Rust data loading straight into DuckDB memory.
  • No Heavy Builds: Compiles directly via Cargo without requiring local DuckDB C++ source builds.
  • Automated Tooling: Complete integration with DuckDB's standard extension CI/CD toolchain.

Cloning

Clone the repository along with its required submodules:

git clone --recurse-submodules https://github.com/oglego/duckdb-hudi.git

Dependencies

In principle, these extensions can be compiled with the Rust toolchain alone. However, this template relies on some additional tooling to make life a little easier and to be able to share CI/CD infrastructure with extension templates for other languages:

  • Python3
  • Python3-venv
  • Make
  • Git

Installing these dependencies will vary per platform:

  • For Linux, these come generally pre-installed or are available through the distro-specific package manager.
  • For MacOS, homebrew.
  • For Windows, chocolatey.

Building

After installing the dependencies, building is a two-step process. Firstly run:

make configure

This will ensure a Python venv is set up with DuckDB and DuckDB's test runner installed. Additionally, depending on configuration, DuckDB will be used to determine the correct platform for which you are compiling.

Then, to build the extension run:

make debug

This delegates the build process to cargo, which will produce a shared library in target/debug/<shared_lib_name>. After this step, a script is run to transform the shared library into a loadable extension by appending a binary footer. The resulting extension is written to the build/debug directory.

To create optimized release binaries, simply run make release instead.

Running the extension

To run the extension code, start duckdb with -unsigned flag. This will allow you to load the local extension file.

duckdb -unsigned

After loading the extension by the file path, you can use the functions provided by the extension (in this case, rusty_quack()).

LOAD './build/debug/extension/duckdb_hudi/duckdb_hudi.duckdb_extension';
SELECT * FROM hudi_scan();

Testing

This extension uses the DuckDB Python client for testing. This should be automatically installed in the make configure step. The tests themselves are written in the SQLLogicTest format, just like most of DuckDB's tests. A sample test can be found in test/sql/<extension_name>.test. To run the tests using the debug build:

make test_debug

or for the release build:

make test_release

Version switching

Testing with different DuckDB versions is really simple:

First, run

make clean_all

to ensure the previous make configure step is deleted.

Then, run

DUCKDB_TEST_VERSION=v1.3.2 make configure

to select a different duckdb version to test with

Finally, build and test with

make debug
make test_debug

Known issues

This is a bit of a footgun, but the extensions produced by this template may (or may not) be broken on windows on python3.11 with the following error on extension load:

IO Error: Extension '<name>.duckdb_extension' could not be loaded: The specified module could not be found

This was resolved by using python 3.12