Browse documentation
Load TPC data into Iceberg
Generate Parquet, import complete tables, and check a DuckDB read.
Use a Linux amd64 host, Docker, Python 3.10–3.12, and a free local port 9092. This single-node image is for disposable evaluation data. Install the published crowdb-tpc-loader package from PyPI.
Start and connect
docker run -d --name crowdb-iceberg \
-p 127.0.0.1:9092:9092 \
-v crowdb-iceberg-data:/opt/crowdb/data \
crowdb/crowdb-iceberg:latest
docker inspect --format '{{.State.Health.Status}}' crowdb-icebergWait for healthy, then get the generated credentials. Keep the file private:
umask 077
docker exec crowdb-iceberg crowdb-monitor credentials show --format env > ./crowdb-iceberg.env
set -a
. ./crowdb-iceberg.env
set +aInstall and load
Install crowdb-tpc-loader from PyPI. Source: github.com/buzzcrow/crowdb-tpc-loader. Its first TPC-H run can download the pinned tpchgen-cli binary. TPC-DS may fetch DuckDB’s tpcds extension. Supply these ahead of time and use --no-download for an offline run.
python3 -m venv .venv
. .venv/bin/activate
python -m pip install crowdb-tpc-loader
crowdb-tpc-loader load --benchmark tpch --sf 1 \
--namespace tpch_demo --report-file ./tpch-demo.jsonThe TPC-H load creates eight tables. To try TPC-DS separately, use a fresh namespace:
crowdb-tpc-loader load --benchmark tpcds --sf 1 \
--namespace tpcds_demo --upload-workers 4 --report-file ./tpcds-demo.jsonThe loader validates the whole dataset before creating tables. TPC-H and TPC-DS share the same flow: files calculate MD5 concurrently and stream through one S3 upload client with up to eight connections. Each request uses its table's catalog credentials and sends Content-MD5; SigV4 authentication stays enabled without file payload SHA256 prereads. --upload-workers N selects 1–24 scheduled table writes, with at most eight active S3 workers. Catalog configuration and namespace setup happen once per run. Recovery records are saved before remote actions; ready records can share one save. Local files are removed after all commits are verified unless --keep-files is set. A failed run retains staging; inspect the JSON report and recovery guide before retrying. Use SF=1 and a fresh namespace for performance tests.
Read from DuckDB
Run the DuckDB CLI with its iceberg and httpfs extensions. The empty warehouse selector is required for this CROWDB profile.
duckdb <<SQL
INSTALL iceberg;
LOAD iceberg;
INSTALL httpfs;
LOAD httpfs;
CREATE SECRET crowdb_catalog (TYPE ICEBERG, TOKEN '$ICEBERG_TOKEN');
ATTACH '' AS crowdb (TYPE ICEBERG, SECRET crowdb_catalog, ENDPOINT '$ICEBERG_URI');
SELECT r_regionkey, r_name FROM crowdb.tpch_demo.region WHERE r_regionkey = 1;
SQLThe query returned (1, AMERICA).
TPC-H and TPC-DS development checks
Open DuckDB again, repeat the extension and ATTACH statements above, then run TPC-H Q1:
SELECT l_returnflag, l_linestatus,
sum(l_quantity) AS sum_qty,
sum(l_extendedprice) AS sum_base_price,
sum(l_extendedprice * (1 - l_discount)) AS sum_disc_price,
sum(l_extendedprice * (1 - l_discount) * (1 + l_tax)) AS sum_charge,
avg(l_quantity) AS avg_qty,
avg(l_extendedprice) AS avg_price,
avg(l_discount) AS avg_disc,
count(*) AS count_order
FROM crowdb.tpch_demo.lineitem
WHERE l_shipdate <= DATE '1998-09-02'
GROUP BY l_returnflag, l_linestatus
ORDER BY l_returnflag, l_linestatus;At SF 0.01, Q1 returned four groups with count_order values 14,876, 348, 29,181, and 14,902. All 22 DuckDB TPC-H queries against the eight container tables matched DuckDB reading the same local Parquet files. A separate TPC-DS load created 24 tables with 277,976 rows using --upload-workers 4; all 99 DuckDB TPC-DS queries matched the same local Parquet data. These are development checks of the container read path, not timed benchmark results or TPC scores.
To inspect all imported table metadata and sample scans, see the loader repository’s verification commands. The published image’s latest tag may move. Keep its digest with any results.
Finish
docker rm -f crowdb-iceberg
rm -f ./crowdb-iceberg.envThe named volume remains. Delete it only when you intend to discard the imported data.