Browse documentation
CROWDB / GUIDES

Load TPC data into Iceberg

Generate Parquet, import complete tables, and check a DuckDB read.

LATEST IMAGE · DEVELOPMENT PREVIEW

Use a Linux amd64 host, Docker, Python 3.10–3.12, and a free local port 9092. This single-node image is for disposable evaluation data. Install the published crowdb-tpc-loader package from PyPI.

Start and connect

sh
docker run -d --name crowdb-iceberg \
  -p 127.0.0.1:9092:9092 \
  -v crowdb-iceberg-data:/opt/crowdb/data \
  crowdb/crowdb-iceberg:latest
docker inspect --format '{{.State.Health.Status}}' crowdb-iceberg

Wait for healthy, then get the generated credentials. Keep the file private:

sh
umask 077
docker exec crowdb-iceberg crowdb-monitor credentials show --format env > ./crowdb-iceberg.env
set -a
. ./crowdb-iceberg.env
set +a

Install and load

Install crowdb-tpc-loader from PyPI. Source: github.com/buzzcrow/crowdb-tpc-loader. Its first TPC-H run can download the pinned tpchgen-cli binary. TPC-DS may fetch DuckDB’s tpcds extension. Supply these ahead of time and use --no-download for an offline run.

sh
python3 -m venv .venv
. .venv/bin/activate
python -m pip install crowdb-tpc-loader
crowdb-tpc-loader load --benchmark tpch --sf 1 \
  --namespace tpch_demo --report-file ./tpch-demo.json

The TPC-H load creates eight tables. To try TPC-DS separately, use a fresh namespace:

sh
crowdb-tpc-loader load --benchmark tpcds --sf 1 \
  --namespace tpcds_demo --upload-workers 4 --report-file ./tpcds-demo.json

The loader validates the whole dataset before creating tables. TPC-H and TPC-DS share the same flow: files calculate MD5 concurrently and stream through one S3 upload client with up to eight connections. Each request uses its table's catalog credentials and sends Content-MD5; SigV4 authentication stays enabled without file payload SHA256 prereads. --upload-workers N selects 1–24 scheduled table writes, with at most eight active S3 workers. Catalog configuration and namespace setup happen once per run. Recovery records are saved before remote actions; ready records can share one save. Local files are removed after all commits are verified unless --keep-files is set. A failed run retains staging; inspect the JSON report and recovery guide before retrying. Use SF=1 and a fresh namespace for performance tests.

Read from DuckDB

Run the DuckDB CLI with its iceberg and httpfs extensions. The empty warehouse selector is required for this CROWDB profile.

sh
duckdb <<SQL
INSTALL iceberg;
LOAD iceberg;
INSTALL httpfs;
LOAD httpfs;
CREATE SECRET crowdb_catalog (TYPE ICEBERG, TOKEN '$ICEBERG_TOKEN');
ATTACH '' AS crowdb (TYPE ICEBERG, SECRET crowdb_catalog, ENDPOINT '$ICEBERG_URI');
SELECT r_regionkey, r_name FROM crowdb.tpch_demo.region WHERE r_regionkey = 1;
SQL

The query returned (1, AMERICA).

TPC-H and TPC-DS development checks

Open DuckDB again, repeat the extension and ATTACH statements above, then run TPC-H Q1:

sql
SELECT l_returnflag, l_linestatus,
       sum(l_quantity) AS sum_qty,
       sum(l_extendedprice) AS sum_base_price,
       sum(l_extendedprice * (1 - l_discount)) AS sum_disc_price,
       sum(l_extendedprice * (1 - l_discount) * (1 + l_tax)) AS sum_charge,
       avg(l_quantity) AS avg_qty,
       avg(l_extendedprice) AS avg_price,
       avg(l_discount) AS avg_disc,
       count(*) AS count_order
FROM crowdb.tpch_demo.lineitem
WHERE l_shipdate <= DATE '1998-09-02'
GROUP BY l_returnflag, l_linestatus
ORDER BY l_returnflag, l_linestatus;

At SF 0.01, Q1 returned four groups with count_order values 14,876, 348, 29,181, and 14,902. All 22 DuckDB TPC-H queries against the eight container tables matched DuckDB reading the same local Parquet files. A separate TPC-DS load created 24 tables with 277,976 rows using --upload-workers 4; all 99 DuckDB TPC-DS queries matched the same local Parquet data. These are development checks of the container read path, not timed benchmark results or TPC scores.

To inspect all imported table metadata and sample scans, see the loader repository’s verification commands. The published image’s latest tag may move. Keep its digest with any results.

Finish

sh
docker rm -f crowdb-iceberg
rm -f ./crowdb-iceberg.env

The named volume remains. Delete it only when you intend to discard the imported data.