Browse documentation
Three layers. One system.
A dependency map from access semantics to protected storage.
The system in one picture
CROWDB keeps access semantics separate from the mechanisms that store and protect bytes. S3 and Iceberg have working implementations. Dataset is a planned model for samples, shards, tensors and batches.
Chunk Stream · Chunk-KV
Placement, bounded streaming, protection and repair
crowdb-kv · Multi-Paxos · WAL · crowdb-tree · RPC
Access: objects and tables stay distinct
The Access layer exposes HTTP interfaces for S3 objects and Iceberg tables. Iceberg includes its own native catalog and FileIO path; it does not require a separate S3 deployment for those files.
Shared infrastructure does not erase the data model. Object operations and table commits have different meanings. Uploading an object through the S3 endpoint does not register a table in the Iceberg catalog.
The planned Dataset model includes HTTP and topology-aware native access. The latter is intended to avoid making an HTTP server the permanent intermediary. That native path is not part of the development preview.
Chunk: one place for storage work
The Chunk layer owns placement, streaming, protection and repair. Chunk Stream supplies durable ordered append. Chunk-KV uses range partitions that can split and rebalance while operations continue.
The repository describes bounded streaming, mirrored small data, strip-level erasure coding and shard repair in the shared chunk foundation. These are implementation descriptions, not a production durability claim or a benchmark result. Specific profiles can disable capabilities; physical reclamation is disabled in the single-node container.
Reusable KV: distributed state underneath
crowdb-kv runs independent Multi-Paxos slots concurrently. It supplies replicated state, write-ahead logging and lease reads with pluggable engines. The storage tree and RPC layer are reusable parts of this foundation.
The KV layer is not the product-level data model. The diagram shows architectural dependencies; it does not say that every object or table payload is routed through KV as another sequential hop.
Direct GPU delivery is a design direction
The design leaves room for a path from DiskIO buffers through RDMA or GPUDirect Storage to GPU memory. That requires control over topology, buffer ownership and data lifetimes, not just another API.
Dataset and direct GPU delivery remain planned. There is no live GPU path, measured speedup, or production-readiness claim attached to this illustration.
Read the source contracts
This page summarizes access models and the storage core. The design documents remain in the CROWDB repository; for setup steps, use the user manual.