Browse documentation
CROWDB / ARCHITECTURE

Three layers. One system.

A dependency map from access semantics to protected storage.

LATEST IMAGE · DEVELOPMENT PREVIEW

The system in one picture

CROWDB keeps access semantics separate from the mechanisms that store and protect bytes. S3 and Iceberg have working implementations. Dataset is a planned model for samples, shards, tensors and batches.

FIG. 01 / THREE LAYERSimplementedplanned
S3HTTP objectsIMPLEMENTED
IcebergCatalog + FileIOIMPLEMENTED
DatasetHTTP + native clientIN DESIGN
02Chunk
One protected storage layer

Chunk Stream · Chunk-KV
Placement, bounded streaming, protection and repair

chunk client → Chunk I/O → ChunkDB → DiskIO → DiskDB
01Reusable KV
Replicated metadata and durable state

crowdb-kv · Multi-Paxos · WAL · crowdb-tree · RPC

PLANNED PATH   DiskIO buffers → RDMA / GDS → GPU memory. This is a design direction, not a capability of the preview.
The KV layer supplies distributed state; the diagram does not imply that every payload is stored in KV. S3 and Iceberg use the shared core without translating through one another. Physical reclamation is disabled in the current single-node container profile.

Access: objects and tables stay distinct

The Access layer exposes HTTP interfaces for S3 objects and Iceberg tables. Iceberg includes its own native catalog and FileIO path; it does not require a separate S3 deployment for those files.

Shared infrastructure does not erase the data model. Object operations and table commits have different meanings. Uploading an object through the S3 endpoint does not register a table in the Iceberg catalog.

The planned Dataset model includes HTTP and topology-aware native access. The latter is intended to avoid making an HTTP server the permanent intermediary. That native path is not part of the development preview.

Chunk: one place for storage work

The Chunk layer owns placement, streaming, protection and repair. Chunk Stream supplies durable ordered append. Chunk-KV uses range partitions that can split and rebalance while operations continue.

The repository describes bounded streaming, mirrored small data, strip-level erasure coding and shard repair in the shared chunk foundation. These are implementation descriptions, not a production durability claim or a benchmark result. Specific profiles can disable capabilities; physical reclamation is disabled in the single-node container.

Reusable KV: distributed state underneath

crowdb-kv runs independent Multi-Paxos slots concurrently. It supplies replicated state, write-ahead logging and lease reads with pluggable engines. The storage tree and RPC layer are reusable parts of this foundation.

The KV layer is not the product-level data model. The diagram shows architectural dependencies; it does not say that every object or table payload is routed through KV as another sequential hop.

Direct GPU delivery is a design direction

The design leaves room for a path from DiskIO buffers through RDMA or GPUDirect Storage to GPU memory. That requires control over topology, buffer ownership and data lifetimes, not just another API.

Dataset and direct GPU delivery remain planned. There is no live GPU path, measured speedup, or production-readiness claim attached to this illustration.

Read the source contracts

This page summarizes access models and the storage core. The design documents remain in the CROWDB repository; for setup steps, use the user manual.

Implementation summary checked against the CROWDB design documents on September 28, 2026. Diagrams explain the design; they are not performance evidence.