Browse documentation
CROWDB / USER MANUAL

Cluster setup and management

Bootstrap a KV cluster and manage replicas and physical nodes.

Development guide · September 28, 2026 snapshot

2. Advanced: Bootstrap a KV Cluster

The remaining sections describe lower-level cluster and server administration. They are not required for the local S3 workflow above.

Management CLI commands omit --system-ip and --system-port for brevity. They default to the system-group discovery endpoint 127.0.0.1:10000; either flag may point to any system-group node because leader discovery is automatic. The CROWDB_SYSTEM_IP/CROWDB_SYSTEM_PORT environment variables provide the same overrides. The following console HTTP curl examples assume:

IP=127.0.0.1
PORT=14000

An S3 mini-cluster already starts crowdb-web with its persisted registry. For a separately managed cluster, before using these commands:

  • Start crowdb-web with crowdb-web --port 14000. Add --test-mode for an in-memory console configuration that is lost on restart.
  • Set CROWDB_KV_SERVER_BIN when crowdb-kv-server is not next to crowdb-web and not available through PATH.
  • Use each machine's reachable hostname or IP instead of 127.0.0.1 for a multi-machine deployment.

Server configuration files

KV server, diskdb, chunkdb, and diskio use TOML startup configuration. Valid templates are shipped in each server's conf/ directory. Values resolve in this order: compiled defaults, then file values, then CLI options that were explicitly supplied. A CLI option you omit does not erase its file value.

KV server and diskio make --config optional; diskdb and chunkdb require a config path. A malformed, unreadable, or invalid named file stops startup instead of silently falling back to defaults. server.rpc_workers controls the inbound RPC worker count (default 2 for the Rust servers and 4 for diskio), must be positive, and takes effect only after restart. File watchers may report static changes before restart, but the active listener does not change.

Console local deployment generates node-specific config files and reuses the same paths when restarting services. Keep manually managed files with the server's data and deployment records; group 0 currently stores topology, not process configuration.

2.1 Register the physical topology

Create a rack, add nodes, and deploy a server on each node. The deploy command starts crowdb-kv-server on the target node (via SSH if ssh_user is set, or as a local subprocess otherwise). No manual start needed.

CLI:

# Create a rack
crowdb-cli rack add --id r1 --name "rack-one"

# Register each node (repeat for n2, n3)
crowdb-cli node add --id n1 --rack r1 --host 127.0.0.1

# Deploy a crowdb-kv-server process on each node (repeat for n2, n3)
crowdb-cli server deploy --node n1 --rest-port 2001 --rpc-port 20001

curl:

# Create a rack
curl -X POST "http://$IP:$PORT/api/racks" -H 'Content-Type: application/json' \
  -d '{"id":"r1"}'

# Register each node (repeat for n2, n3)
curl -X POST "http://$IP:$PORT/api/nodes" -H 'Content-Type: application/json' \
  -d '{"id":"n1","rack_id":"r1","host":"127.0.0.1","ssh_port":22,"ssh_user":""}'

# Deploy a crowdb-kv-server process on each node (repeat for n2, n3)
curl -X POST "http://$IP:$PORT/api/nodes/n1/server/deploy" \
  -H 'Content-Type: application/json' \
  -d '{"rest_port":2001,"rpc_port":20001}'

2.2 Initialize the cluster

Before creating data stores or groups, the cluster must be initialized. This creates the system group (store 0, group 0) which stores cluster topology metadata as KV entries, providing HA for the topology itself.

CLI:

# Initialize with all deployed nodes
crowdb-cli cluster init --nodes n1,n2,n3

curl:

curl -X POST "http://$IP:$PORT/api/cluster/init" \
  -H 'Content-Type: application/json' \
  -d '{"nodes":["n1","n2","n3"]}'

This creates store 0 and group 0 on each selected node, wires remotes for multi-node, persists topology in console config, and writes hardware hierarchy + KV-cluster topology into group 0 via HardwareClient + KVClusterMetaClient (text-path keys, JSON values). After initialization, data store/group creation is unblocked.

For a single-node dev cluster, pass one node:

crowdb-cli cluster init --nodes n1

2.3 Create a store and group

A store is the logical container that owns one or more groups.

CLI:

# Create a store on n1
crowdb-cli store add --store-id 3 --nodes n1

# Create a group with an initial replica on n1
crowdb-cli paxos add \
  --store-id 3 --group-id 3 --replica-id 1 --nodes n1

curl:

curl -X POST "http://$IP:$PORT/api/stores" -H 'Content-Type: application/json' \
  -d '{"store_id":3,"nodes":["n1"]}'

curl -X POST "http://$IP:$PORT/api/stores/3/groups" -H 'Content-Type: application/json' \
  -d '{"group_id":3,"replica_id":1,"nodes":["n1"]}'

If the cluster has not been initialized, store/group creation returns 409 Conflict with a message directing you to run cluster init first.

2.4 Add the remaining replicas

CLI:

crowdb-cli replica add \
  --store-id 3 --group-id 3 --node n2 --replica-id 2

crowdb-cli replica add \
  --store-id 3 --group-id 3 --node n3 --replica-id 3

curl:

curl -X POST "http://$IP:$PORT/api/stores/3/groups/3/replicas" \
  -H 'Content-Type: application/json' \
  -d '{"node_id":"n2","replica_id":2}'

curl -X POST "http://$IP:$PORT/api/stores/3/groups/3/replicas" \
  -H 'Content-Type: application/json' \
  -d '{"node_id":"n3","replica_id":3}'

The service orchestrates the full add-replica flow: creates the local group on the target node, wires remotes bidirectionally, and the new replica catches up via snapshot streaming before joining the voting set.

2.5 Verify and smoke test

CLI:

# Check group health
crowdb-cli paxos inspect --store-id 3 --group-id 3
# Look for "leader=" and replica states

# Put / Get
crowdb-cli kv put --store-id 3 --group-id 3 \
  --key hello --value world

crowdb-cli kv get --store-id 3 --group-id 3 --key hello

curl:

curl "http://$IP:$PORT/api/stores/3/groups/3"

curl -X POST "http://$IP:$PORT/api/stores/3/groups/3/kv/put" \
  -H 'Content-Type: application/json' \
  -d '{"key":"hello","value":"world"}'

curl "http://$IP:$PORT/api/stores/3/groups/3/kv/get?key=hello"

4. Cluster Management

4.1 Check cluster health

CLI:

# High-level summary (servers + store/group counts)
crowdb-cli cluster status

# Full topology (logical stores/groups/replicas + physical nodes/servers)
crowdb-cli cluster topology

# Inspect a specific store, group, or node
crowdb-cli cluster inspect s3          # store 3
crowdb-cli cluster inspect s3/g3       # group 3 in store 3
crowdb-cli cluster inspect n1          # node n1

curl:

# All nodes
curl "http://$IP:$PORT/api/nodes"

# All deployed servers
curl "http://$IP:$PORT/api/servers"

# A specific group
curl "http://$IP:$PORT/api/stores/3/groups/3"
# healthy: all replicas up, leader known
# degraded: some replicas down, quorum + leader available
# unavailable: quorum lost

4.2 Add a read replica

CLI:

crowdb-cli replica add --store-id 3 --group-id 3 --node n4 --replica-id 4

curl:

curl -X POST "http://$IP:$PORT/api/stores/3/groups/3/replicas" \
  -H 'Content-Type: application/json' \
  -d '{"node_id":"n4","replica_id":4}'

The new replica streams a snapshot from the leader, catches up, then joins the voting set automatically.

4.3 Remove a replica

CLI:

crowdb-cli replica remove --store-id 3 --group-id 3 --replica-id 3

curl:

curl -X DELETE "http://$IP:$PORT/api/stores/3/groups/3/replicas/3"

If the target is the leader, the service asks it to step down first, waits for a new leader, then removes the replica.

4.4 Replace a failed node

  1. Provision the new machine with the same node ID, management port, and RPC port.

  2. Deploy the server via the service. The server auto-loads its store/group configuration from conf/node-config.json on startup. No --stores/--groups CLI args needed for normal restart:

    CLI:

    crowdb-cli server deploy --node n1 --rest-port 2001 --rpc-port 20001
    

    curl:

    curl -X POST "http://$IP:$PORT/api/nodes/n1/server/deploy" \
      -H 'Content-Type: application/json' \
      -d '{"rest_port":2001,"rpc_port":20001}'
    

    If node-config.json is lost, fall back to explicit bootstrap args by starting crowdb-kv-server manually with --stores/--groups/ --replica:

    crowdb-kv-server \
      --management-addr 0.0.0.0 --management-port 2001 \
      --ports 20001 --election-profile default \
      --stores 3 --groups 3 --replica 1
    
  3. Verify group health.

If the WAL and config directory were also lost, add the replacement as a new replica with a new replica ID instead of reusing the old one.


Migrated from the CROWDB user manual on September 28, 2026. Documentation content is licensed under Apache-2.0.