Browse documentation
CROWDB / USER MANUAL

Upgrade, recovery and backup

Rolling upgrades, loss of quorum and backup guidance.

Development guide · September 28, 2026 snapshot

5. Rolling Upgrade

Upgrade one node at a time. Wait for each node to rejoin and catch up before moving to the next.

For each node:

  1. Stop:

    CLI:

    crowdb-cli server stop --node n1
    

    curl:

    curl -X POST "http://$IP:$PORT/api/nodes/n1/server/stop"
    
  2. Install the new binary on the node.

  3. Restart the server. The server auto-loads its store/group configuration from conf/node-config.json on startup:

    CLI:

    crowdb-cli server restart --node n1
    

    curl:

    curl -X POST "http://$IP:$PORT/api/nodes/n1/server/restart"
    

    If node-config.json is missing, start crowdb-kv-server manually with explicit args:

    crowdb-kv-server \
      --management-addr 0.0.0.0 --management-port 2001 \
      --ports 20001 --election-profile default \
      --stores 3 --groups 3 --replica 1
    

    --stores/--groups tells the server to reopen the WAL and rejoin as a full member. --replica must match the assigned replica ID.

  4. Wait for healthy:

    crowdb-cli cluster status
    crowdb-cli paxos inspect --store-id 3 --group-id 3
    
  5. Smoke test:

    crowdb-cli kv get --store-id 3 --group-id 3 --key hello
    
  6. Move to the next node.

What to watch: after stopping a node, the remaining nodes elect a new leader. Wait for the group view to show a leader before proceeding. A brief latency spike during leader transition is normal.


6. Emergency: Loss of Quorum

If two of three nodes fail, the remaining node cannot elect itself leader. Writes and linearizable reads block.

  • Restore the failed nodes from backups and restart. The server auto-loads from conf/node-config.json; if the config is lost, fall back to --stores/--groups/--replica args. This is always the safest path.
  • Recover with data loss (last resort): force the surviving node to become leader by manually truncating the log. Only safe when the other nodes are permanently lost.

Do not add a new node to a quorum-less group without first recovering leadership.


7. Backup

CROWDB durability comes from the per-store WAL (--wal-root), the per-node config cache (--config-root), and the durable KV engine (--data-root). For disaster recovery, back up:

  • {wal-root}/store{store_id}/ for each store
  • {config-root}/node-config.json — per-node store/group config cache
  • {data-root}/store{store_id}/group{group_id}/ if using crowdb-tree durable KV engine

Restore by placing these on the replacement node and starting the server. With node-config.json present, no --stores/--groups bootstrap args are needed. If the config is lost, use explicit --stores/--groups/--replica args to recover from WAL.


Migrated from the CROWDB user manual on September 28, 2026. Documentation content is licensed under Apache-2.0.