Node Replacement

Replacing Nodes

Replace dead nodes with replace_address - Maintain cluster size!

🔍 When Do You Replace a Node?

Node replacement is for when a node is DEAD and you want to bring in a NEW node to take its place!

✅ Replace Node When:

  • 💀 Hardware failure - Disk died, server won't boot
  • 💀 Node is permanently dead - Can't be recovered
  • 💀 Want to maintain cluster size - Keep same number of nodes
  • 💀 Node down > 3 hours - Hints expired, too much data to catch up
  • 💀 Data corruption - Beyond repair

❌ Don't Replace When:

  • ✋ Node is still running - Use nodetool decommission
  • ✋ Temporary failure - Just restart it!
  • ✋ Node down < 3 hours - Let hinted handoff catch it up
  • ✋ Scaling down - Use removenode instead

🔀 Replace vs Remove vs Add

Three different operations - Know the difference!

🔄

Replace Node

New node takes dead node's place

When:

  • 💀 Node is dead
  • 🔢 Keep cluster size same
  • 🎫 Inherit dead node's tokens

How:

  • New node starts with replace_address
  • Gets exact same token ranges
  • Streams data from replicas

Result: Same cluster size ✅

➖

Remove Node

Shrink cluster size

When:

  • ⬇️ Scaling down
  • 💰 Cost reduction
  • 🔄 Tokens redistributed

How:

  • Use nodetool decommission
  • OR nodetool removenode
  • Data moves to remaining nodes

Result: Cluster shrinks 📉

➕

Add Node

Grow cluster size

When:

  • ⬆️ Scaling up
  • 💪 More capacity needed
  • 🆕 Gets NEW tokens

How:

  • Start Cassandra normally
  • Auto-bootstrap enabled
  • Streams from neighbors

Result: Cluster grows 📈

Key Decision Tree

# Is the node dead/permanently gone? YES → Do you want to keep cluster size? YES → Replace node (this page!) NO → Remove node (shrink cluster) NO → Node is alive: → Use decommission to remove gracefully → OR just restart it if temporary issue # Want to grow cluster? → Use normal node addition (not replacement!)

✅ Prerequisites Before Replacement

Check these BEFORE starting!

Critical Checks

# 1. Verify node is truly dead ssh dead-node # Can't connect? Good - it's dead! # 2. Check replication factor cqlsh -e "DESCRIBE KEYSPACE my_keyspace" # Output must show RF >= 2: replication = {'class': 'NetworkTopologyStrategy', 'datacenter1': '3'} ← RF=3 ✅ # If RF=1, YOU WILL LOSE DATA! ❌ # 3. Check cluster health (from live node) nodetool status # Should see dead node as DN: UN 192.168.1.10 100 GB 16 33.3% abc123 UN 192.168.1.11 100 GB 16 33.3% def456 DN 192.168.1.15 100 GB 16 33.4% ghi789 ← Dead! # 4. Get the dead node's IP # Note: 192.168.1.15 (we'll need this!) # 5. Have new hardware ready # - Same or better specs # - Cassandra installed (same version!) # - Can be same IP or different IP

🔄 Complete Replacement Process

Step-by-step with commands!

1

Remove Dead Node from Cluster

Tell cluster to forget the dead node

# From a LIVE node, get dead node's Host ID nodetool status # Output: DN 192.168.1.15 100 GB 16 33.4% ghi789-... # ^^^^^^^^^^ # Host ID! # Remove the dead node nodetool removenode ghi789-abc-def-123 # Wait for completion (5-15 minutes) nodetool removenode status # Output: RemovalStatus: No token removals in process. # Verify node is gone nodetool status # Dead node should not appear!
2

Prepare New Node

Install and configure Cassandra on new hardware

# On NEW node, install Cassandra # Must be SAME VERSION as cluster! # Check cluster version first (from live node): nodetool version # ReleaseVersion: 4.1.0 # Install matching version on new node sudo apt install cassandra=4.1.0 # Edit cassandra.yaml on NEW node sudo vim /etc/cassandra/cassandra.yaml # Configure these settings: cluster_name: 'MyCluster' # Must match cluster! seeds: "192.168.1.10,192.168.1.11" # Live nodes (NOT dead node!) listen_address: 192.168.1.15 # Can be same or different IP rpc_address: 192.168.1.15 # Match listen_address endpoint_snitch: GossipingPropertyFileSnitch # Edit cassandra-rackdc.properties sudo vim /etc/cassandra/cassandra-rackdc.properties dc=datacenter1 # Must match cluster! rack=rack1 # Match dead node's rack # DON'T start Cassandra yet!
3

Start Replacement with replace_address

The magic parameter that makes it a replacement!

# METHOD 1: JVM Option (RECOMMENDED) # On NEW node: sudo JVM_OPTS="$JVM_OPTS -Dcassandra.replace_address=192.168.1.15" \ cassandra # Use dead node's IP address! # Even if new node has different IP # METHOD 2: cassandra.yaml # Edit cassandra.yaml: replace_address_first_boot: 192.168.1.15 # Then start: sudo systemctl start cassandra # What happens now: # 1. New node contacts seeds # 2. Announces "I'm replacing 192.168.1.15" # 3. Gets dead node's token ranges # 4. Starts streaming data from replicas

Phase 1: Joining (Instant)

Node announces replacement

# Check logs on NEW node tail -f /var/log/cassandra/system.log # You'll see: INFO Starting replacement of 192.168.1.15 INFO Replacement node taking tokens: [100, 200, 300...] INFO Node state jump to JOINING

Phase 2: Streaming (30 min - 6 hours)

Data streams from replicas

# Monitor streaming progress nodetool netstats # Output shows: Mode: JOINING Streaming from: /192.168.1.10 my_keyspace/users: 125543265/500000000 bytes (25%) Streaming from: /192.168.1.11 my_keyspace/orders: 456782340/1200000000 bytes (38%) # Check from another live node: nodetool status UN 192.168.1.10 100 GB 16 33.3% abc123 UN 192.168.1.11 100 GB 16 33.3% def456 UJ 192.168.1.15 45 GB 16 33.4% xyz999 ← Joining! #^^ UJ = Up and Joining

Phase 3: Normal (Complete!)

Replacement successful

# Logs show completion: INFO All streaming completed successfully INFO Node 192.168.1.15 state jump to NORMAL # Verify status: nodetool status UN 192.168.1.10 100 GB 16 33.3% abc123 UN 192.168.1.11 100 GB 16 33.3% def456 UN 192.168.1.15 100 GB 16 33.4% xyz999 ← Normal! ✅ # Replacement complete!
4

Verify Replacement Success

Ensure everything is working!

# 1. Check node status nodetool status # All nodes should be UN (Up Normal) # 2. Verify token ownership nodetool ring | grep 192.168.1.15 # Should show same tokens as dead node had # 3. Test data access cqlsh 192.168.1.15 SELECT count(*) FROM my_keyspace.users; # Count should match other nodes SELECT * FROM my_keyspace.users WHERE id=123; # Data should be readable # 4. Check data size nodetool tablestats my_keyspace.users | grep "Space used" # Should match other nodes (~100 GB) # 5. Run repair (IMPORTANT!) nodetool repair -pr my_keyspace # Catches any missed writes during replacement

📊 Monitoring Replacement Progress

Watch it carefully!

Key Commands

# Monitor streaming (on NEW node) nodetool netstats # Watch logs (on NEW node) tail -f /var/log/cassandra/system.log | grep -i "stream\|joining\|normal" # Check status from live node watch -n 10 'nodetool status' # Monitor disk space (on NEW node) watch -n 30 'df -h /var/lib/cassandra' # Check network throughput iftop -i eth0 # Check compaction activity nodetool compactionstats

Timeline Estimates

  • 💾 100 GB node: 30 min - 2 hours
  • 💾 500 GB node: 2-6 hours
  • 💾 2 TB node: 8-24 hours
  • ⚡ Factors: Network speed, disk I/O, RF, cluster load

💼 Real-World Scenarios

Complete examples!

Scenario 1: Disk Failure - Same IP

# Context: Production node's SSD died # Goal: Replace disk, keep same IP # BEFORE STATE: $ nodetool status UN 10.50.1.100 500 GB 256 33.3% node1-abc DN 10.50.1.101 500 GB 256 33.3% node2-def ← DEAD UN 10.50.1.102 500 GB 256 33.4% node3-ghi # STEP 1: Remove dead node $ nodetool removenode node2-def-456-789 Removal status: REMOVED # STEP 2: Replace disk hardware # - Install new SSD # - Keep same IP: 10.50.1.101 # STEP 3: Install Cassandra (matching version!) $ cassandra -v 4.1.0 ← Must match cluster # STEP 4: Configure cassandra.yaml $ sudo vim /etc/cassandra/cassandra.yaml cluster_name: 'Production' seeds: "10.50.1.100,10.50.1.102" listen_address: 10.50.1.101 rpc_address: 10.50.1.101 # STEP 5: Start with replace_address $ sudo JVM_OPTS="$JVM_OPTS -Dcassandra.replace_address=10.50.1.101" \ cassandra # STEP 6: Monitor streaming $ tail -f /var/log/cassandra/system.log INFO Replacing node 10.50.1.101 INFO Streaming from /10.50.1.100: 250 GB INFO Streaming from /10.50.1.102: 250 GB ... INFO Node state jump to NORMAL # AFTER STATE: $ nodetool status UN 10.50.1.100 500 GB 256 33.3% node1-abc UN 10.50.1.101 500 GB 256 33.3% node2-NEW ← Replaced! ✅ UN 10.50.1.102 500 GB 256 33.4% node3-ghi # Success! Same IP, new hardware

Scenario 2: Different IP (Cloud Migration)

# Context: Moving from on-prem to AWS # Old IP: 192.168.1.50 (on-prem) # New IP: 10.0.5.200 (AWS) # STEP 1: Dead node info $ nodetool status DN 192.168.1.50 500 GB 256 33.3% old-node # STEP 2: Remove dead node $ nodetool removenode old-node-host-id # STEP 3: Configure NEW node (AWS) # cassandra.yaml on 10.0.5.200: cluster_name: 'MyCluster' seeds: "10.0.5.201,10.0.5.202" ← AWS seeds listen_address: 10.0.5.200 ← NEW IP rpc_address: 10.0.5.200 # STEP 4: Start with OLD IP in replace_address! $ sudo JVM_OPTS="$JVM_OPTS -Dcassandra.replace_address=192.168.1.50" \ cassandra # ^^^^^^^^^^^^^^ # OLD IP here! # What happens: # - New node (10.0.5.200) says "I'm replacing 192.168.1.50" # - Gets 192.168.1.50's token ranges # - Streams data to 10.0.5.200 # - Cluster maps old tokens → new IP # RESULT: $ nodetool status UN 10.0.5.200 500 GB 256 33.3% new-node ← New IP! ✅ UN 10.0.5.201 500 GB 256 33.3% node2 UN 10.0.5.202 500 GB 256 33.4% node3 # Update your application connection strings! # Replace 192.168.1.50 → 10.0.5.200

Scenario 3: Emergency - Node Vanished

# Context: Node completely destroyed (fire, flood, etc) # Need IMMEDIATE replacement # STEP 1: Verify disaster $ ping 10.20.3.75 # No response - node is GONE $ nodetool status DN 10.20.3.75 ??? 256 33.3% disaster-node # STEP 2: Emergency removal with assassinate # (Faster than removenode) $ nodetool assassinate 10.20.3.75 # Immediately removes from gossip # STEP 3: Verify removal $ nodetool status # Node should be GONE instantly # STEP 4: Quick new node setup # Use automation/Ansible/Terraform if available # STEP 5: Start replacement $ sudo JVM_OPTS="$JVM_OPTS -Dcassandra.replace_address=10.20.3.75" \ cassandra # STEP 6: Aggressive monitoring $ watch -n 5 'nodetool netstats' # Timeline: # - assassinate: Instant # - Replacement: 1-4 hours (depends on data size) # - Total: Much faster than normal removenode # CAUTION: Only use assassinate when 100% sure node is dead!

🔧 Troubleshooting

Fix common issues!

❌ "Node already exists" Error

# ERROR: Can't start node with replace_address because node 192.168.1.15 already exists! # CAUSE: Didn't removenode before replacing # FIX: # 1. Stop new node $ sudo systemctl stop cassandra # 2. Remove dead node (from live node) $ nodetool removenode # 3. Verify removal $ nodetool status # Dead node should be gone # 4. Retry replacement $ sudo JVM_OPTS="$JVM_OPTS -Dcassandra.replace_address=192.168.1.15" \ cassandra

❌ Streaming Stuck/Slow

# SYMPTOM: No progress for >1 hour $ nodetool netstats Mode: JOINING Streaming: 145 GB / 500 GB (29%) ← Stuck! # DIAGNOSIS: # 1. Check network $ iftop -i eth0 # Is network saturated? # 2. Check disk on source nodes $ ssh source-node $ iostat -x 5 # %util near 100%? Disk is bottleneck # 3. Check for errors $ grep -i "error\|exception" /var/log/cassandra/system.log # FIXES: # Fix 1: Increase stream throughput # Edit cassandra.yaml: stream_throughput_outbound_megabits_per_sec: 400 # Default: 200 # Restart and try again # Fix 2: Check disk space on new node $ df -h /var/lib/cassandra # Need enough space for all data! # Fix 3: Disable compaction temporarily $ nodetool disableautocompaction # Re-enable after replacement: $ nodetool enableautocompaction

❌ Wrong Data After Replacement

# SYMPTOM: Data count doesn't match $ cqlsh 10.1.0.15 SELECT count(*) FROM users; 3500000 ← Expected: 5000000! # DIAGNOSIS: $ nodetool tablestats my_keyspace.users Space used: 120 GB ← Expected: ~170 GB # CAUSE: Incomplete streaming # FIX: Run full repair $ nodetool repair -full my_keyspace users # This will: # 1. Compare with replicas # 2. Stream missing data # 3. Ensure consistency # Wait for completion (1-3 hours) # Verify: SELECT count(*) FROM users; 5000000 ← Fixed! ✅

❌ Node Keeps Restarting

# SYMPTOM: Cassandra won't stay up $ systemctl status cassandra Active: activating (auto-restart) # CHECK LOGS: $ tail -100 /var/log/cassandra/system.log # Common errors and fixes: # Error 1: replace_address still set ERROR: replace_address is set but node is not replacing! # Fix: Remove from cassandra.yaml after first boot $ sudo sed -i '/replace_address/d' /etc/cassandra/cassandra.yaml $ sudo systemctl restart cassandra # Error 2: Version mismatch ERROR: Cannot start version 4.0.0 in cluster running 4.1.0 # Fix: Install exact matching version! $ sudo apt install cassandra=4.1.0 # Error 3: Out of memory ERROR: java.lang.OutOfMemoryError # Fix: Increase JVM heap in cassandra-env.sh MAX_HEAP_SIZE="16G" # Increase from 8G

💡 Best Practices

Do it right!

✅

DO

  • Verify node is truly dead first
  • Remove dead node before replacing
  • Match Cassandra version exactly
  • Use same cluster_name
  • Monitor streaming progress
  • Run repair after completion
  • Replace during low traffic
  • Take snapshot before replacing
❌

DON'T

  • Replace if node is still alive
  • Skip removenode step
  • Use different Cassandra version
  • Wrong cluster_name
  • Interrupt streaming
  • Skip post-replacement repair
  • Replace multiple nodes at once
  • Forget to remove replace_address

Production Checklist

Complete this before replacing:

  1. ✅ Verify node is dead (can't SSH, won't boot)
  2. ✅ Check RF >= 2 (or data loss!)
  3. ✅ Check cluster health (other nodes UN)
  4. ✅ Note dead node's IP address
  5. ✅ Remove dead node with removenode
  6. ✅ Prepare new hardware (same or better)
  7. ✅ Install matching Cassandra version
  8. ✅ Configure cassandra.yaml (cluster_name, seeds, etc)
  9. ✅ Start with replace_address
  10. ✅ Monitor streaming to completion
  11. ✅ Verify status (UN), data, tokens
  12. ✅ Run nodetool repair -pr
  13. ✅ Remove replace_address from config
  14. ✅ Update documentation/inventory

🎉 You Can Replace Nodes Safely!

Congratulations! You now know how to replace dead nodes!

🎓 What You Learned:

  • 🔍 When to replace: Dead nodes that need swapping
  • 🔀 Replace vs Remove: Keep size vs shrink cluster
  • ✅ Prerequisites: RF >= 2, removenode first
  • 🔄 The process: 4 steps with replace_address
  • 📊 Monitoring: netstats, logs, status
  • 💼 Real scenarios: Same IP, different IP, emergencies
  • 🔧 Troubleshooting: Stuck streaming, wrong data, restarts

💡 Key Takeaways:

  1. Replacement maintains cluster size - Dead node → New node
  2. Always removenode first - Clear out dead node
  3. replace_address is magic - New node gets dead node's tokens
  4. Can change IP - Old IP in replace_address, new IP in config
  5. Monitor carefully - Can take hours to stream data
  6. Run repair after - Catches missed writes

📋 Quick Reference:

# Step 1: Remove dead node (from live node) nodetool status # Get host ID nodetool removenode # Step 2: Prepare new node # Install Cassandra, configure cassandra.yaml # Step 3: Start replacement (on new node) sudo JVM_OPTS="$JVM_OPTS -Dcassandra.replace_address=DEAD_NODE_IP" \ cassandra # Step 4: Monitor and verify nodetool netstats # Watch streaming nodetool status # Check status nodetool repair -pr # After completion

✅ Now you can handle node failures like a pro!
Dead node? No problem - just replace it!

Advertisement

📱 Responsive Ad 📱