Replacing Nodes
Replace dead nodes with replace_address - Maintain cluster size!
🔍 When Do You Replace a Node?
Node replacement is for when a node is DEAD and you want to bring in a NEW node to take its place!
✅ Replace Node When:
- 💀 Hardware failure - Disk died, server won't boot
- 💀 Node is permanently dead - Can't be recovered
- 💀 Want to maintain cluster size - Keep same number of nodes
- 💀 Node down > 3 hours - Hints expired, too much data to catch up
- 💀 Data corruption - Beyond repair
❌ Don't Replace When:
- ✋ Node is still running - Use
nodetool decommission - ✋ Temporary failure - Just restart it!
- ✋ Node down < 3 hours - Let hinted handoff catch it up
- ✋ Scaling down - Use
removenodeinstead
🔀 Replace vs Remove vs Add
Three different operations - Know the difference!
Replace Node
New node takes dead node's place
When:
- 💀 Node is dead
- 🔢 Keep cluster size same
- 🎫 Inherit dead node's tokens
How:
- New node starts with
replace_address - Gets exact same token ranges
- Streams data from replicas
Result: Same cluster size ✅
Remove Node
Shrink cluster size
When:
- ⬇️ Scaling down
- 💰 Cost reduction
- 🔄 Tokens redistributed
How:
- Use
nodetool decommission - OR
nodetool removenode - Data moves to remaining nodes
Result: Cluster shrinks 📉
Add Node
Grow cluster size
When:
- ⬆️ Scaling up
- 💪 More capacity needed
- 🆕 Gets NEW tokens
How:
- Start Cassandra normally
- Auto-bootstrap enabled
- Streams from neighbors
Result: Cluster grows 📈
Key Decision Tree
✅ Prerequisites Before Replacement
Check these BEFORE starting!
Critical Checks
🔄 Complete Replacement Process
Step-by-step with commands!
Remove Dead Node from Cluster
Tell cluster to forget the dead node
Prepare New Node
Install and configure Cassandra on new hardware
Start Replacement with replace_address
The magic parameter that makes it a replacement!
Phase 1: Joining (Instant)
Node announces replacement
Phase 2: Streaming (30 min - 6 hours)
Data streams from replicas
Phase 3: Normal (Complete!)
Replacement successful
Verify Replacement Success
Ensure everything is working!
📊 Monitoring Replacement Progress
Watch it carefully!
Key Commands
Timeline Estimates
- 💾 100 GB node: 30 min - 2 hours
- 💾 500 GB node: 2-6 hours
- 💾 2 TB node: 8-24 hours
- ⚡ Factors: Network speed, disk I/O, RF, cluster load
💼 Real-World Scenarios
Complete examples!
Scenario 1: Disk Failure - Same IP
Scenario 2: Different IP (Cloud Migration)
Scenario 3: Emergency - Node Vanished
🔧 Troubleshooting
Fix common issues!
❌ "Node already exists" Error
❌ Streaming Stuck/Slow
❌ Wrong Data After Replacement
❌ Node Keeps Restarting
💡 Best Practices
Do it right!
DO
- Verify node is truly dead first
- Remove dead node before replacing
- Match Cassandra version exactly
- Use same cluster_name
- Monitor streaming progress
- Run repair after completion
- Replace during low traffic
- Take snapshot before replacing
DON'T
- Replace if node is still alive
- Skip removenode step
- Use different Cassandra version
- Wrong cluster_name
- Interrupt streaming
- Skip post-replacement repair
- Replace multiple nodes at once
- Forget to remove replace_address
Production Checklist
Complete this before replacing:
- ✅ Verify node is dead (can't SSH, won't boot)
- ✅ Check RF >= 2 (or data loss!)
- ✅ Check cluster health (other nodes UN)
- ✅ Note dead node's IP address
- ✅ Remove dead node with
removenode - ✅ Prepare new hardware (same or better)
- ✅ Install matching Cassandra version
- ✅ Configure cassandra.yaml (cluster_name, seeds, etc)
- ✅ Start with
replace_address - ✅ Monitor streaming to completion
- ✅ Verify status (UN), data, tokens
- ✅ Run
nodetool repair -pr - ✅ Remove
replace_addressfrom config - ✅ Update documentation/inventory
🎉 You Can Replace Nodes Safely!
Congratulations! You now know how to replace dead nodes!
🎓 What You Learned:
- 🔍 When to replace: Dead nodes that need swapping
- 🔀 Replace vs Remove: Keep size vs shrink cluster
- ✅ Prerequisites: RF >= 2, removenode first
- 🔄 The process: 4 steps with
replace_address - 📊 Monitoring: netstats, logs, status
- 💼 Real scenarios: Same IP, different IP, emergencies
- 🔧 Troubleshooting: Stuck streaming, wrong data, restarts
💡 Key Takeaways:
- Replacement maintains cluster size - Dead node → New node
- Always removenode first - Clear out dead node
- replace_address is magic - New node gets dead node's tokens
- Can change IP - Old IP in replace_address, new IP in config
- Monitor carefully - Can take hours to stream data
- Run repair after - Catches missed writes
📋 Quick Reference:
✅ Now you can handle node failures like a pro!
Dead node? No problem - just replace it!
📱 Responsive Ad 📱