Removing Nodes
Decommission safely - Never just stop a node!
⚠️ THE GOLDEN RULE
NEVER JUST STOP/KILL A NODE!
❌ WRONG WAY
✅ RIGHT WAY
🎯 Decommission = Graceful exit with data transfer
💀 Just stopping = Data loss!
🔀 Decommission vs Removenode
Two different commands for two different scenarios!
nodetool decommission
Graceful removal (PREFERRED)
When to use:
- ✅ Downsizing cluster
- ✅ Node is healthy and running
- ✅ Planned removal
- ✅ Normal operations
How it works:
- Streams data to other nodes
- Updates token ownership
- Leaves cluster gracefully
- Zero data loss ✅
nodetool removenode
Emergency removal (LAST RESORT)
When to use:
- 🚨 Node is DEAD (hardware failure)
- 🚨 Can't access node
- 🚨 Node won't start
- 🚨 Emergency only!
How it works:
- Forces removal from cluster
- NO data streaming
- Relies on replication (RF)
- Data loss if RF=1 ⚠️
Critical Decision Tree
✅ Graceful Decommission Process
The right way to remove a node!
Pre-Decommission Checks
Run Decommission
Monitor Progress
Phase 1: Status Change (Instant)
Node enters LEAVING state
- Gossip updated across cluster
- Status changes from UN to UL
- No new writes directed to this node
- Reads still served
Phase 2: Data Streaming (30 min - 6 hours)
Transfers all data to other nodes
- Streams SSTables to nodes that will own data
- Multiple parallel streams
- Progress shown in logs and netstats
- This is the longest phase!
Phase 3: Leaving Cluster (1-2 min)
Final cleanup and exit
- Verifies all data transferred
- Updates gossip state
- Removes self from cluster
- Cassandra process stops automatically
Post-Decommission
Decommission Complete!
- ✅ All data transferred to other nodes
- ✅ Zero data loss
- ✅ Cluster rebalanced automatically
- ✅ Node safely removed
- ✅ Can now repurpose/shutdown server
🚨 Emergency Removenode (Dead Node)
When node is permanently dead!
Use Only When
- 💀 Hardware failure (server won't boot)
- 💀 Network partition (can't reach node)
- 💀 Disk failure (can't read data)
- 💀 Node is permanently gone
- ⚠️ You have RF >= 2 (or you WILL lose data!)
Get Dead Node's Host ID
Start Removal Process
Monitor Removal
Handle Stuck Removal (If Needed)
After Emergency Removal
CRITICAL: Run repair on all remaining nodes!
📊 Verification Steps
Ensure removal was successful!
✅ Check Cluster Status
✅ Check Schema Agreement
✅ Verify Data Accessibility
✅ Check Load Distribution
🔧 Troubleshooting
Fix common issues!
❌ Decommission Stuck/Slow
❌ Can't Start Decommission
❌ Node Still Shows After Removal
❌ Data Missing After Removal
💡 Best Practices
Do it right!
DO
- Always decommission first
- Check RF >= 2 before removing
- Monitor decommission progress
- Remove during low traffic
- Verify cluster health after
- Run repair after removenode
- Document the process
- Take snapshot before removal
DON'T
- Just stop/kill the node
- Remove if RF=1 (data loss!)
- Interrupt decommission
- Remove during peak traffic
- Skip verification steps
- Use removenode for healthy nodes
- Remove multiple nodes at once
- Forget to update monitoring
Production Checklist
Complete this checklist:
- ✅ Pre-removal: Check RF >= 2, all nodes UN, take snapshot
- ✅ Decommission: Use nodetool decommission (if node is alive)
- ✅ OR Removenode: Use nodetool removenode (if node is dead)
- ✅ Monitor: Watch logs, check streaming progress
- ✅ Verify: Check status, schema, data access, load distribution
- ✅ Repair: Run repair if used removenode
- ✅ Update docs: Update inventory, monitoring, runbooks
🎉 You Can Safely Remove Nodes!
Congratulations! You now know how to remove nodes safely!
🎓 What You Learned:
- ⚠️ Golden rule: NEVER just stop a node!
- 🔀 Two methods: Decommission (preferred) vs Removenode (emergency)
- ✅ Graceful decommission: 3-phase process with data streaming
- 🚨 Emergency removenode: For dead nodes only
- 📊 Verification: Status, schema, data, load checks
- 🔧 Troubleshooting: Stuck decommission, ghost nodes, missing data
- 💡 Best practices: RF >= 2, monitor, verify, repair
💡 Key Takeaways:
- Decommission = Graceful - Streams data first
- Removenode = Emergency - For dead nodes only
- Check RF >= 2 - Or risk data loss!
- Monitor progress - Can take hours
- Verify after removal - All checks must pass
- Run repair - After emergency removal
📋 Quick Reference:
⚠️ Remember: NEVER just stop a node!
Always decommission first!
Responsive Ad