Adding Nodes
Scale your cluster safely and efficiently!
๐ Why Add Nodes to Your Cluster?
Growing your cluster = More capacity, better performance!
๐ฏ Common Reasons to Add Nodes
- ๐พ Disk capacity: Running out of space (>70% full)
- ๐ Throughput: Need more read/write capacity
- โก Performance: Latency increasing under load
- ๐ก๏ธ Redundancy: Improve fault tolerance (RF=3 minimum)
- ๐ Geographic expansion: New datacenter
- ๐ Growth planning: Before hitting limits
๐ Real Example: Netflix
Netflix's Cassandra clusters scale constantly:
- ๐ฌ Peak traffic: Friday nights, new releases
- โ Scale up: Add 20-30% more nodes ahead of peak
- ๐ Result: Zero downtime during surges
- ๐ฐ Cost: Add nodes incrementally vs over-provisioning
โก Cassandra makes scaling easy - add nodes anytime!
When to Add
Scale proactively!
- ๐พ Disk > 70% full
- ๐ฅ CPU > 80% sustained
- ๐ Latency increasing
- โ ๏ธ Thread pool pending tasks
- ๐ Growth projections
Don't Wait For
Too late!
- ๐พ Disk 95% full
- ๐ฅ Constant OOM errors
- ๐ Severe performance degradation
- ๐จ Production incidents
- ๐ฐ Customer complaints
๐ Planning Your Node Addition
Think before you scale!
Capacity Planning
How many nodes do you need?
Token Allocation
Cassandra handles this automatically (vnodes)
Timeline Expectations
Adding a node takes time - plan accordingly:
- โ๏ธ Preparation: 15-30 minutes (config, network, etc)
- โฑ๏ธ Bootstrap: 1-6 hours (depends on data size)
- ๐ Streaming: Most time-consuming part
- ๐งน Cleanup: 30 minutes - 2 hours (per old node)
- ๐ Total: Plan for 4-12 hours for complete process
๐ง Preparing the New Node
Get everything ready first!
Hardware/VM Setup
Install Cassandra
Configure cassandra.yaml
Clean Data Directories
Configuration Checklist
VERIFY before starting node:
- โ cluster_name matches existing cluster EXACTLY
- โ seeds point to existing seed nodes
- โ listen_address is this node's IP
- โ rpc_address is this node's IP
- โ endpoint_snitch matches cluster
- โ Cassandra version matches cluster
- โ Data directories are empty
- โ If ANY of these are wrong, node won't join correctly!
โ Adding the Node - Bootstrap Process
Let's do this!
Start Cassandra
Monitor Bootstrap Progress
Phase 1: Joining Ring (1-2 min)
Node discovers cluster via gossip
- Connects to seed nodes
- Learns cluster topology
- Gets token assignments
- Status: UJ (Up and Joining)
Phase 2: Streaming Data (1-6 hours)
Node receives data it will own
- Downloads data from multiple nodes
- Receives SSTables for its token ranges
- Progress shown in logs
- Can monitor with nodetool netstats
Phase 3: Finalizing (5-10 min)
Node prepares for production
- Builds indexes
- Compacts received data
- Updates gossip state
- Status changes to UN (Up and Normal)
Bootstrap Complete!
Node is now part of the cluster and serving traffic!
- โ Status shows UN (Up and Normal)
- โ Load column shows data size
- โ Owns approximately equal token %
- โ Accepting read/write requests
- ๐ But wait - you're not done yet! Cleanup needed...
โ Post-Addition Tasks
Critical cleanup steps!
Run Cleanup on OLD Nodes
CRITICAL: Old nodes still have data they no longer own!
โ ๏ธ Don't skip this! Cleanup frees disk space and ensures correct data distribution.
Run Repair (Optional but Recommended)
Verify Cluster Health
๐ข Adding Multiple Nodes
Scale up faster!
Important Rules
When adding multiple nodes:
- โ ๏ธ Add ONE at a time! Wait for each to complete
- โ ๏ธ Don't start multiple simultaneously (overwhelms cluster)
- โฐ Wait for UN status before starting next
- ๐งน Run cleanup after ALL nodes added (not between each)
- โ Order doesn't matter (vnodes handle distribution)
Best Practice: Sequential Addition
๐ Troubleshooting
Fix common issues!
โ Node Won't Join Cluster
Check configuration
โ Bootstrap Stuck/Slow
Network or resource issue
โ Out of Memory During Bootstrap
Increase heap size
โ Schema Disagreement
Nodes have different schemas
๐ก Best Practices
Do it right!
DO
- Plan ahead (capacity, timing)
- Add during low traffic
- Match versions exactly
- Wait for UN before next node
- Run cleanup on old nodes
- Monitor bootstrap progress
- Verify cluster health
- Document the process
DON'T
- Add during peak traffic
- Start multiple nodes at once
- Use different Cassandra versions
- Skip cleanup step
- Interrupt bootstrap process
- Add when disk >90% full
- Forget to enable auto-start
- Mix single-token and vnodes
Production Checklist
Complete this checklist for each node:
- โ Pre-addition: Verify capacity need, plan timing, prepare hardware
- โ Configuration: Match cluster name, set seeds, configure IPs
- โ Bootstrap: Start node, monitor logs, verify UN status
- โ Cleanup: Run on old nodes (not new node!)
- โ Verification: Check status, schema, data access
- โ Documentation: Update inventory, runbooks
๐ You Can Scale Your Cluster!
Congratulations! You now know how to add nodes safely!
๐ What You Learned:
- ๐ Why add nodes: Capacity, performance, redundancy
- ๐ Planning: Capacity calculations, timeline expectations
- ๐ง Preparation: Hardware, installation, configuration
- โ Bootstrap: 3-phase joining process
- โ Post-addition: Cleanup, repair, verification
- ๐ข Multiple nodes: Sequential addition strategy
- ๐ Troubleshooting: Common issues and fixes
- ๐ก Best practices: Production-ready workflow
๐ก Key Takeaways:
- Add one at a time - Never start multiple simultaneously
- Match configuration - cluster_name, seeds, version
- Monitor bootstrap - Watch for UN status
- Run cleanup - On OLD nodes after addition
- Plan for time - 4-12 hours for complete process
- Scale proactively - Before hitting 70% capacity
๐ Quick Reference:
โก Cassandra scales elastically - add capacity anytime!
Responsive Ad