Rolling Restart
Update your cluster node-by-node with ZERO downtime!
📖 The Story: Tom's 3AM Disaster
Tom needed to update Cassandra config on his 6-node production cluster. He restarted all nodes AT ONCE. Cluster went down. 15,000 users got errors. CEO woke him at 3AM screaming. $500K revenue lost in 2 hours. All because he didn't know about rolling restarts.
😱 The Wrong Way (What Tom Did)
Friday, 8PM - "Quick Config Change":
8:03 PM - The Cascade:
The Damage:
- 💥 8:03 PM: All 6 nodes down, cluster unreachable
- 📱 8:05 PM: PagerDuty explodes (500+ alerts)
- 😡 8:07 PM: Customer complaints flood support
- 🔄 8:10 PM: Nodes start coming back up slowly
- ⚠️ 8:30 PM: Cluster still not forming (bootstrap issues)
- 📞 9:00 PM: Tom escalates to senior DBA
- 💰 10:00 PM: Cluster finally stable after 2 hours
Final Cost:
- 💸 $500K: Lost revenue (2 hours downtime)
- 😠 2,000: Angry customers
- 📉 15%: Churn spike
- 💼 Tom: Demoted to junior
✅ The Right Way (Rolling Restart)
What Tom Should Have Done:
The Result:
- ✅ 15 minutes: All nodes updated
- ✅ Zero downtime: Users didn't notice
- ✅ $0 lost: No impact on revenue
- ✅ 0 complaints: Perfect deployment
- ✅ Tom promoted: "Excellent ops skills!"
Rolling restart: The difference between disaster and success! 🎯
🎯 What Is a Rolling Restart?
Restart nodes one-by-one to maintain availability!
Full Restart (BAD)
Restart all nodes at once
What happens:
- 💥 Entire cluster goes down
- 💥 All data unavailable
- 💥 Users get errors
- 💥 Revenue stops
- 💥 SLAs violated
Downtime: 5-30 minutes
NEVER DO THIS IN PRODUCTION!
Rolling Restart (GOOD)
Restart one node at a time
What happens:
- ✅ Only 1 node down at a time
- ✅ RF ensures data available
- ✅ Users don't notice
- ✅ Zero revenue impact
- ✅ SLAs maintained
Downtime: ZERO!
ALWAYS USE IN PRODUCTION!
How Rolling Restart Works
The secret: Replication Factor (RF)!
Key Concepts
1. Quorum Math
2. Wait Between Restarts
3. Order Matters
📋 When Do You Need a Rolling Restart?
Common scenarios!
Config Changes
- cassandra.yaml updates
- JVM settings (jvm.options)
- Timeout adjustments
- Cache sizes
- Compaction settings
Most common reason!
Security Updates
- Enable authentication
- Enable encryption
- SSL certificate update
- Audit logging enable
- Password changes
Critical for security!
JVM Tuning
- Heap size changes
- GC tuning
- JVM flags
- Memory settings
- Thread pool config
Performance tuning!
Bug Fixes
- Apply hotfixes
- Driver updates
- Plugin updates
- Clear corrupted state
- Force clean restart
Emergency fixes!
OS/Hardware
- OS kernel updates
- Disk driver updates
- Network driver updates
- Memory upgrades
- After reboot
Infrastructure changes!
Performance
- Clear caches
- Reset metrics
- Apply optimizations
- Reload config
- Fresh start
Optimization!
What DOESN'T Need a Restart
These changes take effect immediately (no restart):
- ✅ Schema changes: CREATE/ALTER/DROP TABLE
- ✅ User management: CREATE/ALTER ROLE, GRANT/REVOKE
- ✅ Compaction: ALTER TABLE compaction
- ✅ TTL/GC grace: ALTER TABLE gc_grace_seconds
- ✅ Dynamic config: Some settings via nodetool
Always check documentation before restarting!
🔄 The Rolling Restart Process
Step-by-step with commands!
Pre-Flight Check
Verify cluster health BEFORE starting
Update Config Files
Make changes on ALL nodes BEFORE restarting
Restart First Node
Start the rolling restart
Wait and Verify
CRITICAL: Don't rush to next node!
Repeat for All Nodes
One by one, with patience
Post-Restart Verification
Confirm success!
🚀 Advanced Techniques
Pro-level strategies!
1. Rack-Aware Rolling Restart
2. Multi-DC Rolling Restart
3. Canary Restart
4. Parallel Rolling Restart
ADVANCED: Only with RF >= 5
Restart 2 nodes at once (if RF >= 5):
🤖 Automation Scripts
Automate the process safely!
Bash Script: rolling_restart.sh
Usage
🔧 Troubleshooting Rolling Restarts
Fix common issues!
❌ Node Won't Start
❌ Cluster Loses Quorum
❌ Node Stuck in Joining
❌ Increased Latency During Restart
💡 Rolling Restart Best Practices
Do it right every time!
DO
- Check cluster health first
- Update config on ALL nodes
- Restart ONE node at a time
- Wait for UN before next
- Wait 2-5 minutes between
- Test on canary first
- Schedule during low traffic
- Monitor the entire time
DON'T
- Restart all nodes at once
- Restart 2+ with RF=3
- Skip pre-flight checks
- Rush between nodes
- Do during peak traffic
- Forget to monitor
- Leave auto_bootstrap on
- Skip post-verification
Rolling Restart Checklist
Complete this checklist for every rolling restart:
| Phase | Task | Done? |
|---|---|---|
| Pre-Flight | ☐ All nodes UN | |
| ☐ RF >= 3 verified | ||
| ☐ No active repairs | ||
| ☐ Low pending compactions | ||
| Config | ☐ Updated on ALL nodes | |
| ☐ Syntax validated | ||
| ☐ auto_bootstrap: false | ||
| Restart | ☐ One node at a time | |
| ☐ Wait for UN each time | ||
| ☐ 2-5 min wait between | ||
| Post | ☐ All nodes UN | |
| ☐ Config applied | ||
| ☐ Test query works | ||
| ☐ No application errors |
Timeline Expectations
| Cluster Size | Time Per Node | Total Time |
|---|---|---|
| 3 nodes | ~4 minutes | ~12 minutes |
| 6 nodes | ~4 minutes | ~24 minutes |
| 12 nodes | ~4 minutes | ~48 minutes |
| 24 nodes | ~4 minutes | ~96 minutes (1.6 hrs) |
Note: ~2 min restart + 2 min wait = 4 min per node
🎉 Master Rolling Restarts!
You now know how to update Cassandra with ZERO downtime!
🎓 What You Learned:
- 🎯 What: Restart one node at a time
- ❌ Tom's disaster: Restarted all at once = $500K lost
- ✅ The right way: One at a time = zero downtime
- 🔄 The process: 6 steps from pre-flight to verification
- 🚀 Advanced: Rack-aware, multi-DC, canary
- 🤖 Automation: rolling_restart.sh script
- 🔧 Troubleshooting: Fix common issues
💡 Key Takeaways:
- ONE at a time - Never restart 2+ with RF=3
- WAIT between restarts - 2-5 minutes for stability
- RF >= 3 required - Or you'll lose quorum
- Monitor constantly - Watch status, logs, metrics
- Use automation - Scripts prevent human error
- Verify after - Test queries, check config applied
📋 The Golden Rule:
🚀 With RF=3: Restart ONE at a time, WAIT for UN, repeat
Result: ZERO downtime, happy users, safe job! ✅
🔄 Quick Reference:
🔁 Remember: Rolling restart = Zero downtime magic! 🎯
Don't be Tom - do it right!
📱 Responsive Ad 📱