Zero Downtime Updates

Rolling Restart

Update your cluster node-by-node with ZERO downtime!

📖 The Story: Tom's 3AM Disaster

Tom needed to update Cassandra config on his 6-node production cluster. He restarted all nodes AT ONCE. Cluster went down. 15,000 users got errors. CEO woke him at 3AM screaming. $500K revenue lost in 2 hours. All because he didn't know about rolling restarts.

😱 The Wrong Way (What Tom Did)

Friday, 8PM - "Quick Config Change":

-- Tom's cluster (6 nodes, RF=3): $ nodetool status UN 10.0.1.10 500 GB ← Node 1 UN 10.0.1.11 500 GB ← Node 2 UN 10.0.1.12 500 GB ← Node 3 UN 10.0.1.13 500 GB ← Node 4 UN 10.0.1.14 500 GB ← Node 5 UN 10.0.1.15 500 GB ← Node 6 -- Tom: "Let's update read_request_timeout" -- Edits cassandra.yaml on ALL nodes -- Tom's fatal mistake: $ ansible cassandra-cluster -m shell -a "systemctl restart cassandra" ↑ RESTARTS ALL 6 NODES AT ONCE! 💥

8:03 PM - The Cascade:

-- All 6 nodes go DOWN simultaneously: $ nodetool status DN 10.0.1.10 ← DOWN! DN 10.0.1.11 ← DOWN! DN 10.0.1.12 ← DOWN! DN 10.0.1.13 ← DOWN! DN 10.0.1.14 ← DOWN! DN 10.0.1.15 ← DOWN! -- ENTIRE CLUSTER OFFLINE! 💥 -- Application errors: Exception: NoHostAvailable All connection attempts failed! ↑ 15,000 users see this!

The Damage:

  • 💥 8:03 PM: All 6 nodes down, cluster unreachable
  • 📱 8:05 PM: PagerDuty explodes (500+ alerts)
  • 😡 8:07 PM: Customer complaints flood support
  • 🔄 8:10 PM: Nodes start coming back up slowly
  • ⚠️ 8:30 PM: Cluster still not forming (bootstrap issues)
  • 📞 9:00 PM: Tom escalates to senior DBA
  • 💰 10:00 PM: Cluster finally stable after 2 hours

Final Cost:

  • 💸 $500K: Lost revenue (2 hours downtime)
  • 😠 2,000: Angry customers
  • 📉 15%: Churn spike
  • 💼 Tom: Demoted to junior

✅ The Right Way (Rolling Restart)

What Tom Should Have Done:

-- Restart ONE node at a time: # 1. Restart Node 1 $ ssh node1 $ sudo systemctl restart cassandra $ nodetool status ← Wait until UN # Status after Node 1 restart: UN 10.0.1.10 500 GB ← Node 1 (restarted) ✅ UN 10.0.1.11 500 GB ← Node 2 (still running) UN 10.0.1.12 500 GB ← Node 3 (still running) UN 10.0.1.13 500 GB ← Node 4 (still running) UN 10.0.1.14 500 GB ← Node 5 (still running) UN 10.0.1.15 500 GB ← Node 6 (still running) -- RF=3 means data is on 3 nodes -- Even with 1 node down, still have 2 copies! -- USERS EXPERIENCE ZERO DOWNTIME! ✅ # 2. Wait 2 minutes, then restart Node 2 # 3. Wait 2 minutes, then restart Node 3 # ... Continue for all 6 nodes -- Total time: ~15 minutes -- User downtime: ZERO! ✅ -- Revenue lost: $0! ✅ -- Tom's job: SAFE! ✅

The Result:

  • ✅ 15 minutes: All nodes updated
  • ✅ Zero downtime: Users didn't notice
  • ✅ $0 lost: No impact on revenue
  • ✅ 0 complaints: Perfect deployment
  • ✅ Tom promoted: "Excellent ops skills!"

Rolling restart: The difference between disaster and success! 🎯

🎯 What Is a Rolling Restart?

Restart nodes one-by-one to maintain availability!

❌

Full Restart (BAD)

Restart all nodes at once

What happens:

  • 💥 Entire cluster goes down
  • 💥 All data unavailable
  • 💥 Users get errors
  • 💥 Revenue stops
  • 💥 SLAs violated

Downtime: 5-30 minutes

NEVER DO THIS IN PRODUCTION!

✅

Rolling Restart (GOOD)

Restart one node at a time

What happens:

  • ✅ Only 1 node down at a time
  • ✅ RF ensures data available
  • ✅ Users don't notice
  • ✅ Zero revenue impact
  • ✅ SLAs maintained

Downtime: ZERO!

ALWAYS USE IN PRODUCTION!

How Rolling Restart Works

The secret: Replication Factor (RF)!

-- With RF=3, every piece of data is on 3 nodes: Data for key "user_123": Node 1: Has copy ← If we restart this one... Node 3: Has copy ← This still serves! Node 5: Has copy ← And this! -- Restart Node 1: Node 1: DOWN ← Restarting Node 3: UP ✅ ← Still serving users! Node 5: UP ✅ ← Still serving users! Result: Users get data from Node 3 or Node 5! They don't even notice Node 1 is down! ✅

Key Concepts

1. Quorum Math

-- With RF=3, QUORUM = 2 nodes: QUORUM = (RF / 2) + 1 = (3 / 2) + 1 = 2 nodes -- Safe to restart 1 node at a time: Available nodes: 3 total - 1 restarting = 2 available QUORUM needs: 2 nodes 2 >= 2 ← QUORUM still satisfied! ✅ -- NEVER restart 2+ nodes at once with RF=3: Available nodes: 3 - 2 = 1 available QUORUM needs: 2 nodes 1 < 2 ← QUORUM FAILS! ❌

2. Wait Between Restarts

-- After restarting a node, WAIT until: 1. Node status is UN (Up Normal) $ nodetool status UN 10.0.1.10 ← Good! ✅ 2. Node finished catching up $ nodetool tpstats # All pools should have pending = 0 3. Wait 2-5 minutes for stability -- THEN restart next node

3. Order Matters

-- Best practice: Restart in rack order Rack 1: Node 1, Node 2 Rack 2: Node 3, Node 4 Rack 3: Node 5, Node 6 -- Restart order: 1. Node 1 (Rack 1) 2. Node 3 (Rack 2) ← Different rack! 3. Node 5 (Rack 3) ← Different rack! 4. Node 2 (Rack 1) 5. Node 4 (Rack 2) 6. Node 6 (Rack 3) -- Why? Spread risk across racks!

📋 When Do You Need a Rolling Restart?

Common scenarios!

⚙️

Config Changes

  • cassandra.yaml updates
  • JVM settings (jvm.options)
  • Timeout adjustments
  • Cache sizes
  • Compaction settings

Most common reason!

🔒

Security Updates

  • Enable authentication
  • Enable encryption
  • SSL certificate update
  • Audit logging enable
  • Password changes

Critical for security!

☕

JVM Tuning

  • Heap size changes
  • GC tuning
  • JVM flags
  • Memory settings
  • Thread pool config

Performance tuning!

🐛

Bug Fixes

  • Apply hotfixes
  • Driver updates
  • Plugin updates
  • Clear corrupted state
  • Force clean restart

Emergency fixes!

💾

OS/Hardware

  • OS kernel updates
  • Disk driver updates
  • Network driver updates
  • Memory upgrades
  • After reboot

Infrastructure changes!

📊

Performance

  • Clear caches
  • Reset metrics
  • Apply optimizations
  • Reload config
  • Fresh start

Optimization!

What DOESN'T Need a Restart

These changes take effect immediately (no restart):

  • ✅ Schema changes: CREATE/ALTER/DROP TABLE
  • ✅ User management: CREATE/ALTER ROLE, GRANT/REVOKE
  • ✅ Compaction: ALTER TABLE compaction
  • ✅ TTL/GC grace: ALTER TABLE gc_grace_seconds
  • ✅ Dynamic config: Some settings via nodetool

Always check documentation before restarting!

🔄 The Rolling Restart Process

Step-by-step with commands!

1

Pre-Flight Check

Verify cluster health BEFORE starting

-- 1. Check cluster status: $ nodetool status /* All nodes should be UN (Up Normal): UN 10.0.1.10 500 GB 16 33.3% abc123 UN 10.0.1.11 500 GB 16 33.3% def456 UN 10.0.1.12 500 GB 16 33.4% ghi789 */ -- 2. Check for ongoing repairs: $ nodetool compactionstats # Should show no active repairs -- 3. Check for pending compactions: $ nodetool compactionstats pending tasks: 0 ← Should be low/zero -- 4. Verify RF >= 3 (CRITICAL!): $ cqlsh -e "DESCRIBE KEYSPACE my_keyspace" replication = {'class': 'NetworkTopologyStrategy', 'datacenter1': '3'} ← RF=3 ✅ -- If RF < 2, DO NOT PROCEED! -- 5. Note the order: $ nodetool status | grep UN | awk '{print $2}' # Save this list - restart in this order
2

Update Config Files

Make changes on ALL nodes BEFORE restarting

-- Update cassandra.yaml on ALL nodes: $ ansible cassandra-cluster -m lineinfile -a " path=/etc/cassandra/cassandra.yaml regexp='^read_request_timeout' line='read_request_timeout: 10000ms' " -- Or manually on each node: $ ssh node1 $ sudo vim /etc/cassandra/cassandra.yaml # Change: read_request_timeout: 10000ms $ exit -- Repeat for node2, node3, etc. -- Verify changes on all nodes: $ ansible cassandra-cluster -m shell -a " grep read_request_timeout /etc/cassandra/cassandra.yaml " -- IMPORTANT: Config updated, but NOT restarted yet!
3

Restart First Node

Start the rolling restart

-- SSH to first node: $ ssh node1 -- Check it's UP before restart: $ nodetool status | grep 127.0.0.1 UN 127.0.0.1 500 GB ← Currently UP -- Restart Cassandra: $ sudo systemctl restart cassandra -- Watch logs for startup: $ tail -f /var/log/cassandra/system.log /* Look for: INFO Starting listening for CQL clients... INFO Node /10.0.1.10 state jump to NORMAL */ -- Check status (from another node): $ ssh node2 $ nodetool status /* Node1 should go DN briefly, then back to UN: DN 10.0.1.10 500 GB ← Down briefly ... UN 10.0.1.10 500 GB ← Back up! ✅ */
4

Wait and Verify

CRITICAL: Don't rush to next node!

-- Wait for node to be fully UP: $ watch -n 5 'nodetool status | grep 10.0.1.10' /* Wait until: UN 10.0.1.10 500 GB 16 33.3% abc123 ^^ UP and NORMAL */ -- Check thread pools are idle: $ nodetool tpstats /* Verify: Pool Name Active Pending ReadStage 0 0 ← Should be 0 MutationStage 0 0 ← Should be 0 */ -- Check gossip is stable: $ nodetool gossipinfo | grep "STATUS:" # Should show NORMAL for all nodes -- WAIT 2-5 MINUTES for stabilization $ sleep 120 ← 2 minute safety buffer -- NOW ready for next node!
5

Repeat for All Nodes

One by one, with patience

-- For each remaining node: # Node 2: $ ssh node2 $ sudo systemctl restart cassandra $ # Wait for UN + 2 minutes # Node 3: $ ssh node3 $ sudo systemctl restart cassandra $ # Wait for UN + 2 minutes # Node 4: $ ssh node4 $ sudo systemctl restart cassandra $ # Wait for UN + 2 minutes # ... Continue for all nodes -- Timeline for 6-node cluster: # ~2 min restart + 2 min wait = 4 min/node # 6 nodes × 4 min = ~24 minutes total # But ZERO downtime! ✅
6

Post-Restart Verification

Confirm success!

-- 1. Check all nodes UP: $ nodetool status /* Should see: UN 10.0.1.10 500 GB 16 33.3% abc123 ✅ UN 10.0.1.11 500 GB 16 33.3% def456 ✅ UN 10.0.1.12 500 GB 16 33.4% ghi789 ✅ */ -- 2. Verify config applied (check on one node): $ grep read_request_timeout /etc/cassandra/cassandra.yaml read_request_timeout: 10000ms ← New value! ✅ -- 3. Test query: $ cqlsh -e "SELECT * FROM my_keyspace.users LIMIT 1" # Should work! ✅ -- 4. Check application logs: # Should show NO errors during restart -- 5. Check metrics: # Request rate, latency should be normal -- Success! Rolling restart complete! 🎉

🚀 Advanced Techniques

Pro-level strategies!

1. Rack-Aware Rolling Restart

-- Cluster topology: Rack 1: Node 1, Node 2 Rack 2: Node 3, Node 4 Rack 3: Node 5, Node 6 -- Smart restart order (spread across racks): 1. Node 1 (Rack 1) 2. Node 3 (Rack 2) ← Different rack 3. Node 5 (Rack 3) ← Different rack 4. Node 2 (Rack 1) ← Back to Rack 1 5. Node 4 (Rack 2) 6. Node 6 (Rack 3) -- Why? If rack has issue, other racks still have replicas! -- Script to restart by rack: for rack in 1 2 3; do for node in $(nodetool status | grep "rack$rack" | awk '{print $2}'); do echo "Restarting $node in rack $rack" ssh $node "sudo systemctl restart cassandra" sleep 180 # 3 minutes done done

2. Multi-DC Rolling Restart

-- Two datacenters: DC1: Node 1, Node 2, Node 3 DC2: Node 4, Node 5, Node 6 -- CRITICAL: Complete one DC before starting another! # Phase 1: Restart entire DC1 1. Restart Node 1 (DC1) 2. Wait + verify 3. Restart Node 2 (DC1) 4. Wait + verify 5. Restart Node 3 (DC1) 6. Wait + verify 7. **WAIT 10 MINUTES** # Phase 2: Restart entire DC2 8. Restart Node 4 (DC2) 9. Wait + verify ... -- Why wait between DCs? # Allows DC1 to fully stabilize # Cross-DC replication to catch up # Reduces risk if something goes wrong

3. Canary Restart

-- Test config on ONE node first (canary): # Step 1: Update config ONLY on Node 1 $ ssh node1 $ sudo vim /etc/cassandra/cassandra.yaml # Make changes # Step 2: Restart Node 1 (the canary) $ sudo systemctl restart cassandra # Step 3: Monitor for 30 minutes # - Check logs for errors # - Monitor performance # - Watch for issues # Step 4: If canary is healthy, proceed with others if canary is healthy; then # Update config on all remaining nodes # Do rolling restart of remaining nodes else # Rollback Node 1 # Fix the issue fi -- Benefit: Catches bad configs before full rollout!

4. Parallel Rolling Restart

ADVANCED: Only with RF >= 5

Restart 2 nodes at once (if RF >= 5):

-- With RF=5, QUORUM=3 QUORUM = (5 / 2) + 1 = 3 nodes -- Can restart 2 nodes simultaneously: Available: 5 - 2 = 3 nodes 3 >= 3 QUORUM ← Still safe! -- Restart pairs from different racks: 1. Restart Node 1 (Rack 1) + Node 4 (Rack 2) 2. Wait for both UN 3. Restart Node 2 (Rack 1) + Node 5 (Rack 2) 4. Wait for both UN 5. Restart Node 3 -- Benefit: 2x faster! (~12 min instead of ~24 min) -- WARNING: DON'T DO THIS WITH RF=3! # RF=3, 2 down = only 1 available # QUORUM=2 but only 1 available = FAIL! ❌

🤖 Automation Scripts

Automate the process safely!

Bash Script: rolling_restart.sh

#!/bin/bash # rolling_restart.sh - Safe rolling restart set -e # Exit on error # Configuration NODES=("10.0.1.10" "10.0.1.11" "10.0.1.12") WAIT_SECONDS=180 # 3 minutes between nodes MAX_WAIT=600 # 10 minutes max wait for UN echo "🔄 Starting rolling restart of ${#NODES[@]} nodes" # Pre-flight check echo "✅ Pre-flight check..." for node in "${NODES[@]}"; do status=$(ssh $node "nodetool status | grep 127.0.0.1 | awk '{print \$1}'") if [ "$status" != "UN" ]; then echo "❌ Node $node is not UN! Status: $status" exit 1 fi echo " ✅ $node is UP" done # Rolling restart for i in "${!NODES[@]}"; do node="${NODES[$i]}" num=$((i+1)) echo "" echo "🔄 [$num/${#NODES[@]}] Restarting $node..." # Restart Cassandra ssh $node "sudo systemctl restart cassandra" echo " ⏳ Restarted, waiting for UP..." # Wait for node to be UN waited=0 while [ $waited -lt $MAX_WAIT ]; do status=$(ssh $node "nodetool status | grep 127.0.0.1 | awk '{print \$1}'" 2>/dev/null || echo "DN") if [ "$status" == "UN" ]; then echo " ✅ $node is UP after ${waited}s" break fi sleep 10 waited=$((waited+10)) echo -n "." done echo "" if [ $waited -ge $MAX_WAIT ]; then echo "❌ Timeout waiting for $node to be UN!" exit 1 fi # Wait for stability if [ $num -lt ${#NODES[@]} ]; then echo " ⏳ Waiting ${WAIT_SECONDS}s for stability..." sleep $WAIT_SECONDS fi done echo "" echo "🎉 Rolling restart complete!" echo "✅ All nodes restarted successfully" # Final verification echo "" echo "📊 Final status:" nodetool status

Usage

# Make executable: $ chmod +x rolling_restart.sh # Run: $ ./rolling_restart.sh /* Output: 🔄 Starting rolling restart of 3 nodes ✅ Pre-flight check... ✅ 10.0.1.10 is UP ✅ 10.0.1.11 is UP ✅ 10.0.1.12 is UP 🔄 [1/3] Restarting 10.0.1.10... ⏳ Restarted, waiting for UP... ✅ 10.0.1.10 is UP after 45s ⏳ Waiting 180s for stability... 🔄 [2/3] Restarting 10.0.1.11... ⏳ Restarted, waiting for UP... ✅ 10.0.1.11 is UP after 42s ⏳ Waiting 180s for stability... 🔄 [3/3] Restarting 10.0.1.12... ⏳ Restarted, waiting for UP... ✅ 10.0.1.12 is UP after 48s 🎉 Rolling restart complete! ✅ All nodes restarted successfully */

🔧 Troubleshooting Rolling Restarts

Fix common issues!

❌ Node Won't Start

-- SYMPTOM: Node stays DN after restart -- DIAGNOSIS: $ ssh problem-node $ tail -100 /var/log/cassandra/system.log -- COMMON CAUSES: # 1. Config syntax error: ERROR: Cannot start node. Invalid yaml: ... # FIX: Rollback config: $ sudo cp /etc/cassandra/cassandra.yaml.bak \ /etc/cassandra/cassandra.yaml $ sudo systemctl start cassandra # 2. Port already in use: ERROR: Address already in use: /0.0.0.0:9042 # FIX: Kill old process: $ ps aux | grep cassandra $ sudo kill -9 $ sudo systemctl start cassandra # 3. Disk full: ERROR: Cannot create commit log # FIX: Free up space: $ df -h /var/lib/cassandra $ nodetool clearsnapshot $ sudo systemctl start cassandra

❌ Cluster Loses Quorum

-- SYMPTOM: Queries failing with "Unavailable" -- CAUSE: Restarted too many nodes at once -- Example with RF=3: Restarted Node 1 and Node 2 together ← BAD! Available nodes: 3 - 2 = 1 QUORUM needs: 2 1 < 2 ← QUORUM LOST! 💥 -- FIX: WAIT for nodes to come back up! $ watch -n 5 'nodetool status' # Don't restart any more nodes until: # - Both nodes are UN # - Cluster is stable # - Then resume one-at-a-time -- PREVENTION: NEVER restart 2+ nodes with RF=3!

❌ Node Stuck in Joining

-- SYMPTOM: $ nodetool status UJ 10.0.1.10 ← Stuck in "Joining" state -- CAUSE: Node thinks it's new (bootstrap issue) -- CHECK: Is auto_bootstrap enabled? $ grep auto_bootstrap /etc/cassandra/cassandra.yaml # auto_bootstrap: true ← This is the problem! -- FIX: Disable auto_bootstrap for restart: $ sudo systemctl stop cassandra $ sudo vim /etc/cassandra/cassandra.yaml # Add/set: auto_bootstrap: false $ sudo systemctl start cassandra # Node should now join as NORMAL -- PREVENTION: Always set auto_bootstrap: false # for rolling restarts!

❌ Increased Latency During Restart

-- SYMPTOM: P99 latency spikes during restart -- CAUSE: Restarted too fast, not enough wait time -- FIX: Increase wait time between nodes: WAIT_SECONDS=300 # 5 minutes instead of 3 -- Also check: # 1. Is node finished bootstrapping? $ nodetool netstats Mode: NORMAL ← Should say NORMAL, not JOINING # 2. Are compactions caught up? $ nodetool compactionstats pending tasks: 0 ← Should be low/zero # 3. Is GC stable? $ nodetool gcstats # Should not show long pauses -- PREVENTION: Wait longer between restarts!

💡 Rolling Restart Best Practices

Do it right every time!

✅

DO

  • Check cluster health first
  • Update config on ALL nodes
  • Restart ONE node at a time
  • Wait for UN before next
  • Wait 2-5 minutes between
  • Test on canary first
  • Schedule during low traffic
  • Monitor the entire time
❌

DON'T

  • Restart all nodes at once
  • Restart 2+ with RF=3
  • Skip pre-flight checks
  • Rush between nodes
  • Do during peak traffic
  • Forget to monitor
  • Leave auto_bootstrap on
  • Skip post-verification

Rolling Restart Checklist

Complete this checklist for every rolling restart:

Phase Task Done?
Pre-Flight ☐ All nodes UN
☐ RF >= 3 verified
☐ No active repairs
☐ Low pending compactions
Config ☐ Updated on ALL nodes
☐ Syntax validated
☐ auto_bootstrap: false
Restart ☐ One node at a time
☐ Wait for UN each time
☐ 2-5 min wait between
Post ☐ All nodes UN
☐ Config applied
☐ Test query works
☐ No application errors

Timeline Expectations

Cluster Size Time Per Node Total Time
3 nodes ~4 minutes ~12 minutes
6 nodes ~4 minutes ~24 minutes
12 nodes ~4 minutes ~48 minutes
24 nodes ~4 minutes ~96 minutes (1.6 hrs)

Note: ~2 min restart + 2 min wait = 4 min per node

🎉 Master Rolling Restarts!

You now know how to update Cassandra with ZERO downtime!

🎓 What You Learned:

  • 🎯 What: Restart one node at a time
  • ❌ Tom's disaster: Restarted all at once = $500K lost
  • ✅ The right way: One at a time = zero downtime
  • 🔄 The process: 6 steps from pre-flight to verification
  • 🚀 Advanced: Rack-aware, multi-DC, canary
  • 🤖 Automation: rolling_restart.sh script
  • 🔧 Troubleshooting: Fix common issues

💡 Key Takeaways:

  1. ONE at a time - Never restart 2+ with RF=3
  2. WAIT between restarts - 2-5 minutes for stability
  3. RF >= 3 required - Or you'll lose quorum
  4. Monitor constantly - Watch status, logs, metrics
  5. Use automation - Scripts prevent human error
  6. Verify after - Test queries, check config applied

📋 The Golden Rule:

🚀 With RF=3: Restart ONE at a time, WAIT for UN, repeat
Result: ZERO downtime, happy users, safe job! ✅

🔄 Quick Reference:

# For each node: 1. ssh node 2. sudo systemctl restart cassandra 3. Watch: nodetool status (wait for UN) 4. Sleep 120 (2 minute stability) 5. Repeat for next node # Never: ❌ Restart all at once ❌ Restart 2+ with RF=3 ❌ Skip the wait time # Always: ✅ One at a time ✅ Wait for UN ✅ Monitor everything

🔁 Remember: Rolling restart = Zero downtime magic! 🎯
Don't be Tom - do it right!

Advertisement

📱 Responsive Ad 📱