Stay Current & Secure

Upgrading Cassandra

Safely upgrade to newer versions with zero downtime!

📖 The Story: Maria's Upgrade Disaster

Maria ran Cassandra 2.1 for 3 years. Security team finally demanded upgrade. She jumped straight to 4.0 without testing. Incompatible SSTable format. Cluster wouldn't start. 8-hour outage. Lost $1M in revenue. All because she skipped major → major upgrade rules.

😱 What Maria Did Wrong

The Fatal Jump:

-- Maria's cluster (6 nodes running 2.1.20): $ nodetool version ReleaseVersion: 2.1.20 -- Security audit: "You MUST upgrade! 2.1 has vulnerabilities!" -- Maria: "Let's jump to latest! 4.0 is out!" -- Maria's fatal mistake: $ sudo apt install cassandra=4.0.0 $ sudo systemctl restart cassandra -- Node tries to start... ERROR: Cannot read SSTable format version 'ja' ERROR: SSTable created by incompatible version FATAL: Unable to start Cassandra -- Node won't start! 💥

The Cascade:

  • 💥 Hour 1: Upgraded all 6 nodes to 4.0
  • 💥 Hour 1:15: None would start (SSTable incompatible)
  • 😱 Hour 2: Panic - entire cluster down!
  • 📞 Hour 3: Called DataStax support ($$$)
  • ⬇️ Hour 4-6: Downgrading all nodes back to 2.1
  • 🔄 Hour 7: Running `nodetool upgradesstables`
  • ✅ Hour 8: Finally back online

Total Damage:

  • 💰 $1M: Lost revenue (8 hours)
  • 💸 $50K: Emergency support fees
  • 😡 5,000: Angry customers
  • 📰 Press: "Major Outage Due to Failed Upgrade"
  • 💼 Maria: Put on probation

✅ The Right Way (Incremental Upgrades)

What Maria Should Have Done:

-- Rule: Can only upgrade ONE MAJOR version at a time! -- Starting point: Version: 2.1.20 -- Correct upgrade path: 2.1.20 → 2.2.19 → 3.0.28 → 3.11.14 → 4.0.0 ↓ ↓ ↓ ↓ Patch Patch Patch Finally! -- Process for EACH upgrade: 1. Test in staging environment 2. Take full snapshot 3. Upgrade ONE node 4. Verify it works 5. Rolling upgrade remaining nodes 6. Run upgradesstables 7. Wait 24 hours for stability 8. THEN start next upgrade -- Timeline: # 4 major upgrades × 2 days each = 8 days total # BUT: Zero downtime! ✅

The Safe Result:

  • ✅ 8 days: Incremental upgrades
  • ✅ Zero downtime: Rolling upgrades
  • ✅ $0 lost: No revenue impact
  • ✅ Tested: Each step validated
  • ✅ Maria promoted: "Excellent planning!"

Patience = Success: One major version at a time! 🎯

🎯 Why Upgrade Cassandra?

The business case for staying current!

🔒

Security Patches

  • Critical vulnerabilities fixed
  • CVE patches
  • Prevent exploits
  • Compliance requirements
  • Avoid breaches

Most important reason!

🐛

Bug Fixes

  • Data corruption fixes
  • Stability improvements
  • Memory leak fixes
  • GC improvements
  • Crash prevention

Reliability!

🚀

Performance

  • Faster queries
  • Better compaction
  • Improved caching
  • Reduced latency
  • Higher throughput

Speed boost!

✨

New Features

  • Better data types
  • Improved CQL
  • New compaction strategies
  • Enhanced security
  • Better monitoring

Modern capabilities!

📚

Support

  • Old versions EOL
  • No security updates
  • Limited community help
  • Driver incompatibility
  • Integration issues

Stay supported!

⚖️

Compliance

  • PCI-DSS requirements
  • HIPAA compliance
  • SOX audits
  • Security policies
  • Vendor requirements

Legal requirements!

Cost of NOT Upgrading

Risk Example Cost
Security Breach Known CVE exploited $5M+ fine, reputation damage
Data Corruption Known bug corrupts data Lost data, customer churn
No Support Critical issue, EOL version Extended downtime, no help
Failed Audit Compliance requires latest Operations halted

🛣️ Cassandra Upgrade Paths

Know your route!

The Golden Rule

⚠️ You can ONLY upgrade ONE MAJOR version at a time! ⚠️

-- CORRECT: Incremental upgrades 2.1 → 2.2 → 3.0 → 3.11 → 4.0 ✅ -- WRONG: Skip major versions 2.1 → 4.0 ❌ WILL BREAK! 3.0 → 4.1 ❌ WILL BREAK!

Version Upgrade Matrix

From Version To Version Direct? Path Required
2.1.x 2.2.x ✅ Yes Direct upgrade
2.1.x 3.0.x ❌ No 2.1 → 2.2 → 3.0
2.2.x 3.0.x ✅ Yes Direct upgrade
3.0.x 3.11.x ✅ Yes Direct upgrade
3.0.x 4.0.x ❌ No 3.0 → 3.11 → 4.0
3.11.x 4.0.x ✅ Yes Direct upgrade
4.0.x 4.1.x ✅ Yes Direct upgrade
4.1.x 5.0.x ✅ Yes Direct upgrade

Common Upgrade Paths

Path 1: Legacy (2.1) to Modern (4.0)

-- Complete path (4 upgrades): 2.1.20 → 2.2.19 → 3.0.28 → 3.11.14 → 4.0.11 -- Timeline: ~8-10 days total # Each upgrade: 1-2 days # Validation between: 1 day -- Why can't skip? # SSTable format changes # Protocol changes # Schema changes

Path 2: Recent (3.11) to Latest (4.1)

-- Shorter path (2 upgrades): 3.11.14 → 4.0.11 → 4.1.3 -- Timeline: ~4-5 days total -- Note: 3.11 is LTS (Long Term Support) # Safe to stay on 3.11 if not ready for 4.x

Path 3: Patch Upgrades (Easy)

-- Within same major version: 4.0.7 → 4.0.11 ✅ Easy, low risk 4.1.0 → 4.1.3 ✅ Easy, low risk -- Timeline: 1-2 days -- No SSTable format changes -- Just bug fixes & security patches

📋 Pre-Upgrade Preparation

Critical steps BEFORE upgrading!

1

Read Release Notes

Know what's changing!

-- Read EVERYTHING between your version and target: # Upgrading 3.11 → 4.0? Read: - Cassandra 4.0 Release Notes - Cassandra 4.0 NEWS.txt - Cassandra 4.0 CHANGES.txt -- Look for: ✅ Breaking changes ✅ Deprecated features ✅ New requirements ✅ Config changes ✅ Known issues -- Example breaking changes: # 4.0 removed Thrift # 4.0 changed default compaction # 4.0 requires Java 11
2

Take Full Snapshot

Backup EVERYTHING!

-- Take snapshot on ALL nodes: $ nodetool snapshot --tag pre-upgrade-4.0 -- Verify snapshot created: $ nodetool listsnapshots /* Output: Snapshot name Keyspace Table Size pre-upgrade-4.0 my_keyspace users 50 GB pre-upgrade-4.0 my_keyspace orders 120 GB */ -- Copy snapshots to external storage: $ rsync -avz /var/lib/cassandra/data/ \ backup-server:/cassandra-backups/pre-upgrade/ -- CRITICAL: Test restore BEFORE upgrading!
3

Test in Staging

NEVER upgrade production first!

-- Build staging cluster (same version as prod): Staging: 3.11.14 (matching production) -- Copy production data to staging: $ nodetool snapshot --tag staging-data $ rsync -avz prod:/var/lib/cassandra/data/ \ staging:/var/lib/cassandra/data/ -- Perform upgrade on staging: 1. Upgrade one staging node 2. Run upgradesstables 3. Test application thoroughly 4. Upgrade remaining staging nodes 5. Test for 24-48 hours -- Test checklist: ✅ All queries work ✅ Application connects ✅ Performance acceptable ✅ No errors in logs ✅ Monitoring working -- If staging fails, FIX BEFORE production!
4

Check Compatibility

Verify all components compatible

-- Check driver versions: # Upgrading to 4.0? Need driver that supports 4.0! Python driver: 3.25+ supports 4.0 Java driver: 4.13+ supports 4.0 Node.js driver: 4.6+ supports 4.0 -- Check Java version: $ java -version # Cassandra 4.0 requires Java 11 # Cassandra 3.x works with Java 8 -- Check monitoring tools: # Does Prometheus exporter work with 4.0? # Does Grafana dashboard need updates? -- Check integrations: # Spark connector compatible? # Kafka connector compatible?

Pre-Upgrade Checklist

  • ☐ Read release notes (all versions between current and target)
  • ☐ Take snapshots (all nodes, copy offsite)
  • ☐ Test restore (verify backups work!)
  • ☐ Test in staging (at least 24 hours)
  • ☐ Check drivers (compatible versions)
  • ☐ Check Java version (meets requirements)
  • ☐ Schedule maintenance window (even though rolling upgrade)
  • ☐ Notify stakeholders (let them know upgrade happening)
  • ☐ Prepare rollback plan (how to go back if fails)
  • ☐ Check cluster health (all nodes UN, no repairs running)

⬆️ The Upgrade Process

Step-by-step rolling upgrade!

1

Upgrade First Node (Canary)

Test on one node first

-- Choose a canary node (not seed node!): $ ssh node3 ← Non-seed node -- 1. Take snapshot: $ nodetool snapshot --tag pre-upgrade-node3 -- 2. Drain the node: $ nodetool drain # Flushes memtables, stops accepting writes -- 3. Stop Cassandra: $ sudo systemctl stop cassandra -- 4. Install new version: $ sudo apt update $ sudo apt install cassandra=4.0.11 -- 5. Start Cassandra: $ sudo systemctl start cassandra -- 6. Watch logs for startup: $ tail -f /var/log/cassandra/system.log /* Look for: INFO Starting listening for CQL clients... INFO Node state jump to NORMAL */ -- 7. Verify status: $ nodetool version ReleaseVersion: 4.0.11 ← New version! ✅ $ nodetool status UN 10.0.1.13 500 GB ← Node3 is UP! ✅
2

Run upgradesstables

Rewrite SSTables to new format

-- IMPORTANT: Must run after EACH node upgrade! $ nodetool upgradesstables /* What this does: - Rewrites SSTables in new format - Can take 30 min - 6 hours - Increases disk I/O temporarily - Necessary for new features */ -- Monitor progress: $ nodetool compactionstats /* Output: pending tasks: 15 compaction type keyspace table Upgrade my_keyspace users */ -- Wait until complete: pending tasks: 0 ← Finished! ✅
3

Monitor Canary (24 Hours)

Watch for issues before continuing

-- Monitor canary node for 24 hours: # Check logs continuously: $ tail -f /var/log/cassandra/system.log | grep -i error # Check metrics: $ nodetool tablestats my_keyspace.users # Verify latency is normal # Check GC: $ nodetool gcstats # Should not show increase in GC # Test queries: $ cqlsh node3 cqlsh> SELECT * FROM my_keyspace.users LIMIT 10; # Should work! ✅ -- If ANY issues, rollback immediately: $ sudo systemctl stop cassandra $ sudo apt install cassandra=3.11.14 $ sudo systemctl start cassandra -- If canary healthy for 24h, proceed!
4

Rolling Upgrade Remaining Nodes

One at a time, with patience

-- For EACH remaining node: # Node 1: $ ssh node1 $ nodetool snapshot --tag pre-upgrade-node1 $ nodetool drain $ sudo systemctl stop cassandra $ sudo apt install cassandra=4.0.11 $ sudo systemctl start cassandra # Wait for UN $ nodetool upgradesstables # Wait 2-5 minutes for stability # Node 2: # ... Repeat same process # Continue for all nodes -- Important: ✅ ONE node at a time ✅ Wait for UN after each ✅ Run upgradesstables on each ✅ 2-5 min wait between nodes -- Timeline for 6-node cluster: # ~6 hours total # But ZERO downtime! ✅
5

Post-Upgrade Verification

Confirm success!

-- 1. Check all nodes upgraded: $ ansible cassandra-cluster -m shell -a "nodetool version" /* All should show: ReleaseVersion: 4.0.11 */ -- 2. Check cluster health: $ nodetool status /* All should be UN: UN 10.0.1.10 500 GB ✅ UN 10.0.1.11 500 GB ✅ UN 10.0.1.12 500 GB ✅ */ -- 3. Verify upgradesstables completed everywhere: $ ansible cassandra-cluster -m shell -a "nodetool compactionstats" # All should show: pending tasks: 0 -- 4. Test queries: $ cqlsh -e "SELECT * FROM my_keyspace.users LIMIT 1" # Should work! ✅ -- 5. Check application: # Verify app still working # Check app logs for errors # Monitor metrics -- 6. Run repair (recommended): $ nodetool repair -pr -- 7. Clean up snapshots (after 1 week): $ nodetool clearsnapshot pre-upgrade-4.0 -- Success! Upgrade complete! 🎉

🔄 Major vs Minor Upgrades

Understand the difference!

⚠️

Major Upgrade

3.11 → 4.0 (First digit change)

Changes:

  • SSTable format changes
  • New features
  • Breaking changes
  • Protocol changes
  • Config changes

Risk: HIGH

Time: 1-2 days

Testing: MANDATORY

Plan carefully!

✅

Minor/Patch Upgrade

4.0.7 → 4.0.11 (Patch version)

Changes:

  • Bug fixes
  • Security patches
  • Performance improvements
  • No breaking changes
  • No format changes

Risk: LOW

Time: 4-6 hours

Testing: Recommended

Safer, do often!

Upgrade Complexity Comparison

Aspect Major (3.11 → 4.0) Patch (4.0.7 → 4.0.11)
upgradesstables ✅ Required (hours) ❌ Not needed
Staging test ✅ Mandatory ⚠️ Recommended
Rollback ⚠️ Complex ✅ Easy
Driver update ✅ Often required ❌ Usually not needed
Downtime Zero (rolling) Zero (rolling)
Total time 1-2 days 4-6 hours

🔧 Troubleshooting Upgrades

Fix common issues!

❌ Node Won't Start After Upgrade

-- SYMPTOM: $ sudo systemctl status cassandra Failed to start cassandra -- CHECK LOGS: $ tail -100 /var/log/cassandra/system.log -- COMMON CAUSES: # 1. Wrong Java version: ERROR: Java 8 not supported, need Java 11+ # FIX: Install Java 11 $ sudo apt install openjdk-11-jdk $ sudo update-alternatives --config java # 2. Config incompatibility: ERROR: Unknown configuration parameter 'thrift_framed_transport_size_in_mb' # FIX: Remove deprecated config $ sudo vim /etc/cassandra/cassandra.yaml # Comment out or remove deprecated settings # 3. SSTable format error: ERROR: Cannot read SSTable version # FIX: Restore from snapshot + rollback $ nodetool refresh $ sudo systemctl restart cassandra

❌ Cluster Loses Quorum

-- SYMPTOM: Queries failing with "Unavailable" -- CAUSE: Upgraded too many nodes at once -- FIX: WAIT for nodes to come back online $ watch -n 5 'nodetool status' # Don't upgrade more until cluster stable! -- PREVENTION: ONE node at a time with RF=3

❌ Driver Incompatibility

-- SYMPTOM: Application can't connect after upgrade ERROR: Protocol version 4 not supported by server -- CAUSE: Driver too old for new Cassandra version -- FIX: Upgrade driver BEFORE upgrading Cassandra # Python example: pip install --upgrade cassandra-driver # Java example: com.datastax.oss java-driver-core 4.15.0 -- PREVENTION: Check driver compatibility BEFORE upgrade

❌ Performance Degradation

-- SYMPTOM: Latency increased after upgrade -- CHECK: Did you run upgradesstables? $ nodetool compactionstats pending tasks: 0 ← Should be 0 -- CHECK: Mixed versions in cluster? $ ansible cassandra-cluster -m shell -a "nodetool version" # All should match! -- CHECK: Compaction strategy changed? $ cqlsh -e "DESC TABLE my_keyspace.users" # 4.0 changed default compaction -- FIX: May need to tune new settings

Emergency Rollback

If upgrade goes catastrophically wrong:

-- 1. Stop Cassandra on problem node: $ sudo systemctl stop cassandra -- 2. Remove new version: $ sudo apt remove cassandra -- 3. Install old version: $ sudo apt install cassandra=3.11.14 -- 4. Restore snapshot if needed: $ nodetool refresh my_keyspace users -- 5. Start Cassandra: $ sudo systemctl start cassandra -- 6. Verify: $ nodetool version $ nodetool status -- IMPORTANT: Only rollback if SEVERE issues! # Prefer to fix forward when possible

💡 Upgrade Best Practices

Do it right!

✅

DO

  • Read ALL release notes
  • Test in staging first
  • Take full snapshots
  • One major version at a time
  • Upgrade one node at a time
  • Run upgradesstables
  • Wait 24h on canary
  • Monitor continuously
❌

DON'T

  • Skip major versions
  • Upgrade all at once
  • Skip staging tests
  • Forget snapshots
  • Skip upgradesstables
  • Ignore release notes
  • Upgrade during peak
  • Rush the process

Upgrade Checklist

Complete this for every major upgrade:

Phase Task Done?
Planning ☐ Read release notes
☐ Check compatibility
☐ Plan upgrade path
Preparation ☐ Take snapshots (all nodes)
☐ Copy backups offsite
☐ Test restore
Staging ☐ Build staging cluster
☐ Perform upgrade in staging
☐ Test for 24-48 hours
Production ☐ Upgrade canary node
☐ Run upgradesstables
☐ Monitor canary 24h
☐ Rolling upgrade remaining
Verification ☐ All nodes same version
☐ upgradesstables complete
☐ Application working
☐ No errors in logs
☐ Run repair

Upgrade Timeline Estimates

Upgrade Type Testing Production Total
Patch (4.0.7 → 4.0.11) 4-8 hours 4-6 hours 1 day
Minor (4.0 → 4.1) 1 day 6-8 hours 2 days
Major (3.11 → 4.0) 1-2 days 1 day 2-3 days
Multi-major (2.1 → 4.0) 2-3 days 4-5 days 8-10 days

🎉 Master Cassandra Upgrades!

You now know how to safely upgrade Cassandra!

🎓 What You Learned:

  • 📖 Maria's disaster: Jumped 2.1 → 4.0 = $1M lost
  • 🎯 Why upgrade: Security, bugs, performance, features
  • 🛣️ Upgrade paths: ONE major version at a time!
  • 📋 Preparation: Snapshots, staging, testing
  • ⬆️ Process: Canary → monitor → rolling upgrade
  • 🔄 Major vs minor: Major = risky, patch = safe
  • 🔧 Troubleshooting: Fix common upgrade issues
  • 💡 Best practices: Test, snapshot, be patient

💡 Key Takeaways:

  1. ONE major version at a time - 2.1 → 2.2 → 3.0 (not 2.1 → 3.0)
  2. Test in staging FIRST - Never upgrade production first
  3. Take snapshots - Backup everything before starting
  4. Canary approach - Test on one node, wait 24h
  5. Run upgradesstables - Required after major upgrades
  6. Be patient - Don't rush, zero downtime is guaranteed

⚠️ The Golden Rule:

🚫 NEVER skip major versions!
2.1 → 4.0 = DISASTER 💥
2.1 → 2.2 → 3.0 → 3.11 → 4.0 = SUCCESS ✅

📋 Quick Upgrade Process:

# For EACH major version upgrade: ## 1. PREPARATION (Day 1) Read release notes Take snapshots on all nodes Test in staging environment ## 2. CANARY (Day 2) Upgrade one non-seed node Run upgradesstables Monitor for 24 hours ## 3. ROLLING UPGRADE (Day 3) for each remaining node: nodetool drain Stop Cassandra Install new version Start Cassandra Wait for UN Run upgradesstables Wait 2-5 minutes ## 4. VERIFICATION All nodes same version All upgradesstables complete Application working No errors in logs ## 5. STABILIZE Monitor for 24-48 hours Run repair Clean old snapshots # Then start next major version upgrade!

🎯 Upgrade Decision Tree:

## Question: What am I upgrading? if patch upgrade (e.g., 4.0.7 → 4.0.11): Risk: LOW Time: 4-6 hours upgradesstables: NO Testing: Recommended → Safe, do it often! elif one major version (e.g., 3.11 → 4.0): Risk: MEDIUM Time: 1-2 days upgradesstables: YES (required) Testing: MANDATORY → Plan carefully, test thoroughly elif multiple major versions (e.g., 2.1 → 4.0): Risk: HIGH Time: 8-10 days upgradesstables: YES (each step) Testing: MANDATORY (each step) Path: 2.1 → 2.2 → 3.0 → 3.11 → 4.0 → Multiple upgrades, be patient else: → Read release notes first!

⚡ Pro Tips:

  • 🕐 Schedule during low traffic - Even though zero downtime
  • 📊 Monitor metrics - Watch latency, throughput, errors
  • 🔄 Upgrade drivers first - Before upgrading Cassandra
  • ☕ Check Java version - 4.0+ requires Java 11
  • 📸 Keep snapshots 1 week - In case of delayed issues
  • 🏃 Don't rush - Patience prevents disasters

🚨 When to Rollback:

  • ❌ Node won't start after multiple attempts
  • ❌ Data corruption detected
  • ❌ Critical application features broken
  • ❌ Severe performance degradation
  • ✅ Minor issues? Fix forward, don't rollback

📚 Recommended Upgrade Schedule:

Upgrade Type Frequency Reason
Patch Every 3-6 months Security fixes, bug fixes
Minor Every 6-12 months New features, improvements
Major Every 1-2 years Stay current, avoid EOL

✅ Success Metrics:

Your upgrade is successful when:

  • ✅ All nodes show same version
  • ✅ All nodes status = UN
  • ✅ upgradesstables completed on all nodes
  • ✅ Zero application errors
  • ✅ Latency within normal range
  • ✅ No errors in Cassandra logs
  • ✅ Repair completes successfully
  • ✅ Monitoring dashboards green
  • ✅ Users didn't notice anything

⬆️ Remember Maria's lesson:
Patience + Planning = Zero Downtime Success! 🎯

🎓 Final Wisdom

The difference between a $1M disaster and a smooth upgrade?
Reading the release notes.
Testing in staging.
Taking snapshots.
Following the rules.
Being patient.

Don't be Maria. Plan your upgrades! 🚀

Advertisement

📱 Responsive Ad 📱