Data Protection

Backup & Restore

Protect your data - Because disasters happen!

🚨 The $1.7 Million Lesson

GitLab's 2017 Disaster

January 31, 2017 - 300GB of production data accidentally deleted

What happened:

  • 🔴 Engineer removed wrong database directory
  • 🔴 6 hours of production data lost
  • 🔴 Backup 1: Failed (replication lag)
  • 🔴 Backup 2: Failed (accidentally disabled)
  • 🔴 Backup 3: Failed (never tested)
  • 🔴 Backup 4: Failed (wrong path)
  • 🔴 Backup 5: Worked! (LVM snapshots saved them)

💰 The Cost

  • 💸 $1.7M in lost revenue
  • 😰 18 hours of downtime
  • 📉 Brand damage and customer trust
  • ⚠️ Lesson: 4 out of 5 backup methods failed!

✅ Don't Be Like GitLab

  • 💾 Multiple backup methods
  • 🏗️ Separate storage locations
  • ✅ Test restores regularly
  • ⚙️ Automate everything
  • 📝 Document procedures
  • 🚨 Monitor backup success

💾 Backups aren't real until you've tested a restore!

📂 Types of Backups

Choose the right strategy!

📸

Snapshots

Point-in-time backups

  • ✅ Fast to create (seconds)
  • ✅ Minimal performance impact
  • ✅ Full table backup
  • ⚠️ Uses disk space
  • 💡 Best for: Daily backups
📦

Incremental

Only changed data

  • ✅ Space efficient
  • ✅ Automatic with commit logs
  • ✅ Point-in-time recovery
  • ⚠️ Harder to restore
  • 💡 Best for: Continuous backup
💾

Full Cluster

Complete cluster copy

  • ✅ Complete recovery
  • ✅ Easiest to restore
  • ⚠️ Slow to create
  • ⚠️ Large storage needs
  • 💡 Best for: Major migrations

Recommended Strategy: 3-2-1 Rule

  • 3 copies of your data (production + 2 backups)
  • 2 different media types (local disk + cloud)
  • 1 copy offsite (S3, different datacenter)

Example: Production data + Daily snapshots on local disk + Weekly backups to S3

📸 Creating Snapshots

Fast, reliable backups!

1

Flush Memtables First

Ensure all data is on disk

# Flush all keyspaces to disk nodetool flush # Flush specific keyspace nodetool flush my_keyspace # This ensures snapshot is complete!
2

Take Snapshot

# Snapshot all keyspaces with timestamp nodetool snapshot -t backup-$(date +%Y%m%d-%H%M%S) # Snapshot specific keyspace nodetool snapshot -t backup-$(date +%Y%m%d) my_keyspace # Snapshot specific table nodetool snapshot -t backup-$(date +%Y%m%d) my_keyspace.users # Snapshots stored in: # /var/lib/cassandra/data/keyspace/table/snapshots/backup-name/
3

Verify Snapshot

# List all snapshots nodetool listsnapshots # Output shows: # - Snapshot name # - Keyspace # - Table # - Size # - Creation time # Check snapshot directory ls -lh /var/lib/cassandra/data/my_keyspace/users-*/snapshots/

Important: Snapshots Take Disk Space!

Snapshots are hard links, but still use space:

  • ⚠️ Initial snapshot: Uses ~same space as table
  • ⚠️ Data updates: New SSTables created, old kept by snapshot
  • ⚠️ Can fill disk if not cleaned up!
  • ✅ Always monitor disk space
  • ✅ Clear old snapshots regularly

☁️ Storing Backups

Where to keep your backups safe!

📦

Copy to Archive Directory

# Create archive directory mkdir -p /backup/cassandra/$(date +%Y%m%d) # Copy snapshots for keyspace in /var/lib/cassandra/data/*; do for table in $keyspace/*; do snapshot_dir="$table/snapshots/backup-20250106" if [ -d "$snapshot_dir" ]; then cp -al $snapshot_dir/* /backup/cassandra/20250106/ fi done done # -a = preserve attributes # -l = hard links (save space)
☁️

Upload to S3 / Cloud Storage

# Compress snapshot tar -czf backup-20250106.tar.gz /backup/cassandra/20250106/ # Upload to S3 aws s3 cp backup-20250106.tar.gz \ s3://my-cassandra-backups/node1/backup-20250106.tar.gz # Or use rclone for any cloud provider rclone copy /backup/cassandra/20250106/ \ remote:cassandra-backups/node1/20250106/ # Or use Google Cloud Storage gsutil -m cp -r /backup/cassandra/20250106/ \ gs://my-cassandra-backups/node1/20250106/
🗑️

Clean Up Old Snapshots

# Clear specific snapshot nodetool clearsnapshot -t backup-20250106 # Clear all snapshots nodetool clearsnapshot # Automated cleanup (keep last 7 days) find /var/lib/cassandra/data -type d -name "snapshots" \ -mtime +7 -exec rm -rf {} \; # Clean up local archives (keep 30 days) find /backup/cassandra/ -type d -mtime +30 -delete
💰

Amazon S3

Most popular choice

  • ✅ Very cheap ($0.023/GB)
  • ✅ 99.999999999% durability
  • ✅ Easy to use (AWS CLI)
  • ✅ Lifecycle policies
☁️

Google Cloud Storage

Google's alternative

  • ✅ Comparable pricing
  • ✅ Good for GCP users
  • ✅ Fast transfers
  • ✅ Nearline/Coldline tiers
🔷

Azure Blob Storage

For Azure environments

  • ✅ Azure integration
  • ✅ Cool/Archive tiers
  • ✅ Similar pricing
  • ✅ Good redundancy

🔄 Restoring from Backup

Get your data back!

Before You Restore

CRITICAL CHECKLIST:

  • ⚠️ Stop Cassandra on target node
  • ⚠️ Backup current data (in case restore fails)
  • ⚠️ Verify backup integrity
  • ⚠️ Have rollback plan ready
  • ⚠️ Test on non-production first!
1

Stop Cassandra

# Stop Cassandra service sudo systemctl stop cassandra # Or Docker docker stop cassandra-node1 # Verify it's stopped ps aux | grep cassandra
2

Clear Old Data

# Backup current data (just in case!) mv /var/lib/cassandra/data/my_keyspace/users-* \ /var/lib/cassandra/data/my_keyspace/users-old-$(date +%s) # Or remove completely (careful!) rm -rf /var/lib/cassandra/data/my_keyspace/users-* # Clear commit logs rm -rf /var/lib/cassandra/commitlog/* # Clear saved caches rm -rf /var/lib/cassandra/saved_caches/*
3

Download and Extract Backup

# Download from S3 aws s3 cp s3://my-cassandra-backups/node1/backup-20250106.tar.gz /tmp/ # Extract tar -xzf /tmp/backup-20250106.tar.gz -C /tmp/ # Copy SSTables to data directory cp -r /tmp/backup-20250106/* \ /var/lib/cassandra/data/my_keyspace/users-UUID/ # Fix ownership chown -R cassandra:cassandra /var/lib/cassandra/data/
4

Start Cassandra and Verify

# Start Cassandra sudo systemctl start cassandra # Watch logs tail -f /var/log/cassandra/system.log # Wait for node to be ready (2-5 minutes) # Check status nodetool status # Verify data cqlsh -e "SELECT COUNT(*) FROM my_keyspace.users;" # Run repair to sync with other nodes nodetool repair my_keyspace

Full Cluster Restore

If entire cluster is lost:

  1. Restore snapshots to ALL nodes
  2. Start nodes one at a time
  3. Let each node bootstrap completely
  4. Run repair on each node
  5. Verify data consistency

⚙️ Automating Backups

Set it and forget it!

📝

Backup Script

#!/bin/bash # cassandra-backup.sh # Configuration BACKUP_NAME="backup-$(date +%Y%m%d-%H%M%S)" BACKUP_DIR="/backup/cassandra" S3_BUCKET="s3://my-cassandra-backups" NODE_NAME="node1" RETENTION_DAYS=7 # Create backup directory mkdir -p $BACKUP_DIR/$BACKUP_NAME # Step 1: Flush memtables echo "Flushing memtables..." nodetool flush # Step 2: Take snapshot echo "Creating snapshot..." nodetool snapshot -t $BACKUP_NAME # Step 3: Copy snapshots echo "Copying snapshots..." for keyspace in /var/lib/cassandra/data/*; do for table in $keyspace/*; do snapshot_dir="$table/snapshots/$BACKUP_NAME" if [ -d "$snapshot_dir" ]; then cp -al $snapshot_dir/* $BACKUP_DIR/$BACKUP_NAME/ fi done done # Step 4: Compress echo "Compressing..." tar -czf $BACKUP_DIR/$BACKUP_NAME.tar.gz -C $BACKUP_DIR $BACKUP_NAME # Step 5: Upload to S3 echo "Uploading to S3..." aws s3 cp $BACKUP_DIR/$BACKUP_NAME.tar.gz \ $S3_BUCKET/$NODE_NAME/$BACKUP_NAME.tar.gz # Step 6: Clear snapshot echo "Clearing snapshot..." nodetool clearsnapshot -t $BACKUP_NAME # Step 7: Clean up local files rm -rf $BACKUP_DIR/$BACKUP_NAME rm -f $BACKUP_DIR/$BACKUP_NAME.tar.gz # Step 8: Clean old backups echo "Cleaning old backups..." find $BACKUP_DIR -type f -mtime +$RETENTION_DAYS -delete echo "Backup complete: $BACKUP_NAME"
⏰

Schedule with Cron

# Edit crontab crontab -e # Add daily backup at 2 AM 0 2 * * * /usr/local/bin/cassandra-backup.sh >> /var/log/cassandra-backup.log 2>&1 # Weekly full backup on Sunday at 3 AM 0 3 * * 0 /usr/local/bin/cassandra-full-backup.sh >> /var/log/cassandra-backup.log 2>&1 # Hourly incremental (commit logs) 0 * * * * /usr/local/bin/cassandra-incremental.sh >> /var/log/cassandra-backup.log 2>&1
📧

Monitor Backup Success

# Add to backup script # Send notification on success if [ $? -eq 0 ]; then echo "Backup successful: $BACKUP_NAME" | \ mail -s "Cassandra Backup Success" admin@example.com else echo "Backup FAILED: $BACKUP_NAME" | \ mail -s "ALERT: Cassandra Backup Failed" admin@example.com fi # Or use monitoring service curl -X POST https://healthchecks.io/ping/your-uuid

✅ Testing Your Backups

The most important part!

Untested Backups = No Backups

You don't have a backup until you've verified you can restore it!

Monthly Test Restores

Schedule regular restore drills

  1. Spin up fresh test cluster
  2. Restore latest backup
  3. Verify data integrity
  4. Check row counts match
  5. Test queries work correctly
  6. Measure restore time
  7. Document any issues

Verify Backup Integrity

# Check backup exists aws s3 ls $S3_BUCKET/$NODE_NAME/ # Download and test extraction aws s3 cp $S3_BUCKET/$NODE_NAME/backup-latest.tar.gz /tmp/ tar -tzf /tmp/backup-latest.tar.gz > /dev/null # Check file sizes are reasonable du -sh /tmp/backup-latest.tar.gz

Compare Row Counts

# Production cluster cqlsh prod -e "SELECT COUNT(*) FROM my_keyspace.users;" # Test restore cluster cqlsh test -e "SELECT COUNT(*) FROM my_keyspace.users;" # Counts should match! # If not, investigate discrepancy

Document Recovery Time

Know your RTO (Recovery Time Objective)

  • Download time from S3
  • Extraction time
  • Cassandra startup time
  • Repair time
  • Total: Your RTO

Example: 100GB backup = ~15 min download + 5 min extract + 10 min startup + 30 min repair = 1 hour RTO

🚨 Disaster Recovery Scenarios

Be prepared for the worst!

💻

Single Node Failure

One node dies

Recovery:

  1. Replace hardware
  2. Install Cassandra
  3. Restore snapshot
  4. Run repair
  5. ⏱️ Time: 1-2 hours
🏢

Datacenter Loss

Entire DC offline

Recovery:

  1. Failover to backup DC
  2. Restore new DC from backup
  3. Rebuild from surviving DC
  4. Update application config
  5. ⏱️ Time: 2-6 hours
💥

Total Cluster Loss

Everything gone

Recovery:

  1. Provision new cluster
  2. Restore ALL nodes from S3
  3. Recreate schema
  4. Start nodes sequentially
  5. ⏱️ Time: 4-12 hours

DR Runbook Checklist

Have this documented and ready:

  1. ✅ Emergency contact list
  2. ✅ Backup locations (S3 buckets, paths)
  3. ✅ Restore procedure (step-by-step)
  4. ✅ Node IP addresses / DNS names
  5. ✅ Cassandra configuration files
  6. ✅ Required credentials / keys
  7. ✅ Monitoring dashboard URLs
  8. ✅ Application connection strings
  9. ✅ Post-recovery verification tests

🎉 You're a Backup Expert!

Congratulations! You now know how to protect your Cassandra data!

🎓 What You Learned:

  • 🚨 Why backups matter: GitLab's $1.7M lesson
  • 📂 Backup types: Snapshots, incremental, full cluster
  • 📸 Taking snapshots: Flush, snapshot, verify
  • ☁️ Storage options: S3, GCS, Azure Blob
  • 🔄 Restoration: Complete recovery procedure
  • ⚙️ Automation: Scripts, cron, monitoring
  • ✅ Testing: Monthly test restores, verification
  • 🚨 Disaster recovery: Node, DC, cluster loss

💡 The Golden Rules:

  1. 3-2-1 Rule: 3 copies, 2 media types, 1 offsite
  2. Test regularly: Monthly restore drills
  3. Automate everything: Daily snapshots, weekly uploads
  4. Monitor backups: Alert on failures
  5. Document procedures: DR runbook ready
  6. Measure RTO: Know your recovery time

📋 Quick Reference:

# Daily backup nodetool flush nodetool snapshot -t backup-$(date +%Y%m%d) # Copy to S3 # Clear snapshot # Test restore # 1. Stop Cassandra # 2. Clear data # 3. Extract backup # 4. Start Cassandra # 5. Verify data # 6. Run repair

💾 Remember: Untested backups = No backups!

Advertisement

Responsive Ad