Backup & Restore
Protect your data - Because disasters happen!
🚨 The $1.7 Million Lesson
GitLab's 2017 Disaster
January 31, 2017 - 300GB of production data accidentally deleted
What happened:
- 🔴 Engineer removed wrong database directory
- 🔴 6 hours of production data lost
- 🔴 Backup 1: Failed (replication lag)
- 🔴 Backup 2: Failed (accidentally disabled)
- 🔴 Backup 3: Failed (never tested)
- 🔴 Backup 4: Failed (wrong path)
- 🔴 Backup 5: Worked! (LVM snapshots saved them)
💰 The Cost
- 💸 $1.7M in lost revenue
- 😰 18 hours of downtime
- 📉 Brand damage and customer trust
- ⚠️ Lesson: 4 out of 5 backup methods failed!
✅ Don't Be Like GitLab
- 💾 Multiple backup methods
- 🏗️ Separate storage locations
- ✅ Test restores regularly
- ⚙️ Automate everything
- 📝 Document procedures
- 🚨 Monitor backup success
💾 Backups aren't real until you've tested a restore!
📂 Types of Backups
Choose the right strategy!
Snapshots
Point-in-time backups
- ✅ Fast to create (seconds)
- ✅ Minimal performance impact
- ✅ Full table backup
- ⚠️ Uses disk space
- 💡 Best for: Daily backups
Incremental
Only changed data
- ✅ Space efficient
- ✅ Automatic with commit logs
- ✅ Point-in-time recovery
- ⚠️ Harder to restore
- 💡 Best for: Continuous backup
Full Cluster
Complete cluster copy
- ✅ Complete recovery
- ✅ Easiest to restore
- ⚠️ Slow to create
- ⚠️ Large storage needs
- 💡 Best for: Major migrations
Recommended Strategy: 3-2-1 Rule
- 3 copies of your data (production + 2 backups)
- 2 different media types (local disk + cloud)
- 1 copy offsite (S3, different datacenter)
Example: Production data + Daily snapshots on local disk + Weekly backups to S3
📸 Creating Snapshots
Fast, reliable backups!
Flush Memtables First
Ensure all data is on disk
Take Snapshot
Verify Snapshot
Important: Snapshots Take Disk Space!
Snapshots are hard links, but still use space:
- ⚠️ Initial snapshot: Uses ~same space as table
- ⚠️ Data updates: New SSTables created, old kept by snapshot
- ⚠️ Can fill disk if not cleaned up!
- ✅ Always monitor disk space
- ✅ Clear old snapshots regularly
☁️ Storing Backups
Where to keep your backups safe!
Copy to Archive Directory
Upload to S3 / Cloud Storage
Clean Up Old Snapshots
Amazon S3
Most popular choice
- ✅ Very cheap ($0.023/GB)
- ✅ 99.999999999% durability
- ✅ Easy to use (AWS CLI)
- ✅ Lifecycle policies
Google Cloud Storage
Google's alternative
- ✅ Comparable pricing
- ✅ Good for GCP users
- ✅ Fast transfers
- ✅ Nearline/Coldline tiers
Azure Blob Storage
For Azure environments
- ✅ Azure integration
- ✅ Cool/Archive tiers
- ✅ Similar pricing
- ✅ Good redundancy
🔄 Restoring from Backup
Get your data back!
Before You Restore
CRITICAL CHECKLIST:
- ⚠️ Stop Cassandra on target node
- ⚠️ Backup current data (in case restore fails)
- ⚠️ Verify backup integrity
- ⚠️ Have rollback plan ready
- ⚠️ Test on non-production first!
Stop Cassandra
Clear Old Data
Download and Extract Backup
Start Cassandra and Verify
Full Cluster Restore
If entire cluster is lost:
- Restore snapshots to ALL nodes
- Start nodes one at a time
- Let each node bootstrap completely
- Run repair on each node
- Verify data consistency
⚙️ Automating Backups
Set it and forget it!
Backup Script
Schedule with Cron
Monitor Backup Success
✅ Testing Your Backups
The most important part!
Untested Backups = No Backups
You don't have a backup until you've verified you can restore it!
Monthly Test Restores
Schedule regular restore drills
- Spin up fresh test cluster
- Restore latest backup
- Verify data integrity
- Check row counts match
- Test queries work correctly
- Measure restore time
- Document any issues
Verify Backup Integrity
Compare Row Counts
Document Recovery Time
Know your RTO (Recovery Time Objective)
- Download time from S3
- Extraction time
- Cassandra startup time
- Repair time
- Total: Your RTO
Example: 100GB backup = ~15 min download + 5 min extract + 10 min startup + 30 min repair = 1 hour RTO
🚨 Disaster Recovery Scenarios
Be prepared for the worst!
Single Node Failure
One node dies
Recovery:
- Replace hardware
- Install Cassandra
- Restore snapshot
- Run repair
- ⏱️ Time: 1-2 hours
Datacenter Loss
Entire DC offline
Recovery:
- Failover to backup DC
- Restore new DC from backup
- Rebuild from surviving DC
- Update application config
- ⏱️ Time: 2-6 hours
Total Cluster Loss
Everything gone
Recovery:
- Provision new cluster
- Restore ALL nodes from S3
- Recreate schema
- Start nodes sequentially
- ⏱️ Time: 4-12 hours
DR Runbook Checklist
Have this documented and ready:
- ✅ Emergency contact list
- ✅ Backup locations (S3 buckets, paths)
- ✅ Restore procedure (step-by-step)
- ✅ Node IP addresses / DNS names
- ✅ Cassandra configuration files
- ✅ Required credentials / keys
- ✅ Monitoring dashboard URLs
- ✅ Application connection strings
- ✅ Post-recovery verification tests
🎉 You're a Backup Expert!
Congratulations! You now know how to protect your Cassandra data!
🎓 What You Learned:
- 🚨 Why backups matter: GitLab's $1.7M lesson
- 📂 Backup types: Snapshots, incremental, full cluster
- 📸 Taking snapshots: Flush, snapshot, verify
- ☁️ Storage options: S3, GCS, Azure Blob
- 🔄 Restoration: Complete recovery procedure
- ⚙️ Automation: Scripts, cron, monitoring
- ✅ Testing: Monthly test restores, verification
- 🚨 Disaster recovery: Node, DC, cluster loss
💡 The Golden Rules:
- 3-2-1 Rule: 3 copies, 2 media types, 1 offsite
- Test regularly: Monthly restore drills
- Automate everything: Daily snapshots, weekly uploads
- Monitor backups: Alert on failures
- Document procedures: DR runbook ready
- Measure RTO: Know your recovery time
📋 Quick Reference:
💾 Remember: Untested backups = No backups!
Responsive Ad