Repair & Consistency
Master Cassandra repair mechanisms, understand read repair vs anti-entropy repair, and learn best practices for maintaining data consistency across your cluster!
🔧 What is Repair in Cassandra?
The Library Book Problem 📚
Imagine you have 3 library branches, each with a copy of the same book catalog. Sometimes:
- 📕 Branch A's copy gets a coffee stain (data corruption)
- 📗 Branch B misses adding a new book (missed write)
- 📘 Branch C has an outdated edition (stale data)
Repair is like sending librarians to compare all 3 copies and make sure they match perfectly!
What is Repair?
Repair is Cassandra's process of ensuring data consistency across replicas. Even though Cassandra is designed to be eventually consistent, repair actively syncs data to fix any inconsistencies.
Two Types of Repair:
1. Read Repair (Automatic)
Happens during reads. Cassandra checks if replicas match and fixes differences automatically.
2. Anti-Entropy Repair (Manual)
You run this manually with nodetool. Compares ALL data across replicas.
❓ Why Does Repair Matter?
Real Production Disaster 💥
Company: Social Media Platform with 10M users
Problem: Skipped repairs for 6 months
Result: During a server restart, they discovered:
- ❌ 15% of user profiles had inconsistent data
- ❌ Some users saw different follower counts depending on which node answered
- ❌ Deleted posts reappeared (tombstones weren't propagated!)
- ❌ Fix took 3 DAYS of emergency maintenance
Cost: $250,000 in lost revenue + angry users
Fix: Started running repairs weekly (should've been doing this all along!)
When Repair is Critical
🚨 Critical Situations
- After node failure - Node was down, missed writes
- After network partition - Nodes couldn't communicate
- Data corruption - Disk errors, bit rot
- Low consistency writes - CL=ONE, data might differ
✅ Regular Maintenance
- Within gc_grace_seconds - Before tombstones deleted (default 10 days)
- Weekly schedule - Industry best practice
- After bulk loads - Imported lots of data
- Preventive care - Stay ahead of problems
The gc_grace_seconds Rule
CRITICAL: You MUST run repair at least once within gc_grace_seconds (default 10 days) on every table. Here's why:
- Day 1: You delete a row, Cassandra creates a tombstone (marker that says "this is deleted")
- Days 2-9: Tombstone spreads to other replicas during reads/repairs
- Day 10: gc_grace_seconds expires, tombstone is permanently deleted
- Problem: If a node missed the tombstone and repair didn't run, deleted data can COME BACK TO LIFE!
⏰ Best Practice: Run repair every 7 days (gives you 3-day buffer before gc_grace_seconds expires)
📖 Read Repair (Automatic & Transparent)
What is Read Repair?
Read Repair happens automatically during SELECT queries. Cassandra compares replicas in the background and fixes any differences. You don't have to do anything!
How Read Repair Works
Read Repair Benefits
- Automatic: No manual intervention needed
- Fast: Fixes data as you query it
- Transparent: Client doesn't wait for repair
- Targeted: Only repairs what you read
Read Repair Limitations
- Only repairs what's read: If you never query row X, it never gets repaired!
- Depends on read_repair_chance: Default is 10% (0.1), so only 10% of reads trigger repair
- Not enough alone: You still MUST run anti-entropy repair regularly
Configuring Read Repair
🔄 Anti-Entropy Repair (Manual, Full Sync)
What is Anti-Entropy Repair?
Anti-Entropy Repair is the complete, thorough comparison of ALL data across replicas. Unlike read repair (which only fixes what you query), this checks EVERYTHING.
Analogy: Library Inventory 📚
Read Repair: When someone checks out a book, librarian notices it's damaged and orders replacement.
Anti-Entropy Repair: Once a week, librarians close the library and check EVERY SINGLE BOOK on EVERY SHELF to make sure everything matches perfectly.
How Anti-Entropy Repair Works
Merkle Trees Explained Simply
Instead of comparing 1 million rows one-by-one (slow!), Cassandra uses Merkle Trees:
- Hash groups of rows: Rows 1-100 = hash ABC123
- Build tree: Combine hashes into tree structure
- Compare tree tops: If top hash matches, ALL data matches!
- Drill down: If mismatch, check which branch differs
- Stream only differences: Only send mismatched data
Result: Compare 1M rows in seconds instead of hours!
⚙️ nodetool repair Commands
Common Repair Options
-pr (Primary Range)
Repairs only data this node is responsible for. Use this!
-hosts
Repair only with specific nodes.
-dc
Repair within specific datacenter only.
-local
Repair only within local datacenter.
Recommended Repair Command
Run this command on EVERY node in your cluster once per week.
✅ Repair Best Practices
✅ DO These
- Run repair every 7 days
- Use -pr flag (primary range)
- Schedule during off-peak hours
- Monitor repair progress
- Use Cassandra Reaper (automation tool)
- Run on one node at a time
❌ DON'T Do These
- Skip repairs for months
- Run repair on all nodes simultaneously
- Repair during peak traffic
- Ignore repair errors
- Rely only on read repair
- Run without monitoring
🤖 Cassandra Reaper (Automated Repair)
What is Reaper?
Cassandra Reaper is an open-source tool that automatically schedules and runs repairs for you. It's like having a robot librarian who checks your books on schedule!
Why Use Reaper?
- ✅ Automatic scheduling: Set it and forget it
- ✅ Web UI: See repair status at a glance
- ✅ Incremental repair: Only repairs what needs it
- ✅ Multi-cluster: Manage multiple clusters
- ✅ Production proven: Used by Netflix, Spotify, Apple
🎓 Repair Summary
Key Takeaways:
- 📖 Read Repair: Automatic, fixes data during queries (10% by default)
- 🔄 Anti-Entropy Repair: Manual, checks ALL data using Merkle trees
- ⏰ Schedule: Run repair every 7 days (before gc_grace_seconds)
- ⚙️ Command: nodetool repair -pr -full (on each node)
- 🤖 Automation: Use Cassandra Reaper for production
Remember: Repair is like regular health checkups - skip them and problems get worse! 🏥
Responsive Ad