COMPACTION
Complete Beginner's Guide
The cleanup process that keeps Cassandra fast and efficient! 🗜️✨
📖 Prerequisites - What You Should Know First
Before learning about Compaction, let's understand some basic concepts. Don't worry - we'll explain everything from scratch!
What is an SSTable?
SSTable = Sorted String Table = A file on disk that stores data
Real-world analogy:
Think of photo albums:
• SSTable: One photo album with sorted pictures
• Multiple SSTables: Multiple albums piling up over time
• Problem: Too many albums = hard to find photos!
In Cassandra:
• Every time you write data → Creates new SSTable
• SSTables are IMMUTABLE (never modified, only created)
• Over time: Hundreds of SSTables pile up!
• Problem: Reading data = checking ALL SSTables! 🐌
Example:
SSTable-1.db: {user_id=1: "Alice", user_id=2: "Bob"}
SSTable-2.db: {user_id=3: "Charlie", user_id=1: "Alice Smith"}
SSTable-3.db: {user_id=2: "Bob Jones", user_id=4: "Diana"}
Query user_id=1 → Must check ALL 3 SSTables! 🐌
Latest value wins: "Alice Smith" from SSTable-2
What are Tombstones?
Tombstone = A marker that says "this data was deleted"
Why do we need them?
• Remember: SSTables are immutable!
• Can't delete data from existing SSTable
• Solution: Write a "deletion marker" (tombstone)
Library book analogy:
• You can't erase text from a library book
• Instead: Add a sticky note saying "IGNORE THIS PAGE"
• That sticky note = Tombstone!
Example:
SSTable-1: {user_id=100: "John"}
DELETE user_id=100; ← User runs delete
SSTable-2: {user_id=100: TOMBSTONE} ← Marker created!
Now when reading: See tombstone → Return "NOT FOUND"
But disk space still used! Both SSTables exist! 💾
What is Disk Space Amplification?
Space Amplification = Using more disk space than your actual data needs
Closet analogy:
• You buy new clothes → Old clothes stay in closet
• Update profile photo → Old photos stay on disk
• Delete account → Tombstone markers stay on disk
• Result: Closet full of stuff you don't wear! 📦
The problem:
• Actual data: 100 GB
• Old versions: 50 GB
• Tombstones: 30 GB
• Duplicate data across SSTables: 70 GB
• Total disk used: 250 GB!
• Space amplification: 2.5x (using 2.5x more space!)
Why it matters:
• Wasting expensive SSD space 💰
• Slower reads (more files to check) 🐌
• Higher costs ❌
Ready to Continue?
Great! Now you know:
✅ SSTable = File on disk with data
✅ Tombstone = Deletion marker
✅ Space Amplification = Wasting disk space
✅ Problem: Too many SSTables + old data = SLOW! 🐌
Now let's learn how Compaction solves all these problems! 🚀
📱 WhatsApp: Managing 2 Billion Users with Compaction
WhatsApp stores 100+ billion messages per day using Cassandra.
The Challenge (Without Compaction):
• Every message write → New SSTable created
• 100 billion writes/day = Millions of SSTables!
• Each query must check ALL SSTables
• Message load time: 5 seconds! 🐌
• Disk space: 10 PB used for 4 PB of actual data!
• Space amplification: 2.5x ❌
Example scenario:
User sends 100 messages:
• 100 SSTables created
• User updates profile → 1 more SSTable
• User deletes 10 messages → 10 tombstone SSTables
• Total: 111 SSTables for one user! 😱
To load chat:
• Check SSTable-1, SSTable-2, ... SSTable-111
• Merge results
• Remove deleted messages (tombstones)
• Time: 5+ seconds!
The Solution: Compaction
• Automatically merge SSTables in background
• 111 SSTables → 1 merged SSTable!
• Remove tombstones permanently
• Remove old versions of data
After Compaction:
• Load chat: Check 1 SSTable instead of 111
• Message load time: 0.5 seconds (10x faster!) ⚡
• Disk space: 4 PB (actual data, no waste!)
• Space amplification: 1.1x ✓
• Cost savings: Reduced storage by 60%! 💰
WhatsApp's Compaction Strategy:
• Runs continuously in background
• Merges old SSTables every few hours
• Keeps system performant for 2 billion users
• This is why WhatsApp is instant! ⚡
This is the power of Compaction! 🗜️✨
❓ What is Compaction?
Simple Definition
Compaction is the process of merging multiple SSTables into fewer, larger SSTables while:
• Removing deleted data (tombstones)
• Removing old versions
• Keeping only latest data
• Freeing disk space
Spring cleaning analogy:
• Your room has clothes everywhere (many SSTables)
• Spring cleaning: Organize everything into closet (merge SSTables)
• Throw away old clothes (remove tombstones)
• Donate duplicates (remove old versions)
• Result: Clean room with only what you need! ✨
Photo Album Analogy
Imagine you have 10 photo albums:
Album 1: Vacation 2020 (50 photos)
Album 2: Vacation 2020 - retakes (10 better photos)
Album 3: Vacation 2020 - deleted (5 photos to remove)
Album 4-10: More updates and deletions
Problem: To see vacation photos, you must check all 10 albums! 🐌
Compaction = Creating ONE master album:
• Take all albums
• Keep only latest versions of each photo
• Remove deleted photos
• Organize into ONE album
• Throw away old albums
Result:
• 1 album instead of 10 ✓
• Only 45 final photos (removed 5 deleted ones) ✓
• No duplicate photos ✓
• Easy to browse! ⚡
🤔 Why Do We Need Compaction?
Problem 1: Slow Reads
Too many SSTables = Slow queries
Without Compaction:
• 100 SSTables must be checked
• Each: 10ms read time
• Total: 1000ms (1 second!) 🐌
With Compaction:
• 10 merged SSTables
• Total: 100ms ⚡
• 10x faster!
Problem 2: Wasted Space
Old data wastes disk space
• Actual data: 100 GB
• Old versions: 80 GB
• Tombstones: 20 GB
• Total: 200 GB wasted!
After Compaction:
• Saves 90 GB (45%) ✓
Problem 3: Tombstones
Deletions don't free space
• Delete 1M records
• 1M tombstones created
• Still using disk! 💾
Compaction removes them:
• Space freed! ✓
• Queries faster! ⚡
📝 Simple Compaction Example
BEFORE Compaction
user_id=200: "Bob v1"
user_id=300: "Charlie"
user_id=400: "Diana"
Total: 3 SSTables, 23 MB, duplicates + tombstones
AFTER Compaction
user_id=300: "Charlie"
user_id=400: "Diana"
✓ Total: 1 SSTable, 12 MB
✓ Removed: Old "Alice v1", Deleted "Bob", Tombstone
✓ Saved: 11 MB (48% space saved!)
✓ Speed: 3x faster queries! ⚡
⚙️ How Compaction Works
The Compaction Process
Step 1: Select SSTables
Strategy decides which SSTables to merge
Step 2: Read all data
Read selected SSTables into memory
Step 3: Merge & deduplicate
• Keep latest version of each key
• Remove tombstones (if old enough)
• Sort all data
Step 4: Write new SSTable
Create merged SSTable on disk
Step 5: Atomically swap
• Mark new SSTable as active
• Delete old SSTables
• Free disk space! ✓
📊 3 Compaction Strategies
Cassandra offers 3 main strategies. Each optimized for different workloads!
Size Tiered (STCS)
Merges similar-sized SSTables
Best for: Write-heavy workloads
Pro: Fast writes ⚡
Con: Uses more disk space
Example: Logs, time-series
Leveled (LCS)
Organizes data into levels
Best for: Read-heavy workloads
Pro: Fast reads, less space
Con: More I/O overhead
Example: User profiles
Time Window (TWCS)
Groups by time buckets
Best for: Time-series data with TTL
Pro: Easy to delete old data
Con: Only for time-series
Example: IoT sensors, metrics
📈 Size Tiered Compaction Strategy (STCS)
How STCS Works
Merges SSTables of similar size together
Process:
1. Group SSTables by size
2. When 4+ similar size → Merge them
3. Creates larger SSTable
4. Repeat for larger sizes
Best for: Write-heavy, append-only workloads
Used by: Twitter (tweets), logging systems
⏰ Leveled Compaction Strategy (LCS)
How LCS Works
Organizes SSTables into levels by size
Levels:
• L0: Small SSTables (new writes)
• L1-L8: Progressively larger, non-overlapping
Best for: Read-heavy workloads
Used by: Facebook (user profiles), Instagram
🕐 Time Window Compaction Strategy (TWCS)
How TWCS Works
Groups data into time buckets
Process:
• Creates buckets (hourly, daily, etc.)
• Each bucket = one SSTable
• Old buckets deleted entirely when TTL expires
Best for: Time-series with TTL
Used by: Netflix (viewing metrics), IoT platforms
⚖️ Which Strategy to Choose?
-- Strategy Comparison
📈 Size Tiered (STCS) - DEFAULT
Best for: Write-heavy workloads
Pros: Fast writes, simple
Cons: Uses 2-3x more disk space
Example: INSERT INTO logs VALUES (...)
⏰ Leveled (LCS)
Best for: Read-heavy workloads
Pros: 90% less space, faster reads
Cons: More write amplification
Example: SELECT * FROM users WHERE id=?
🕐 Time Window (TWCS)
Best for: Time-series data
Pros: Easy old data deletion
Cons: Only for time-bucketed data
Example: INSERT INTO metrics (time, value) ...
Quick Decision Guide
Choose STCS if:
• Mostly writes, few reads
• Have plenty of disk space
• Need simple setup (default)
Choose LCS if:
• Mostly reads
• Limited disk space
• Need predictable performance
Choose TWCS if:
• Time-series data
• Data has TTL
• Need to delete old data efficiently
🏢 How Real Companies Use Compaction
🐦 Twitter - STCS for Tweets
Workload: 500M tweets/day, write-heavy
Strategy: Size Tiered (STCS)
Result: Handles massive write volume efficiently
📘 Facebook - LCS for Profiles
Workload: 3B users, read-heavy
Strategy: Leveled (LCS)
Result: 90% less disk space, faster profile loads
🎬 Netflix - TWCS for Metrics
Workload: Billions of viewing metrics/hour
Strategy: Time Window (TWCS)
Result: Auto-deletes old data, perfect for time-series
💼 Interview Questions & Answers
1 What is compaction in Cassandra and why is it needed?
Answer:
Compaction is the process of merging multiple SSTables into fewer, larger SSTables while removing deleted data (tombstones) and old versions.
Why it's needed:
- Performance: Too many SSTables = slow reads (must check all)
- Space: Tombstones and old versions waste disk space
- Efficiency: Merging reduces files from 100s → 10s
Without compaction: System degrades over time, queries get slower, disk fills up.
2 Compare STCS, LCS, and TWCS compaction strategies
Answer:
Size Tiered (STCS):
- Merges similar-sized SSTables
- Best for: Write-heavy workloads
- Pro: Fast writes, simple
- Con: 2-3x space amplification
Leveled (LCS):
- Organizes SSTables into levels
- Best for: Read-heavy workloads
- Pro: 90% less space, faster reads
- Con: More I/O during compaction
Time Window (TWCS):
- Groups data by time buckets
- Best for: Time-series with TTL
- Pro: Easy deletion of old data
- Con: Only for time-bucketed data
3 What are tombstones and how does compaction handle them?
Answer:
Tombstones are markers indicating deleted data. Since SSTables are immutable, Cassandra can't delete data in-place - instead it writes a tombstone.
How compaction handles them:
- During compaction, check tombstone age
- If older than gc_grace_seconds (default: 10 days) → Permanently remove
- If newer → Keep (needed for distributed deletes)
Why wait 10 days? Ensures all replicas see the deletion before removing tombstone.
4 What is space amplification and how does compaction reduce it?
Answer:
Space amplification is the ratio of disk space used vs actual data size.
Example:
- Actual data: 100 GB
- Disk used: 250 GB
- Space amplification: 2.5x
Causes:
- Old versions of updated data
- Tombstones from deletions
- Duplicate keys across SSTables
How compaction helps: Merges SSTables, keeps only latest version, removes tombstones → Reduces to 1.1-1.5x amplification
5 When would you choose LCS over STCS?
Answer:
Choose LCS (Leveled) over STCS (Size Tiered) when:
1. Read-heavy workload:
- LCS ensures data is in fewer, non-overlapping SSTables
- Reads check fewer files → Faster
2. Limited disk space:
- LCS: 1.1x space amplification
- STCS: 2-3x space amplification
3. Predictable performance needed:
- LCS has consistent read latency
- STCS can have variable performance
Trade-off: LCS requires more I/O during compaction, so writes are slightly slower.
Real example: Facebook uses LCS for user profiles (read-heavy, limited space)
Responsive Ad