Compaction Strategies in Cassandra

Keep your database neat and fast! Compaction strategies merge and reorganize SSTables in the background, reducing redundancy, improving read performance, and ensuring your data stays efficient over time.

📖 The Story: Mike's Disk Space Disaster

Mike's IoT sensor platform collected 1TB of data daily. Everything worked fine for 3 months, then disaster struck. Here's what went wrong...

💥 The Disaster: Wrong Compaction Strategy

The Setup:

  • 📊 Data Pattern: Time-series sensor data (append-only)
  • 💾 Write Rate: 100K writes/second
  • 📈 Retention: 90 days, then delete old data
  • ⚙️ Compaction: Default SizeTieredCompactionStrategy (STCS)

What Went Wrong:

  1. Month 1: Everything fine - 30TB storage used
  2. Month 2: 60TB storage (expected growth)
  3. Month 3: 180TB storage! 💥 (should be 90TB!)
  4. Day 91: Tried to delete old data with TTL...
  5. Result: Data deleted but disk space NOT reclaimed!
  6. Crisis: Disks 95% full! Cluster at risk!

Why It Happened:

  • 🔄 STCS Problem: Creates 4 equal-sized SSTables, merges them into 1 big one
  • 💾 Space Amplification: Needs 2x space during compaction (5 SSTables exist temporarily)
  • ⏰ TTL Data: Old data marked as deleted but stays in old SSTables
  • 🐌 Slow Cleanup: Takes months to compact away expired data
  • 💥 Disk Full: Can't compact because no space to create new SSTables!

✅ The Fix: TimeWindowCompactionStrategy!

Mike Changed Compaction Strategy:

-- Changed from STCS to TWCS ALTER TABLE sensor_data WITH compaction = { 'class': 'TimeWindowCompactionStrategy', 'compaction_window_size': '1', 'compaction_window_unit': 'DAYS' };

How TWCS Fixed Everything:

  1. Time Windows: Creates separate SSTables per day
  2. No Cross-Window Compaction: Day 1 data never mixed with Day 2
  3. Easy TTL Cleanup: When Day 91 arrives, entire Day 1 SSTable deleted!
  4. No Space Amplification: Just delete old SSTable files - instant!
  5. Predictable Storage: 90 days × 1TB/day = 90TB (exactly!)

The Results:

  • 💾 Storage: 180TB → 90TB (50% reduction!)
  • ⚡ TTL Cleanup: Months → Instant (just delete file!)
  • 📉 Space Amplification: 2x → 1x (no temp space!)
  • 🚀 Write Performance: Improved (less compaction overhead)
  • 😊 Predictable Costs: Exactly 90TB always

Mike learned: Right compaction strategy = Critical for production! 🎉

🔄 Compaction Fundamentals

Understand how compaction works and why it matters!

🎯 What is Compaction?

Compaction is the background process that merges multiple SSTables into fewer, larger SSTables. Think of it like defragmenting your hard drive - it reorganizes data for better performance.

Why Compaction is Necessary

1. Remove Deleted Data

Cassandra doesn't delete immediately - marks as tombstone. Compaction actually removes deleted data.

2. Merge Updates

Multiple updates to same row exist in different SSTables. Compaction merges them into latest version.

3. Reduce SSTable Count

Fewer SSTables = faster reads! Reading from 50 SSTables vs 5 SSTables is 10x slower.

4. Reclaim Disk Space

Expired TTL data and tombstones take up space. Compaction removes them and frees disk.

Compaction Process Before Compaction SSTable 1 5MB SSTable 2 5MB SSTable 3 5MB Compact After Compaction Merged SSTable 12MB (not 15MB!) Benefits ✅ 3 SSTables → 1 SSTable ✅ Faster reads (1 file vs 3) ✅ Removed duplicates ✅ Removed tombstones ✅ Reclaimed 3MB space

📦 SizeTieredCompactionStrategy (STCS)

The default strategy - good for write-heavy workloads

How STCS Works

STCS groups SSTables of similar size together and compacts them into one larger SSTable. Typically merges 4 SSTables at a time.

STCS Algorithm

-- STCS Example Configuration CREATE TABLE user_activity ( user_id UUID, activity_time TIMESTAMP, activity_type TEXT, PRIMARY KEY (user_id, activity_time) ) WITH compaction = { 'class': 'SizeTieredCompactionStrategy', 'min_threshold': '4', -- Compact when 4 similar SSTables exist 'max_threshold': '32' -- Max SSTables to compact at once };

STCS Process

Step 1: Flush Creates SSTables

Memtable flushes create small SSTables (5-100MB)

📁 SSTable-1 (5MB), SSTable-2 (5MB), SSTable-3 (5MB), SSTable-4 (5MB)

Step 2: First Compaction (Tier 1)

4 similar-sized SSTables → Compact into 1

🔄 Compact → SSTable-5 (18MB)

Step 3: More Flushes

More writes create more small SSTables

📁 SSTable-6 (5MB), SSTable-7 (5MB), SSTable-8 (5MB), SSTable-9 (5MB)

Step 4: Second Compaction (Tier 1)

Another 4 small SSTables compact

🔄 Compact → SSTable-10 (18MB)

Step 5: Tier 2 Compaction

Eventually 4 medium SSTables (~20MB) compact

🔄 Compact 4×20MB → SSTable-11 (72MB)

When to Use STCS

✅

Best For

  • Write-Heavy: High write throughput
  • Insert-Only: No updates/deletes
  • Short TTL: Or no TTL at all
  • General Purpose: Mixed workloads
  • Low Read Latency Not Critical: Reads can tolerate 10-50ms
❌

Avoid For

  • Time-Series: With TTL (use TWCS!)
  • Read-Heavy: Many reads (use LCS!)
  • Update-Heavy: Frequent updates
  • Long TTL: > 30 days with deletes
  • Disk Space Limited: Needs 2x temporary space

STCS Limitations

  • ⚠️ Space Amplification: Needs 50-100% extra disk space during compaction
  • ⚠️ Read Performance: May need to read from 20+ SSTables
  • ⚠️ TTL Cleanup: Takes months to reclaim expired data
  • ⚠️ Large SSTables: Eventually creates multi-GB SSTables (slow to compact)

📚 LeveledCompactionStrategy (LCS)

For read-heavy workloads with frequent updates

🎯 How LCS Works

LCS organizes SSTables into levels (0, 1, 2, ...). Level 0 has small overlapping SSTables. Higher levels have fixed-size (160MB) non-overlapping SSTables. This guarantees reads touch ≤ 10 SSTables!

LCS Level Structure

-- LCS Configuration ALTER TABLE user_profiles WITH compaction = { 'class': 'LeveledCompactionStrategy', 'sstable_size_in_mb': '160' -- Default: 160MB per SSTable }; -- Level Structure: /* Level 0: 4-8 SSTables (overlapping, from memtable flush) Level 1: ~10 SSTables (160MB each, non-overlapping) Level 2: ~100 SSTables (160MB each, non-overlapping) Level 3: ~1000 SSTables (160MB each, non-overlapping) ... */

LCS Compaction Process

Level 0 → Level 1

When Level 0 reaches 4+ SSTables, compact them with overlapping Level 1 SSTables

L0: [SSTable-1, SSTable-2, SSTable-3, SSTable-4] // 20MB each ↓ Compact with overlapping L1 L1: [SSTable-5, SSTable-6, SSTable-7] // 160MB each, non-overlapping

Level 1 → Level 2

When Level 1 grows > 10 SSTables, promote to Level 2

L1: 12 SSTables × 160MB = 1.92GB // Exceeds limit! ↓ Compact with overlapping L2 L2: [Many 160MB SSTables, all non-overlapping]

LCS Read Performance

⚡ Why LCS Reads are Fast

Key Guarantee: Each level has non-overlapping SSTables. This means for any given key, you only need to check:

  • ✅ All Level 0 SSTables (4-8 max)
  • ✅ At most 1 SSTable per level (1-9)
  • ✅ Total: ≤ 10 SSTables maximum!

Compare to STCS: Might need to read from 50+ SSTables!

When to Use LCS

✅

Best For

  • Read-Heavy: 80%+ reads
  • Update-Heavy: Frequent updates to same rows
  • Low Read Latency: Need < 10ms reads
  • User Profiles: Frequently updated data
  • Predictable Performance: Consistent query times
❌

Avoid For

  • Write-Heavy: > 10K writes/sec
  • Time-Series: With TTL (use TWCS!)
  • Insert-Only: No updates (STCS better)
  • Low Disk IOPS: LCS needs good disks
  • High Churn: Constant compaction overhead

LCS Performance Characteristics

Metric STCS LCS
Read Performance 10-50ms (50+ SSTables) 2-10ms (≤10 SSTables)
Write Amplification 5-10x 10-20x (more compaction!)
Space Amplification 50-100% extra 10-20% extra (better!)
Compaction I/O Periodic spikes Continuous background

⏰ TimeWindowCompactionStrategy (TWCS)

Perfect for time-series data with TTL!

🎯 How TWCS Works

TWCS creates separate SSTables for each time window (hour/day/week). SSTables from different windows are NEVER compacted together. When TTL expires, simply delete the entire old SSTable file - instant space reclamation!

TWCS Configuration

-- TWCS for IoT sensor data (1-day windows) CREATE TABLE sensor_readings ( sensor_id UUID, reading_time TIMESTAMP, temperature DECIMAL, PRIMARY KEY (sensor_id, reading_time) ) WITH compaction = { 'class': 'TimeWindowCompactionStrategy', 'compaction_window_size': '1', 'compaction_window_unit': 'DAYS' } AND default_time_to_live = 7776000; -- 90 days -- For metrics (1-hour windows) ALTER TABLE application_metrics WITH compaction = { 'class': 'TimeWindowCompactionStrategy', 'compaction_window_size': '1', 'compaction_window_unit': 'HOURS' };

TWCS Time Windows

TWCS Time Windows Day 1 Window SSTable-D1 Jan 1, 2025 Day 2 Window SSTable-D2 Jan 2, 2025 Day 3 Window SSTable-D3 Jan 3, 2025 Today SSTable-D90 Mar 31, 2025 TTL Expired! DELETE SSTable-D1.db ✅ Instant Cleanup! No compaction needed Just delete file Space freed immediately

TWCS vs STCS for Time-Series

❌

STCS with TTL

  • Old data mixed with new in same SSTable
  • Must compact to remove expired data
  • Takes months to reclaim space
  • 2x disk space during compaction
  • Unpredictable storage usage

Example: 90 days data = 180TB disk!

✅

TWCS with TTL

  • Old data isolated in separate SSTables
  • Just delete expired SSTable file
  • Instant space reclamation
  • No extra disk space needed
  • Predictable storage usage

Example: 90 days data = 90TB disk! ⚡

When to Use TWCS

Perfect For

  • ✅ Time-Series Data: Logs, metrics, events, sensor readings
  • ✅ TTL Data: Automatic expiration (7-365 days)
  • ✅ Append-Only: Insert-only, no updates to old data
  • ✅ Recent Data Queries: Mostly query last 7 days
  • ✅ Predictable Storage: Know exactly how much space needed

Don't Use For

  • ❌ No TTL: Data kept forever (use STCS or LCS)
  • ❌ Update Old Data: Frequent updates to historical records
  • ❌ Random Access: Query any time period equally
  • ❌ Non-Time-Series: User profiles, products, etc.

TWCS Best Practices

Choose Right Window Size

-- High-frequency metrics (millions/sec): 1 hour windows compaction_window_unit = 'HOURS' -- IoT sensors (thousands/sec): 1 day windows compaction_window_unit = 'DAYS' -- Application logs (hundreds/sec): 1 day windows compaction_window_unit = 'DAYS' -- Rule: Window should contain 50-200GB of data

Set Appropriate TTL

-- Always set TTL with TWCS! ALTER TABLE sensor_data WITH default_time_to_live = 2592000; -- 30 days -- Or per-row TTL INSERT INTO sensor_data (...) VALUES (...) USING TTL 2592000;

Monitor Expired SSTables

-- Check for dropped SSTables nodetool tablestats keyspace.table -- Look for: /* Space used (live): 450GB Space used (total): 455GB ← Should be close! Number of partitions (estimate): 10000000 Pending tombstone collection: 5GB ← Should be small! */

⚖️ Strategy Comparison & Decision Guide

Choose the right strategy for your use case!

Complete Comparison Table

Feature STCS LCS TWCS
Best For Write-heavy, insert-only Read-heavy, updates Time-series with TTL
Read Performance 10-50ms (many SSTables) 2-10ms (≤10 SSTables) 5-20ms (recent data fast)
Write Amplification 5-10x 10-20x (high!) 2-5x (best!)
Space Amplification 50-100% extra 10-20% extra 5-10% extra
TTL Cleanup Months (slow!) Days-weeks Instant (delete file!)
Compaction I/O Periodic spikes Continuous background Low (per-window)
Default Choice ✅ Yes (general) No (specialized) No (specialized)

Decision Flowchart

🤔 Which Strategy Should I Use?

Question 1: Is this time-series data with TTL?

✅ YES → Use TWCS! (IoT, logs, metrics)

❌ NO → Go to Question 2

Question 2: Is this read-heavy with frequent updates?

✅ YES → Use LCS! (User profiles, product catalog)

❌ NO → Go to Question 3

Question 3: Is read latency critical (< 10ms required)?

✅ YES → Use LCS!

❌ NO → Use STCS! (Default, general purpose)

Real-World Examples

📦

Use STCS

  • Social Media Posts: Insert-only, no TTL
  • Order History: Append-only records
  • Audit Logs: No expiration, kept forever
  • Comments/Reviews: Rarely updated
  • Transaction Records: Immutable data
📚

Use LCS

  • User Profiles: Frequent updates, read-heavy
  • Product Catalog: Prices/inventory change
  • Account Balances: Updated frequently
  • Shopping Carts: Add/remove items
  • Session Data: Short-lived, updated often
⏰

Use TWCS

  • IoT Sensor Data: 90-day retention
  • Application Metrics: 30-day retention
  • Server Logs: 7-day retention
  • Event Streams: Time-ordered, expires
  • Analytics Data: Rolling window

📊 Monitoring Compaction

Essential commands and metrics to watch!

Key Nodetool Commands

-- Check compaction status nodetool compactionstats -- Output: /* pending tasks: 127 ← High = compaction falling behind! Active compaction remaining time : 2h30m */ -- Check table statistics nodetool tablestats keyspace.table -- Look for: /* SSTable count: 247 ← High = slow reads! Space used (live): 450GB Space used (total): 900GB ← 2x = PROBLEM! Space amplification */ -- Force compaction (major compaction) nodetool compact keyspace table -- ⚠️ Use sparingly! Creates one huge SSTable -- Set compaction throughput nodetool setcompactionthroughput 64 -- MB/sec (default: 16) -- Stop compaction (emergency) nodetool stop COMPACTION

Healthy vs Unhealthy Metrics

Metric Healthy Unhealthy
Pending compactions < 20 > 100 (falling behind!)
SSTable count < 50 > 200 (slow reads!)
Space amplification < 1.3x (30% extra) > 2x (100%+ extra!)
Pending tombstones < 5% > 20% (wasting space!)

Warning Signs

  • ⚠️ Pending tasks > 100: Compaction can't keep up - increase throughput or add nodes
  • ⚠️ SSTable count > 200: Reads will be very slow - force compaction
  • ⚠️ Space 2x expected: Wrong strategy (STCS with TTL?) - switch to TWCS
  • ⚠️ Disk > 80% full: Risk of out-of-space during compaction!

💼 Interview Questions & Expert Answers

Master compaction strategies for your interview!

1 Explain why TimeWindowCompactionStrategy is essential for time-series data with TTL. What problem does it solve? ▼

Answer: TWCS solves the space amplification problem with TTL data by isolating time windows into separate SSTables, enabling instant space reclamation when data expires.

The Problem with STCS + TTL:

  1. Data Mixing: Old (expired) and new data get merged into same SSTables during compaction
  2. Tombstone Waiting: Expired data becomes tombstones but stays in SSTables
  3. Slow Cleanup: Must wait for compaction to process entire SSTable to remove tombstones
  4. Space Amplification: Can use 2-3x expected space (100TB → 200-300TB!)
  5. Compaction Overhead: Needs 50-100% extra space during compaction

How TWCS Solves It:

  1. Time Windows: Creates separate SSTables per time period (hour/day/week)
  2. No Cross-Window Mixing: Day 1 data never compacted with Day 2 data
  3. Whole-File Expiration: When Day 91 arrives, entire Day 1 SSTable expires
  4. Instant Deletion: Just delete the SSTable file - no compaction needed!
  5. Predictable Space: 90 days retention = exactly 90 days worth of space

Real Numbers:

-- Scenario: 1TB/day, 90-day TTL -- With STCS: // Expected: 90TB (90 days × 1TB) // Actual: 180-270TB! (2-3x space amplification) // Reason: Old data mixed with new in SSTables -- With TWCS: // Expected: 90TB // Actual: 90-95TB (5-10% overhead only) // Reason: Clean window boundaries, instant deletion

Key Takeaway: For time-series with TTL, TWCS is not optional - it's essential to avoid disk space disasters!

2 When would you choose LeveledCompactionStrategy over SizeTieredCompactionStrategy? ▼

Answer: Choose LCS for read-heavy workloads with frequent updates where read latency must be consistently low (< 10ms).

LCS is Better When:

  • ✅ Read-Heavy: 80%+ reads, < 20% writes
  • ✅ Update-Heavy: Same rows updated frequently
  • ✅ Predictable Latency: Need consistent < 10ms reads
  • ✅ Space-Constrained: Can't afford 2x space amplification

Why LCS is Faster for Reads:

LCS organizes SSTables into non-overlapping levels. For any given key:

  • Check all Level 0 SSTables (4-8 max)
  • Check at most 1 SSTable per level (Levels 1-9)
  • Total: ≤ 10 SSTables maximum!

STCS vs LCS Performance:

-- STCS: May need to check 50+ SSTables // Read latency: 20-50ms (checking many files) -- LCS: Maximum 10 SSTables // Read latency: 2-10ms (fewer files)

Trade-offs:

Metric STCS LCS
Read latency 20-50ms 2-10ms ✅
Write amplification 5-10x ✅ 10-20x ❌
Space amplification 50-100% 10-20% ✅

Examples:

  • Use LCS: User profiles, product catalog, account balances
  • Use STCS: Social posts, orders, immutable logs
3 What is write amplification and how does it relate to compaction strategy choice? ▼

Answer: Write amplification is the ratio of data written to disk vs data written by application. Compaction causes write amplification because data is rewritten multiple times during merging.

How Write Amplification Works:

  1. Application writes: 100MB of data
  2. Cassandra writes: 100MB to commit log + memtable
  3. Memtable flush: 100MB written to SSTable (1st time)
  4. Tier 1 compaction: 100MB rewritten in larger SSTable (2nd time)
  5. Tier 2 compaction: 100MB rewritten again (3rd time)
  6. Total disk writes: 300MB for 100MB application data = 3x amplification

Write Amplification by Strategy:

Strategy Amplification Why
TWCS 2-5x ✅ Best Limited per-window compaction
STCS 5-10x Data rewritten multiple tiers
LCS 10-20x ❌ Worst Continuous level promotion

Why It Matters:

  • 💾 Disk Wear: SSDs have limited write endurance
  • 🔥 I/O Bandwidth: High amplification = saturated disks
  • ⚡ Write Performance: More disk writes = slower writes
  • 💰 Cloud Costs: Provisioned IOPS = pay per write

Example:

-- Application writes 1TB/day // TWCS (3x amplification): // Actual disk writes: 3TB/day // SSD lifespan: 5+ years ✅ // LCS (15x amplification): // Actual disk writes: 15TB/day // SSD lifespan: 1-2 years ❌

Key Takeaway: For write-heavy workloads, prefer TWCS or STCS over LCS to minimize write amplification and extend disk life!

4 How would you troubleshoot a Cassandra cluster where compaction is falling behind? ▼

Answer: Check pending compactions and SSTable count, then increase throughput, tune strategy parameters, or add nodes depending on root cause.

Step 1: Diagnose the Problem

-- Check compaction status nodetool compactionstats -- Red flags: /* pending tasks: 500+ ← CRITICAL! Way behind! Active compaction remaining time: 12h */ -- Check table stats nodetool tablestats keyspace.table -- Red flags: /* SSTable count: 847 ← CRITICAL! Should be < 50! Space used (total): 2.5TB Space used (live): 1.2TB ← 2x amplification! */

Step 2: Immediate Fixes

Fix 1: Increase Compaction Throughput

-- Default is 16 MB/sec - too conservative! nodetool setcompactionthroughput 64 -- or 128 -- Or in cassandra.yaml: compaction_throughput_mb_per_sec: 64

Fix 2: Increase Concurrent Compactors

-- In cassandra.yaml (requires restart): concurrent_compactors: 4 -- Default: # of disks -- Increase to 8-16 if you have CPU/disk capacity

Step 3: Tune Strategy Parameters

For STCS:

-- Lower threshold = more frequent compaction ALTER TABLE table_name WITH compaction = { 'class': 'SizeTieredCompactionStrategy', 'min_threshold': '4', -- Default, can lower to 2 'max_threshold': '32' -- Default };

Step 4: Long-Term Solutions

  • ✅ Wrong Strategy? Time-series with STCS → Switch to TWCS
  • ✅ Add Nodes: Distribute load across more nodes
  • ✅ Upgrade Disks: Faster SSDs = faster compaction
  • ✅ Archive Old Data: Reduce total data volume

Emergency: Force Major Compaction

-- Last resort: Force major compaction nodetool compact keyspace table -- ⚠️ WARNING: -- • Creates one huge SSTable -- • Takes hours/days on large tables -- • Needs 100% extra disk space -- • Only use if critically behind!

Key Metrics to Monitor:

  • Pending compactions (target: < 20)
  • SSTable count (target: < 50)
  • Space amplification (target: < 1.5x)
  • Compaction throughput (MB/sec)
5 What happens during a major compaction and why should you avoid it in production? ▼

Answer: Major compaction merges ALL SSTables into one giant SSTable. It's dangerous because it needs massive disk space, takes hours/days, and creates unbalanced data distribution.

What Happens During Major Compaction:

  1. Read All SSTables: Every SSTable on disk (could be 500+ files)
  2. Merge All Data: Combine into single sorted order
  3. Write Giant SSTable: One massive file (could be 500GB-1TB!)
  4. Delete Old SSTables: Remove all the original files
  5. Duration: Hours to days depending on data size

Why It's Dangerous:

Problem 1: Massive Disk Space Required

-- Example: 500GB of SSTables // During compaction: // • Old SSTables: 500GB (must keep until done) // • New SSTable: 500GB (being written) // • Total: 1000GB needed! (2x!) // If disk only 80% full → OUT OF SPACE! 💥

Problem 2: Breaks LCS Strategy

LCS carefully maintains levels. Major compaction destroys this structure:

  • Creates one huge SSTable (not 160MB files!)
  • Breaks non-overlapping guarantee
  • Takes days to rebuild proper level structure
  • Reads become slow until fixed

Problem 3: Unbalanced Distribution

One giant SSTable means:

  • All data locked in one file
  • Future compactions must process entire file
  • Harder to stream during repairs
  • Difficult to delete old data

When Major Compaction is OK:

  • ✅ One-time data cleanup after bulk delete
  • ✅ Pre-production testing on small datasets
  • ✅ Maintenance window with low traffic
  • ✅ Plenty of extra disk space (2x minimum)

Better Alternatives:

-- Instead of major compaction: -- 1. Increase compaction throughput nodetool setcompactionthroughput 128 -- 2. Let normal compaction catch up -- (Takes longer but safer!) -- 3. For TWCS: Just wait for TTL cleanup -- (Automatic space reclamation) -- 4. Add nodes to cluster -- (Distribute load)

Key Takeaway: Major compaction is a last resort. In production, rely on automatic compaction strategies and only force major compaction during scheduled maintenance windows!

🎓 Chapter Summary: Compaction Strategy Mastery

You now understand compaction strategies at a production level!

Quick Decision Guide:

  • ⏰ Time-series with TTL? → Use TWCS (saves 50% disk!)
  • 📚 Read-heavy with updates? → Use LCS (10x faster reads!)
  • 📦 Write-heavy, insert-only? → Use STCS (default, works great!)

Key Metrics:

  • ✅ Pending compactions < 20
  • ✅ SSTable count < 50
  • ✅ Space amplification < 1.5x

Remember Mike's lesson: Right strategy = Production success! 🚀

Advertisement

Responsive Ad