TOMBSTONES - Deletion Markers
ðŠĶ Master Cassandra's deletion mechanism with interactive examples, animations, and real-world scenarios. Learn from absolute basics!
ð Fundamentals (Start Here If You're Brand New!)
Before learning about tombstones, let's understand ALL the basic terms. No prior knowledge needed!
What is "Deletion"?
Deletion = Removing data
Like throwing away paper: You don't want it anymore
In regular databases:
âĒ Find the data
âĒ Erase it immediately
âĒ Gone forever!
In Cassandra (distributed):
âĒ Can't delete immediately!
âĒ Data lives on multiple servers
âĒ Need coordination
âĒ Use tombstones instead â
Why different? Multiple copies need to know!
What is a Tombstone?
Tombstone = Marker for deleted data
Like a gravestone: Shows something was here, now it's gone
Instead of erasing:
âĒ Cassandra writes a marker
âĒ Says "This data is deleted"
âĒ Keeps marker for a while
âĒ Eventually removes marker too
Why markers:
All servers need to see deletion
Marker spreads to all copies
Ensures everyone knows!
Distributed System
Distributed = Data on multiple computers
Example: 3 servers
âĒ Server A has copy 1
âĒ Server B has copy 2
âĒ Server C has copy 3
Problem with deletion:
âĒ Delete on Server A
âĒ But Server B offline!
âĒ When B comes back: still has data!
âĒ Data "resurrects" (zombie data!)
Solution: Tombstones prevent resurrection!
Replication
Replication = Making copies
Like photocopying a document
Why replicate:
âĒ Backup (if one server fails)
âĒ Faster reads (read from nearest)
âĒ No single point of failure
Replication Factor (RF):
âĒ RF=3: 3 copies of every row
âĒ Common in production
Deletion challenge:
Delete 1 copy? Other 2 remain!
Need to delete ALL copies!
Timestamp
Timestamp = When something happened
Like a clock reading
Cassandra uses microseconds:
âĒ Example: 1735488000000000
âĒ Very precise timing
Why important:
âĒ Determines order of operations
âĒ Delete at time 100
âĒ Write at time 90
âĒ Delete wins! (happened later)
Tombstone has timestamp:
Marks WHEN deletion happened!
Compaction
Compaction = Cleanup process
Like organizing a messy desk
What it does:
âĒ Merges multiple files
âĒ Removes old versions
âĒ Deletes expired tombstones
âĒ Frees up disk space
When tombstones removed:
âĒ After gc_grace_seconds (10 days)
âĒ During compaction
âĒ When safe to delete
Compaction is the janitor!
GC Grace Seconds
GC Grace = How long to keep tombstones
"Grace period" before removal
Default: 864000 seconds
= 864000 ÷ 60 ÷ 60 ÷ 24
= 10 days
Why 10 days:
âĒ Gives offline servers time
âĒ Server comes back online
âĒ Sees tombstone
âĒ Deletes its copy too!
After 10 days:
Tombstone itself deleted!
Zombie Data
Zombie Data = Deleted data comes back!
Like a horror movie - data won't stay dead
How it happens:
1. Delete data (Server A)
2. Server B offline (missed deletion)
3. Tombstone expires (after 10 days)
4. Server B comes back online
5. Repairs: B has data, A doesn't
6. Copies data back to A!
7. Zombie! Data resurrected! ðŧ
Prevention: Keep tombstones long enough!
SSTable (Data File)
SSTable = Sorted String Table
File on disk with data
Immutable = Never changes!
âĒ Once written, can't modify
âĒ Can only create new SSTables
Why tombstones needed:
âĒ Can't erase from SSTable
âĒ File is immutable!
âĒ Must write tombstone in NEW file
Compaction merges:
âĒ Old data + tombstone = deleted
âĒ New SSTable without data!
TTL (Time To Live)
TTL = Auto-delete after time
Like milk expiration date
Example:
INSERT ... USING TTL 3600
(Delete after 3600 seconds = 1 hour)
How it works:
âĒ Write data with TTL
âĒ After TTL expires
âĒ Automatically becomes tombstone!
âĒ No DELETE needed
Perfect for:
Session data, cache, temporary records
ð Quick Reference: All Terms
Now you know the vocabulary!
Core Concepts:
âĒ Deletion = Removing data
âĒ Tombstone = Marker for deleted data ðŠĶ
âĒ Distributed System = Data on multiple servers
âĒ Replication = Making copies (RF=3 typical)
Time & Process:
âĒ Timestamp = When something happened (microseconds)
âĒ Compaction = Cleanup process (merges files)
âĒ GC Grace Seconds = How long to keep tombstones (10 days)
âĒ TTL = Auto-delete after time (expiration)
Problems:
âĒ Zombie Data = Deleted data comes back ðŧ
âĒ SSTable = Immutable file (can't modify)
Ready to learn tombstones! ðŠĶ
ðĪ What are Tombstones?
Let's understand tombstones with a perfect real-world analogy!
ðïļ The Cemetery Analogy (Perfect Explanation!)
Imagine a city with 3 record offices in different locations:
WITHOUT tombstones (naive approach - BROKEN!):
âĒ Person "John Smith" exists in all 3 offices' records
âĒ John passes away
âĒ Office A removes his record immediately
âĒ But Office B is closed for renovation!
âĒ And Office C's computer is down!
Problem:
âĒ 2 weeks later, Office B reopens
âĒ They still have John's record (alive!)
âĒ They send updates to Office A: "Hey, we have John!"
âĒ Office A: "Oh, we must have lost it, thanks!"
âĒ John is "resurrected" in records! ðŧ
âĒ This is ZOMBIE DATA!
WITH tombstones (smart way - WORKS!):
âĒ Person "John Smith" exists in all 3 offices
âĒ John passes away
âĒ Office A doesn't erase the record!
âĒ Instead, they write: John Smith - Deceased Dec 2025
âĒ This is a TOMBSTONE!
What happens:
âĒ Office B reopens
âĒ Syncs with Office A
âĒ Sees: "John Smith - Deceased"
âĒ Marks their record as deceased too!
âĒ Office C comes back online
âĒ Also gets the tombstone marker
âĒ All 3 offices now know John is deceased â
After 10 days (gc_grace_seconds):
âĒ All offices have seen the tombstone
âĒ Safe to remove the tombstone marker itself
âĒ Record finally completely removed
âĒ But NO resurrection possible!
Key insight:
Tombstone = Death certificate
âĒ Proves someone is deceased
âĒ Prevents "coming back to life"
âĒ Spreads to all record offices
âĒ Eventually removed when no longer needed
This is EXACTLY how Cassandra works!
Definition
Tombstone = Special marker for deleted data
What it is:
âĒ NOT actual data deletion
âĒ A marker that says "deleted"
âĒ Has timestamp (when deleted)
âĒ Stored like normal data
Looks like:
Key: "user123"
Value: <TOMBSTONE>
Timestamp: 1735488000
Think: Gravestone marking deleted data!
What It Does
Tombstone overrides old data
When query reads:
âĒ Finds data: "John" age 30
âĒ Also finds: TOMBSTONE
âĒ Tombstone newer? Return nothing â
âĒ Data newer? Return data
Timestamp wins:
âĒ Data written at time 100
âĒ Tombstone at time 200
âĒ Tombstone wins! (newer)
Prevents seeing deleted data!
Lifecycle
Tombstone has a lifespan:
Phase 1: Creation (instant)
âĒ DELETE executed
âĒ Tombstone written
Phase 2: Active (10 days)
âĒ Spreads to all replicas
âĒ Blocks old data
âĒ Participates in queries
Phase 3: Removal (compaction)
âĒ After gc_grace_seconds
âĒ Compaction runs
âĒ Tombstone deleted!
Total: ~10 days existence
â Why Are Tombstones Needed?
Understanding the distributed systems challenge that tombstones solve!
The Distributed Deletion Problem
Scenario: You have data replicated on 3 servers (RF=3)
Initial State:
âĒ Server A: user123 = "Alice"
âĒ Server B: user123 = "Alice"
âĒ Server C: user123 = "Alice"
All 3 servers have the same data â
â Attempt 1: Direct Deletion (BROKEN!)
âĒ Client sends: DELETE FROM users WHERE id='user123'
âĒ DELETE reaches Server A (deletes immediately)
âĒ Server B is down for maintenance ð§
âĒ Server C has network issue ð
Current state:
âĒ Server A: (nothing - deleted) â
âĒ Server B: user123 = "Alice" (still has it!)
âĒ Server C: user123 = "Alice" (still has it!)
2 days later - Server B comes back:
âĒ Repair process runs (anti-entropy)
âĒ B compares with A: "I have user123, you don't!"
âĒ B thinks A lost the data
âĒ B sends data to A: "Here, I'll restore it for you!"
âĒ ZOMBIE! user123 = "Alice" is back on A! ðŧ
This is catastrophic!
âĒ Deleted data reappears
âĒ GDPR compliance violated (can't guarantee deletion)
âĒ Financial records reappear (audit nightmare)
âĒ Deleted user accounts come back
â
Solution: Tombstones (WORKS!)
âĒ Client sends: DELETE FROM users WHERE id='user123'
âĒ Server A writes: user123 = TOMBSTONE(timestamp: 100)
âĒ Server B down (missed it)
âĒ Server C network issue (missed it)
Current state:
âĒ Server A: user123 = TOMBSTONE(100) â
âĒ Server B: user123 = "Alice" (timestamp: 50)
âĒ Server C: user123 = "Alice" (timestamp: 50)
When Server B comes back:
âĒ Repair process runs
âĒ B sees: A has TOMBSTONE(100), I have Alice(50)
âĒ Timestamp comparison: 100 > 50
âĒ B adopts tombstone: "Oh, this was deleted!"
âĒ B writes: user123 = TOMBSTONE(100)
âĒ Server C also gets tombstone
âĒ All 3 servers now know data is deleted! â
After 10 days (gc_grace_seconds):
âĒ All servers have seen the tombstone
âĒ Compaction removes tombstone
âĒ Data truly gone
âĒ No resurrection possible!
Key Insight: Tombstones convert "absence of data" into "presence of deletion marker" - something that can be replicated and wins timestamp battles!
Timestamp Conflicts
Tombstones use timestamps to win!
Example conflict:
âĒ Server A: DELETE at time 100
âĒ Server B: WRITE at time 90
Who wins?
âĒ Tombstone(100) vs Data(90)
âĒ 100 > 90
âĒ Tombstone wins! â
âĒ Data is deleted
Reverse scenario:
âĒ Server A: DELETE at time 100
âĒ Server B: WRITE at time 110
âĒ Data(110) wins!
âĒ Deletion overridden (intended!)
Last Write Wins (LWW) principle!
Immutable SSTables
SSTables can't be modified!
Problem:
âĒ Data written to SSTable
âĒ File closed (immutable)
âĒ Can't go back and erase!
Solution:
âĒ Write tombstone in NEW SSTable
âĒ During compaction:
- Old SSTable: user123 = "Alice"
- New SSTable: user123 = TOMBSTONE
- Merged result: (deleted)
Tombstones work with immutability!
Eventual Consistency
Deletions take time to propagate
Timeline:
âĒ T=0: DELETE executed
âĒ T=1ms: Tombstone on Server A
âĒ T=100ms: Replicated to Server B
âĒ T=5min: Server C comes online
âĒ T=10min: All servers have it â
During propagation:
âĒ Queries might still see data
âĒ Tombstone eventually wins
âĒ Consistent eventually!
This is "eventual consistency"
ðïļ Types of Tombstones
Cassandra has 4 different types of tombstones!
1. Cell Tombstone
Deletes a single column value
Query:
DELETE email FROM users WHERE id='user123'
What happens:
âĒ Only 'email' column deleted
âĒ Other columns remain
âĒ Tombstone for one cell only
Before:
user123: {name: "Alice", email: "alice@email.com"}
After:
user123: {name: "Alice", email: DELETED}
Most granular tombstone type!
2. Row Tombstone
Deletes entire row
Query:
DELETE FROM users WHERE id='user123'
What happens:
âĒ Entire row deleted
âĒ All columns gone
âĒ Single tombstone for whole row
Before:
user123: {name: "Alice", email: "alice@email.com", age: 30}
After:
user123: ROW DELETED
Most common tombstone type!
3. Partition Tombstone
Deletes entire partition
Query:
DELETE FROM events WHERE user_id='user123'
(Partition key = user_id)
What happens:
âĒ All rows in partition deleted
âĒ Could be thousands of rows!
âĒ Single partition tombstone
Example:
âĒ user123 has 10,000 events
âĒ One tombstone deletes all!
âĒ Very efficient â
Efficient for large deletions!
4. Range Tombstone
Deletes range of clustering keys
Query:
DELETE FROM events
WHERE user_id='user123'
AND timestamp >= '2024-01-01'
AND timestamp < '2024-02-01'
What happens:
âĒ Deletes date range
âĒ One tombstone for whole range
âĒ Marks startâend as deleted
Example:
Range: Jan 1 - Feb 1
Could be 1000s of rows
Single range tombstone!
Perfect for time-series data!
5. TTL Tombstone
Auto-created when TTL expires
Query:
INSERT INTO sessions (id, data)
VALUES ('session123', 'data')
USING TTL 3600; -- 1 hour
What happens:
âĒ After 1 hour passes
âĒ Automatically becomes tombstone
âĒ No DELETE needed!
Perfect for:
âĒ Session data
âĒ Cache entries
âĒ Temporary records
Set-it-and-forget-it deletion!
Tombstone Sizes
Different types, different overhead:
Cell Tombstone: ~10 bytes
âĒ Column name + timestamp
Row Tombstone: ~20 bytes
âĒ Partition key + timestamp
Partition Tombstone: ~15 bytes
âĒ Partition key only
âĒ Deletes 1000s of rows!
Range Tombstone: ~30 bytes
âĒ Start + end range + timestamp
Insight: Tombstones are tiny but powerful!
ðŪ Interactive Tombstone Simulator - Try It!
See tombstone creation and zombie prevention in action!
Simulation scenario:
âĒ 3 Cassandra servers (A, B, C) with RF=3
âĒ GC grace seconds: 10 days (simplified to 5 operations)
Try this sequence:
1. Insert Data â Creates user on all servers
2. Server B Goes Down â Simulates offline node
3. Delete Data â Creates tombstone (B misses it!)
4. Server B Comes Back â Watch zombie prevention!
5. Run Compaction â After grace period expires
ðŊ Goal: See how tombstones prevent zombie data resurrection!
ðĄ What You're Learning
This simulator demonstrates the critical zombie data problem:
WITHOUT tombstones:
1. Delete data on Server A
2. Server B is offline (misses deletion)
3. B comes back, sees it has data A doesn't
4. B "helpfully" restores data to A
5. Zombie! Deleted data is back! ðŧ
WITH tombstones (what you'll see):
1. Delete creates tombstone on A & C
2. Server B offline (misses tombstone)
3. B comes back with old data
4. Repair: B sees tombstone from A/C
5. Tombstone wins (newer timestamp)
6. B adopts tombstone
7. No resurrection! Deletion preserved! â
After grace period:
âĒ All servers have seen tombstone
âĒ Compaction removes it safely
âĒ Data truly gone forever
Key insight: Tombstones turn "absence" into "presence of deletion marker" that can be replicated and wins conflicts!
âïļ How Tombstones Work (Complete Flow)
Let's trace the complete tombstone lifecycle!
â° GC Grace Seconds (The Magic Number)
What It Is
gc_grace_seconds = Tombstone lifespan
Default: 864,000 seconds
âĒ 864000 ÷ 60 = 14,400 minutes
âĒ 14400 ÷ 60 = 240 hours
âĒ 240 ÷ 24 = 10 days
Purpose:
âĒ Time for offline servers
âĒ Come back and see deletion
âĒ Prevents zombie data
After 10 days: safe to remove!
Why 10 Days?
Assumptions about server downtime:
Typical scenarios:
âĒ Planned maintenance: 1-4 hours
âĒ Network issues: Minutes to hours
âĒ Hardware failure: 1-3 days
âĒ Catastrophic failure: 3-7 days
10 days provides:
âĒ Plenty of buffer â
âĒ Time for weekend repairs
âĒ Margin for errors
Conservative but safe!
Configuration
Set per table:
View current:
DESC TABLE users;
Change value:
ALTER TABLE users
WITH gc_grace_seconds = 86400;
(1 day = 86400 seconds)
Common values:
âĒ 10 days: Default (safe)
âĒ 1 day: Fast-moving data
âĒ 0: Special cases only!
Too Short = Danger!
Setting gc_grace_seconds = 0:
What happens:
âĒ Tombstone removed immediately
âĒ Server offline for 1 hour
âĒ Tombstone already gone!
âĒ Server comes back with old data
âĒ Zombie data! ðŧ
Only use 0 when:
âĒ Single node cluster (testing)
âĒ Never run repair
âĒ Understand the risks!
Generally: DON'T DO THIS!
Too Long = Problems!
Setting gc_grace_seconds = 30 days:
What happens:
âĒ Tombstones live 30 days
âĒ Accumulate on disk
âĒ Slow down queries
âĒ Waste disk space
Problems:
âĒ Many tombstones in reads
âĒ Performance degradation
âĒ Larger SSTables
âĒ Slower compaction
Balance safety vs performance!
Best Practice
Follow these guidelines:
Default (10 days):
âĒ Most tables â
âĒ Safe and proven
Shorter (1-3 days):
âĒ Stable cluster
âĒ Frequent deletions
âĒ Monitor closely
Never use 0 unless:
âĒ Testing only
âĒ Single node
âĒ You're an expert
When in doubt: keep default!
â ïļ Tombstone Problems (Performance Impact)
The Tombstone Accumulation Problem
Worst-case scenario: Millions of tombstones!
Example: Time-series data with TTL
âĒ Table: sensor_data (temperature readings)
âĒ Inserts: 1 million rows/day with TTL=7 days
âĒ After 7 days: 1M rows become tombstones
âĒ After 10 days (gc_grace): Tombstones eligible for removal
âĒ But compaction hasn't run yet!
State after 10 days:
âĒ Active data: 7M rows (7 days à 1M/day)
âĒ Tombstones: 3M+ (days 8-10)
âĒ Ratio: 30% tombstones!
When you query:
âĒ Must read ALL tombstones
âĒ Filter them out
âĒ Return only live data
âĒ Massive overhead!
Performance impact:
âĒ Query scans 10M rows (7M live + 3M tombstones)
âĒ 43% waste (3M ÷ 10M)
âĒ P99 latency: 50ms â 200ms (4x slower!)
âĒ Disk I/O: Unnecessary reads
âĒ Memory: Tombstones in bloom filters
âĒ Network: Tombstones sent between nodes
Warning threshold:
Cassandra warns at 100,000 tombstones in a single query!
WARN [ReadStage-2] Read 250000 live rows and 500000 tombstone cells for query SELECT * FROM sensor_data LIMIT 1000This means: You have a problem!
1. Read Performance
Tombstones slow queries!
Why:
âĒ Must read tombstones from disk
âĒ Process and filter them
âĒ Then return live data
Example query:
SELECT * FROM users LIMIT 100
Without tombstones:
âĒ Read 100 rows (5ms)
With 10k tombstones:
âĒ Read 10,000 tombstones (100ms)
âĒ Read 100 live rows (5ms)
âĒ Total: 105ms (21x slower!)
More tombstones = slower queries!
2. Disk Space
Tombstones consume space!
Example table:
âĒ 1B rows deleted
âĒ Each tombstone: 20 bytes
âĒ Total: 20GB of tombstones!
Problem:
âĒ Disk fills up
âĒ Can't add new data
âĒ Expensive storage costs
Solution:
âĒ Run compaction
âĒ Remove expired tombstones
âĒ Free up space â
3. Compaction Overhead
Many tombstones = slow compaction!
Why:
âĒ Must process all tombstones
âĒ Check if expired
âĒ Merge with data
Example:
âĒ 100GB data
âĒ 50GB tombstones
âĒ Compaction processes 150GB
âĒ Takes 3x longer!
Impact:
âĒ Higher CPU usage
âĒ More disk I/O
âĒ Slower overall system
ðĒ Real Company Examples (Production Stories)
A. Discord: TTL Tombstone Disaster (500M Tombstones!)
The Problem:
Discord stores message read status with TTL=30 days. Every time a user reads a message, write with TTL. 150M users à 1000 messages/day = 150 billion operations. After 30 days, massive tombstone accumulation!
What went wrong:
âĒ Table: user_read_status (user_id, channel_id, message_id, TTL 30 days)
âĒ Writes: 150 billion/month
âĒ After TTL expires: All become tombstones
âĒ gc_grace_seconds: 10 days
âĒ Total tombstone period: 30 days (TTL) + 10 days (grace) = 40 days
âĒ Tombstone accumulation: 150B à 40/30 = 200 billion tombstones!
Impact:
âĒ Query: "Get unread messages for user"
âĒ Scanned 50M tombstones per query
âĒ P99 latency: 15ms â 3000ms (200x slower!)
âĒ Warnings: "Read 100000 tombstones" constantly
âĒ Disk usage: 2TB of tombstones (50% of cluster!)
âĒ Compaction: Couldn't keep up
Solution:
1. Reduced TTL: 30 days â 7 days (less accumulation)
2. Reduced gc_grace: 10 days â 1 day (faster removal)
3. Aggressive compaction: Forced major compaction
4. Rewrote approach: Moved to time-bucketed partitions
- New partition each day
- Drop entire partition after 7 days (no tombstones!)
Results after fix:
âĒ Tombstone count: 200B â 2M (99.999% reduction!)
âĒ P99 latency: 3000ms â 18ms (167x faster!)
âĒ Disk usage: 2TB freed
âĒ Warnings: Gone â
Lesson learned:
"TTL creates tombstones! High-volume TTL writes = tombstone explosion. Time-bucketed partitions + partition drops avoid tombstones entirely. This saved our cluster from collapsing."
B. Instagram: Deletion Pattern Optimization
Use case:
User deletes photos/posts. 2B users, average 500 photos each. Deletion rate: 0.1% daily = 1M photo deletions/day. Need efficient deletion without tombstone accumulation.
Initial approach (WRONG):
âĒ DELETE FROM user_photos WHERE user_id=X AND photo_id=Y
âĒ Created cell tombstone for each photo
âĒ 1M deletions/day = 1M tombstones/day
âĒ After 10 days (gc_grace): 10M tombstones
Problem:
âĒ Query: "Get user's photos"
âĒ Some users deleted 1000s of photos
âĒ Scan 1000s of tombstones per user
âĒ P99: 50ms â 400ms (8x slower!)
Optimized approach:
1. Soft delete instead of hard delete:
- UPDATE photos SET deleted=true, deleted_at=now()
- No tombstone created! â
- Query: WHERE deleted=false
2. Async cleanup job:
- Nightly job processes deleted=true
- Batches deletions into range tombstones
- One range tombstone for 1000 photos!
3. Short gc_grace for cleanup table:
- gc_grace_seconds = 86400 (1 day)
- Rapid tombstone removal
Results:
âĒ Tombstones per query: 1000s â 0 (live queries see no tombstones!)
âĒ P99 latency: 400ms â 35ms (11x faster!)
âĒ Cleanup efficiency: 1000 photos = 1 range tombstone
âĒ Compaction load: 90% reduction
Instagram engineer quote:
"Never delete in the hot path! Soft delete + async cleanup = best of both worlds. Users see instant deletion (deleted=true), but actual tombstones created in background batches. This keeps read queries fast while still removing data."
C. Uber: Time-Series gc_grace Tuning
Scenario:
Trip location history - massive time-series writes. Table stores GPS coordinates every 4 seconds during trips. TTL=30 days for GDPR compliance. 20M trips/day à 450 points/trip à 30 days = 270 billion rows!
Challenge:
âĒ After 30 days: 9B rows/day expire â tombstones
âĒ gc_grace_seconds: 864000 (10 days default)
âĒ Tombstone accumulation: 9B/day à 10 days = 90 billion tombstones!
âĒ Queries scanning expired data regions: massive tombstone overhead
Initial problem:
âĒ P99 read latency: 25ms â 800ms
âĒ Disk usage: 80% tombstones
âĒ Warning logs: 50k/day about tombstones
âĒ Compaction: Running 24/7, still couldn't catch up
Uber's solution - Aggressive tuning:
1. Reduced gc_grace_seconds: 10 days â 1 day
- Rationale: "Our cluster is stable, nodes rarely down >24h"
- Monitoring: Alert if node down >12 hours
- Risk accepted: Worth it for performance
2. Increased repair frequency:
- Run repair every 12 hours (instead of weekly)
- Ensures deletions propagate quickly
3. Partition strategy:
- Partition by (trip_id, day_bucket)
- Old partitions naturally age out
4. Compaction tuning:
- More frequent minor compactions
- Aggressive tombstone_threshold: 0.1 (vs 0.2 default)
Results:
âĒ Tombstone lifespan: 10 days â 1 day (90% reduction)
âĒ Tombstone count: 90B â 9B (10x less!)
âĒ P99 latency: 800ms â 45ms (18x faster!)
âĒ Disk usage: 80% â 15% tombstones
âĒ Compaction: Caught up within 48 hours
Trade-off:
âĒ Risk: If node down >24 hours, possible zombie data
âĒ Mitigation: Monitoring + alerts + 12h repair cycles
âĒ Reality: In 2 years of production, zero zombie incidents
Key insight:
"Default gc_grace_seconds (10 days) is conservative. For stable clusters with good monitoring, 1-3 days is often better. The performance gain from faster tombstone removal outweighs the small zombie risk. But: requires operational excellence!"
â Best Practices (Production-Ready Tips!)
1. Monitor Tombstone Warnings
Set up monitoring!
Watch for:
âĒ "Read X tombstones" warnings
âĒ Threshold: 100,000 tombstones
Alert when:
âĒ >10 warnings/hour
âĒ Consistent growth
Check with nodetool:
nodetool tablestats keyspace.table
(Look for "Tombstoned cells")
Early detection prevents disasters!
2. Use Time-Bucketed Partitions
For time-series data:
Instead of TTL:
â USING TTL 86400
(Creates tombstones)
Use partitions:
â Partition key: (sensor_id, day)
âĒ One partition per day
âĒ Drop entire partition when old
âĒ No tombstones created!
Example:
DROP TABLE sensor_data_20241201;
(Deletes whole partition instantly)
3. Tune gc_grace Carefully
Guidelines by scenario:
Stable cluster (3 data centers):
âĒ gc_grace: 3-5 days â
âĒ Aggressive repair schedule
Volatile environment:
âĒ gc_grace: 10 days (default) â
âĒ Conservative, safe
High-delete workload:
âĒ gc_grace: 1-2 days
âĒ Fast tombstone removal
âĒ Requires monitoring!
Never use 0 in production!
4. Soft Delete When Possible
Instead of DELETE:
UPDATE SET deleted=true
Benefits:
âĒ No tombstone in hot path!
âĒ Instant for user
âĒ Query: WHERE deleted=false
Async cleanup:
âĒ Background job
âĒ Batch actual DELETEs
âĒ Range tombstones
Perfect for:
User-facing deletions!
5. Regular Compaction
Ensure compaction runs!
Check:
nodetool compactionstats
If tombstones accumulate:
âĒ Force compaction:
nodetool compact keyspace table
Tune strategy:
âĒ tombstone_threshold: 0.2â0.1
âĒ More aggressive removal
Monitor:
âĒ Pending compactions
âĒ Should be near 0
6. Avoid Frequent Range Deletes
Range deletes create range tombstones:
Example (BAD):
DELETE FROM events
WHERE user_id=X
AND time < '2024-01-01'
Creates:
âĒ Large range tombstone
âĒ Affects all queries in range
Better:
âĒ Use time-bucketed partitions
âĒ Drop old partitions
âĒ Or design for no deletes!
ðĻ Common Mistakes to Avoid
â Setting gc_grace_seconds = 0
â Zombie data guaranteed! Only for testing.
â Production requires time for offline nodes
â Using TTL without understanding tombstones
â High-volume TTL writes = tombstone explosion
â Consider time-bucketed partitions instead
â Monitor tombstone warnings closely
â Ignoring tombstone warnings
â "Read 100000 tombstones" = serious problem
â Performance degrading, fix immediately!
â Don't wait until cluster fails
â Frequent small range deletes
â Creates many range tombstones
â Slows all queries in that range
â Design around deletions when possible
â Not monitoring compaction
â Compaction removes tombstones
â If not running: tombstones accumulate forever
â Check nodetool compactionstats regularly
â Deleting in the hot path
â User clicks "delete", waits for tombstone
â Use soft delete + async cleanup instead
â Better user experience + fewer live tombstones
â Same gc_grace for all tables
â Different tables have different needs
â High-delete tables: shorter gc_grace
â Rarely-deleted tables: longer is fine
â Remember: Tombstones are necessary but need management!
ðž Interview Questions & Answers
Complete Answer:
Definition:
Tombstones are special deletion markers that Cassandra writes instead of immediately removing data. They are metadata entries with a timestamp indicating when data was deleted, stored alongside regular data in SSTables.
Why tombstones exist (the fundamental problem):
In a distributed system with data replicated across multiple nodes, you cannot simply erase data immediately because:
- Nodes can be offline: If you delete on Server A but Server B is down for maintenance, B never sees the deletion. When B comes back, it still has the old data and will "restore" it to A through anti-entropy repair, causing zombie data resurrection.
- Immutable SSTables: Cassandra's SSTables are immutable - once written, they cannot be modified. You can't go back and erase data from an existing file.
- Eventual consistency: Operations don't happen simultaneously across all replicas. Deletions take time to propagate.
How tombstones solve this:
Instead of erasing data, Cassandra writes a tombstone marker that says "this data was deleted at timestamp X". This marker:
- Propagates to all replicas just like regular data
- Wins timestamp conflicts (if tombstone timestamp > data timestamp)
- Prevents zombie data by explicitly marking deletion
- Eventually gets removed after gc_grace_seconds (default 10 days)
Example scenario:
- Data: user123="Alice" exists on Servers A, B, C (timestamp: 100)
- Server B goes offline for maintenance
- Client: DELETE user123 â Creates tombstone on A & C (timestamp: 200)
- Server B comes back 2 days later, still has user123="Alice"(100)
- Repair process: B sees tombstone(200) from A/C vs data(100)
- Timestamp comparison: 200 > 100 â Tombstone wins
- B adopts tombstone, deletion preserved, no zombie!
When tombstones are removed:
After gc_grace_seconds (default 864,000 seconds = 10 days), compaction removes expired tombstones. This grace period ensures even offline nodes have time to see the deletion before the marker disappears.
Key takeaway for interview:
"Tombstones solve the distributed deletion problem by converting 'absence of data' into 'presence of a deletion marker' that can be replicated and wins timestamp conflicts. They prevent zombie data resurrection in distributed systems where nodes can be temporarily offline."
Complete Answer:
Definition:
gc_grace_seconds is a table-level setting that specifies how long Cassandra keeps tombstones before they're eligible for removal during compaction. The default is 864,000 seconds (10 days).
Purpose:
This grace period gives offline or partitioned nodes time to come back online and see the deletion tombstone before it's removed. It's insurance against zombie data resurrection.
How it works:
- Tombstone created with deletion timestamp T
- Current time C is tracked during compaction
- If (C - T) > gc_grace_seconds, tombstone is eligible for removal
- Compaction removes eligible tombstones
Setting it TOO LOW (dangerous!):
Example: gc_grace_seconds = 0 (worst case)
- What happens:
- Tombstones removed immediately during compaction
- Node goes offline for 1 hour
- Tombstone already deleted from other nodes
- Node comes back with old data, no tombstone to stop it
- Result: Zombie data! Deleted data resurrects
- Scenarios this can happen:
- Network partition (nodes can't communicate)
- Hardware failure requiring replacement
- Extended maintenance window
- Data center failure
- Real impact: GDPR compliance violation (can't guarantee deletion), financial records reappear (audit nightmare), deleted user accounts come back
When you might use shorter gc_grace:
- Stable, well-monitored cluster
- Frequent repair runs (every 12-24 hours)
- High tombstone accumulation issues
- Example: 1-3 days for production with excellent ops
- But never 0 in production!
Setting it TOO HIGH (performance issues!):
Example: gc_grace_seconds = 30 days or more
- What happens:
- Tombstones accumulate for 30 days
- High-delete workload: millions of tombstones
- Queries must read and filter all tombstones
- Result: Severe performance degradation
- Specific problems:
- Read latency: Scanning 100,000+ tombstones per query
- Disk space: Tombstones consuming GBs of storage
- Compaction overhead: Processing huge volumes
- Memory pressure: Bloom filters include tombstones
- Warning signs: Logs showing "Read 100000 tombstones", P99 latency spikes, disk usage growing despite deletions
The 10-day default rationale:
- Covers typical failure scenarios (1-7 days)
- Includes weekends for manual intervention
- Provides safety margin for unexpected issues
- Balances safety vs. performance
Tuning guidelines by scenario:
- Default workload: Keep 10 days (864,000 sec)
- High-delete, stable cluster: 1-3 days (86,400-259,200 sec)
- Rarely deleted data: 10-14 days is fine
- Time-series with TTL: Consider shorter (1-2 days) with aggressive repair
Monitoring requirements:
If using shorter gc_grace_seconds, you MUST:
- Monitor node downtime (alert if >50% of gc_grace)
- Run repair frequently (at minimum every gc_grace_seconds)
- Have operational excellence and rapid response
- Accept small zombie risk for performance gain
Key takeaway for interview:
"gc_grace_seconds is the tombstone lifespan - too short risks zombie data (if nodes offline longer), too long accumulates tombstones and degrades performance. The 10-day default balances safety and performance. Tuning lower requires operational excellence and frequent repairs. The core trade-off: safety of data deletion vs. performance overhead of keeping tombstones."
Complete Answer:
Performance impacts:
1. Read performance degradation:
- Problem: Queries must read tombstones from disk, process them, filter them out, then return live data
- Example: SELECT * FROM users LIMIT 100
- Without tombstones: Read 100 rows (5ms)
- With 10,000 tombstones: Read 10,000 tombstones (100ms) + 100 rows (5ms) = 105ms (21x slower!)
- Warning threshold: Cassandra warns at 100,000 tombstones per query
- Severe case: Millions of tombstones can make queries timeout entirely
2. Disk space consumption:
- Each tombstone: ~10-30 bytes
- 1 billion tombstones = 10-30 GB disk space
- Can represent 50-80% of total data size in high-delete workloads
- Disk fills up, preventing new writes
3. Compaction overhead:
- Compaction must process all tombstones
- Check expiration, merge with data
- High tombstone volume â slower compaction â tombstones accumulate further (vicious cycle)
- CPU and I/O overhead even when tombstones just sitting there
4. Memory pressure:
- Bloom filters include tombstones
- Larger bloom filters = more memory
- Key cache overhead for tombstone partition keys
Mitigation strategies:
Strategy 1: Avoid tombstones entirely (best!):
- Time-bucketed partitions:
- Instead of: TTL on rows
- Use: Partition per time bucket (day/week/month)
- DROP entire partition when old
- No tombstones created!
- Example: sensor_data_20241201, sensor_data_20241202, etc.
- Design without deletions:
- Append-only architecture where possible
- Soft deletes (deleted=true flag) instead of hard deletes
- Filter at application layer
Strategy 2: Soft delete + async cleanup:
- User action: UPDATE SET deleted=true (no tombstone!)
- Queries: WHERE deleted=false (never see deleted rows)
- Background job: Batch actual DELETEs into range tombstones
- Benefit: Hot path (user queries) never sees tombstones
- Used by Instagram for photo deletions
Strategy 3: Optimize gc_grace_seconds:
- Reduce from 10 days to 1-3 days for stable clusters
- Tombstones removed 3-10x faster
- Requires: Frequent repair (every 12-24h), good monitoring, operational discipline
- Trade-off: Small zombie risk for big performance gain
Strategy 4: Aggressive compaction tuning:
- tombstone_threshold: Lower from 0.2 to 0.1
- Triggers compaction when 10% tombstones (vs 20% default)
- More frequent compaction = faster removal
- unchecked_tombstone_compaction: Enable to compact tombstone-heavy SSTables even without triggering normal compaction
- Manual intervention: Force compact when needed: nodetool compact keyspace table
Strategy 5: Monitoring and alerting:
- Alert on tombstone warnings in logs
- Track tombstone ratio: nodetool tablestats
- Monitor compaction pending tasks
- Query latency correlation with tombstone counts
- Early detection prevents severe issues
Strategy 6: Query optimization:
- Avoid full partition scans when many tombstones
- Add LIMIT to bound tombstone scanning
- Use partition-restricted queries
- Filter at application when possible
Real-world example (Discord):
- Problem: 150B TTL writes/month â 200B tombstones, queries reading 50M tombstones, P99: 15ms â 3000ms (200x slower!)
- Solution:
- Reduced TTL: 30 â 7 days
- Reduced gc_grace: 10 â 1 day
- Rewrote to time-bucketed partitions (drop partitions instead of TTL)
- Results: Tombstones: 200B â 2M (99.999% reduction!), P99: 3000ms â 18ms (167x faster!)
When to use each strategy:
- Time-series data: Time-bucketed partitions (avoid tombstones entirely)
- User-facing deletes: Soft delete + async cleanup
- High-volume deletes: Reduce gc_grace + aggressive compaction
- Any workload: Monitoring is mandatory
Key takeaway for interview:
"Tombstones impact performance by requiring reads, processing, and storage while providing no value. Best mitigation: avoid them entirely through time-bucketed partitions or soft deletes. When unavoidable: reduce gc_grace_seconds (with operational excellence), tune compaction aggressively, and monitor closely. The pattern: prevention > optimization > monitoring."
Responsive Ad