Tombstones

TOMBSTONES - Deletion Markers

ðŸŠĶ Master Cassandra's deletion mechanism with interactive examples, animations, and real-world scenarios. Learn from absolute basics!

📖 Fundamentals (Start Here If You're Brand New!)

Before learning about tombstones, let's understand ALL the basic terms. No prior knowledge needed!

🗑ïļ

What is "Deletion"?

Deletion = Removing data
Like throwing away paper: You don't want it anymore

In regular databases:
â€Ē Find the data
â€Ē Erase it immediately
â€Ē Gone forever!

In Cassandra (distributed):
â€Ē Can't delete immediately!
â€Ē Data lives on multiple servers
â€Ē Need coordination
â€Ē Use tombstones instead ✓

Why different? Multiple copies need to know!

ðŸŠĶ

What is a Tombstone?

Tombstone = Marker for deleted data
Like a gravestone: Shows something was here, now it's gone

Instead of erasing:
â€Ē Cassandra writes a marker
â€Ē Says "This data is deleted"
â€Ē Keeps marker for a while
â€Ē Eventually removes marker too

Why markers:
All servers need to see deletion
Marker spreads to all copies
Ensures everyone knows!

🌐

Distributed System

Distributed = Data on multiple computers

Example: 3 servers
â€Ē Server A has copy 1
â€Ē Server B has copy 2
â€Ē Server C has copy 3

Problem with deletion:
â€Ē Delete on Server A
â€Ē But Server B offline!
â€Ē When B comes back: still has data!
â€Ē Data "resurrects" (zombie data!)

Solution: Tombstones prevent resurrection!

🔄

Replication

Replication = Making copies
Like photocopying a document

Why replicate:
â€Ē Backup (if one server fails)
â€Ē Faster reads (read from nearest)
â€Ē No single point of failure

Replication Factor (RF):
â€Ē RF=3: 3 copies of every row
â€Ē Common in production

Deletion challenge:
Delete 1 copy? Other 2 remain!
Need to delete ALL copies!

⏰

Timestamp

Timestamp = When something happened
Like a clock reading

Cassandra uses microseconds:
â€Ē Example: 1735488000000000
â€Ē Very precise timing

Why important:
â€Ē Determines order of operations
â€Ē Delete at time 100
â€Ē Write at time 90
â€Ē Delete wins! (happened later)

Tombstone has timestamp:
Marks WHEN deletion happened!

🔧

Compaction

Compaction = Cleanup process
Like organizing a messy desk

What it does:
â€Ē Merges multiple files
â€Ē Removes old versions
â€Ē Deletes expired tombstones
â€Ē Frees up disk space

When tombstones removed:
â€Ē After gc_grace_seconds (10 days)
â€Ē During compaction
â€Ē When safe to delete

Compaction is the janitor!

âģ

GC Grace Seconds

GC Grace = How long to keep tombstones
"Grace period" before removal

Default: 864000 seconds
= 864000 ÷ 60 ÷ 60 ÷ 24
= 10 days

Why 10 days:
â€Ē Gives offline servers time
â€Ē Server comes back online
â€Ē Sees tombstone
â€Ē Deletes its copy too!

After 10 days:
Tombstone itself deleted!

ðŸ‘ŧ

Zombie Data

Zombie Data = Deleted data comes back!
Like a horror movie - data won't stay dead

How it happens:
1. Delete data (Server A)
2. Server B offline (missed deletion)
3. Tombstone expires (after 10 days)
4. Server B comes back online
5. Repairs: B has data, A doesn't
6. Copies data back to A!
7. Zombie! Data resurrected! ðŸ‘ŧ

Prevention: Keep tombstones long enough!

ðŸ“Ķ

SSTable (Data File)

SSTable = Sorted String Table
File on disk with data

Immutable = Never changes!
â€Ē Once written, can't modify
â€Ē Can only create new SSTables

Why tombstones needed:
â€Ē Can't erase from SSTable
â€Ē File is immutable!
â€Ē Must write tombstone in NEW file

Compaction merges:
â€Ē Old data + tombstone = deleted
â€Ē New SSTable without data!

ðŸ”Ē

TTL (Time To Live)

TTL = Auto-delete after time
Like milk expiration date

Example:
INSERT ... USING TTL 3600
(Delete after 3600 seconds = 1 hour)

How it works:
â€Ē Write data with TTL
â€Ē After TTL expires
â€Ē Automatically becomes tombstone!
â€Ē No DELETE needed

Perfect for:
Session data, cache, temporary records

🎓 Quick Reference: All Terms

Now you know the vocabulary!

Core Concepts:
â€Ē Deletion = Removing data
â€Ē Tombstone = Marker for deleted data ðŸŠĶ
â€Ē Distributed System = Data on multiple servers
â€Ē Replication = Making copies (RF=3 typical)

Time & Process:
â€Ē Timestamp = When something happened (microseconds)
â€Ē Compaction = Cleanup process (merges files)
â€Ē GC Grace Seconds = How long to keep tombstones (10 days)
â€Ē TTL = Auto-delete after time (expiration)

Problems:
â€Ē Zombie Data = Deleted data comes back ðŸ‘ŧ
â€Ē SSTable = Immutable file (can't modify)

Ready to learn tombstones! ðŸŠĶ

ðŸĪ” What are Tombstones?

Let's understand tombstones with a perfect real-world analogy!

🏛ïļ The Cemetery Analogy (Perfect Explanation!)

Imagine a city with 3 record offices in different locations:

WITHOUT tombstones (naive approach - BROKEN!):
â€Ē Person "John Smith" exists in all 3 offices' records
â€Ē John passes away
â€Ē Office A removes his record immediately
â€Ē But Office B is closed for renovation!
â€Ē And Office C's computer is down!

Problem:
â€Ē 2 weeks later, Office B reopens
â€Ē They still have John's record (alive!)
â€Ē They send updates to Office A: "Hey, we have John!"
â€Ē Office A: "Oh, we must have lost it, thanks!"
â€Ē John is "resurrected" in records! ðŸ‘ŧ
â€Ē This is ZOMBIE DATA!

WITH tombstones (smart way - WORKS!):
â€Ē Person "John Smith" exists in all 3 offices
â€Ē John passes away
â€Ē Office A doesn't erase the record!
â€Ē Instead, they write: John Smith - Deceased Dec 2025
â€Ē This is a TOMBSTONE!

What happens:
â€Ē Office B reopens
â€Ē Syncs with Office A
â€Ē Sees: "John Smith - Deceased"
â€Ē Marks their record as deceased too!
â€Ē Office C comes back online
â€Ē Also gets the tombstone marker
â€Ē All 3 offices now know John is deceased ✓

After 10 days (gc_grace_seconds):
â€Ē All offices have seen the tombstone
â€Ē Safe to remove the tombstone marker itself
â€Ē Record finally completely removed
â€Ē But NO resurrection possible!

Key insight:
Tombstone = Death certificate
â€Ē Proves someone is deceased
â€Ē Prevents "coming back to life"
â€Ē Spreads to all record offices
â€Ē Eventually removed when no longer needed

This is EXACTLY how Cassandra works!

ðŸŠĶ

Definition

Tombstone = Special marker for deleted data

What it is:
â€Ē NOT actual data deletion
â€Ē A marker that says "deleted"
â€Ē Has timestamp (when deleted)
â€Ē Stored like normal data

Looks like:
Key: "user123"
Value: <TOMBSTONE>
Timestamp: 1735488000

Think: Gravestone marking deleted data!

❌

What It Does

Tombstone overrides old data

When query reads:
â€Ē Finds data: "John" age 30
â€Ē Also finds: TOMBSTONE
â€Ē Tombstone newer? Return nothing ✓
â€Ē Data newer? Return data

Timestamp wins:
â€Ē Data written at time 100
â€Ē Tombstone at time 200
â€Ē Tombstone wins! (newer)

Prevents seeing deleted data!

⏰

Lifecycle

Tombstone has a lifespan:

Phase 1: Creation (instant)
â€Ē DELETE executed
â€Ē Tombstone written

Phase 2: Active (10 days)
â€Ē Spreads to all replicas
â€Ē Blocks old data
â€Ē Participates in queries

Phase 3: Removal (compaction)
â€Ē After gc_grace_seconds
â€Ē Compaction runs
â€Ē Tombstone deleted!

Total: ~10 days existence

❓ Why Are Tombstones Needed?

Understanding the distributed systems challenge that tombstones solve!

🌐

The Distributed Deletion Problem

Scenario: You have data replicated on 3 servers (RF=3)

Initial State:
â€Ē Server A: user123 = "Alice"
â€Ē Server B: user123 = "Alice"
â€Ē Server C: user123 = "Alice"
All 3 servers have the same data ✓

❌ Attempt 1: Direct Deletion (BROKEN!)
â€Ē Client sends: DELETE FROM users WHERE id='user123'
â€Ē DELETE reaches Server A (deletes immediately)
â€Ē Server B is down for maintenance 🔧
â€Ē Server C has network issue 🌐

Current state:
â€Ē Server A: (nothing - deleted) ✓
â€Ē Server B: user123 = "Alice" (still has it!)
â€Ē Server C: user123 = "Alice" (still has it!)

2 days later - Server B comes back:
â€Ē Repair process runs (anti-entropy)
â€Ē B compares with A: "I have user123, you don't!"
â€Ē B thinks A lost the data
â€Ē B sends data to A: "Here, I'll restore it for you!"
â€Ē ZOMBIE! user123 = "Alice" is back on A! ðŸ‘ŧ

This is catastrophic!
â€Ē Deleted data reappears
â€Ē GDPR compliance violated (can't guarantee deletion)
â€Ē Financial records reappear (audit nightmare)
â€Ē Deleted user accounts come back

✅ Solution: Tombstones (WORKS!)
â€Ē Client sends: DELETE FROM users WHERE id='user123'
â€Ē Server A writes: user123 = TOMBSTONE(timestamp: 100)
â€Ē Server B down (missed it)
â€Ē Server C network issue (missed it)

Current state:
â€Ē Server A: user123 = TOMBSTONE(100) ✓
â€Ē Server B: user123 = "Alice" (timestamp: 50)
â€Ē Server C: user123 = "Alice" (timestamp: 50)

When Server B comes back:
â€Ē Repair process runs
â€Ē B sees: A has TOMBSTONE(100), I have Alice(50)
â€Ē Timestamp comparison: 100 > 50
â€Ē B adopts tombstone: "Oh, this was deleted!"
â€Ē B writes: user123 = TOMBSTONE(100)
â€Ē Server C also gets tombstone
â€Ē All 3 servers now know data is deleted! ✓

After 10 days (gc_grace_seconds):
â€Ē All servers have seen the tombstone
â€Ē Compaction removes tombstone
â€Ē Data truly gone
â€Ē No resurrection possible!

Key Insight: Tombstones convert "absence of data" into "presence of deletion marker" - something that can be replicated and wins timestamp battles!

⚔ïļ

Timestamp Conflicts

Tombstones use timestamps to win!

Example conflict:
â€Ē Server A: DELETE at time 100
â€Ē Server B: WRITE at time 90

Who wins?
â€Ē Tombstone(100) vs Data(90)
â€Ē 100 > 90
â€Ē Tombstone wins! ✓
â€Ē Data is deleted

Reverse scenario:
â€Ē Server A: DELETE at time 100
â€Ē Server B: WRITE at time 110
â€Ē Data(110) wins!
â€Ē Deletion overridden (intended!)

Last Write Wins (LWW) principle!

🔒

Immutable SSTables

SSTables can't be modified!

Problem:
â€Ē Data written to SSTable
â€Ē File closed (immutable)
â€Ē Can't go back and erase!

Solution:
â€Ē Write tombstone in NEW SSTable
â€Ē During compaction:
- Old SSTable: user123 = "Alice"
- New SSTable: user123 = TOMBSTONE
- Merged result: (deleted)

Tombstones work with immutability!

🔄

Eventual Consistency

Deletions take time to propagate

Timeline:
â€Ē T=0: DELETE executed
â€Ē T=1ms: Tombstone on Server A
â€Ē T=100ms: Replicated to Server B
â€Ē T=5min: Server C comes online
â€Ē T=10min: All servers have it ✓

During propagation:
â€Ē Queries might still see data
â€Ē Tombstone eventually wins
â€Ē Consistent eventually!

This is "eventual consistency"

🗂ïļ Types of Tombstones

Cassandra has 4 different types of tombstones!

📝

1. Cell Tombstone

Deletes a single column value

Query:
DELETE email FROM users WHERE id='user123'

What happens:
â€Ē Only 'email' column deleted
â€Ē Other columns remain
â€Ē Tombstone for one cell only

Before:
user123: {name: "Alice", email: "alice@email.com"}

After:
user123: {name: "Alice", email: DELETED}

Most granular tombstone type!

🗑ïļ

2. Row Tombstone

Deletes entire row

Query:
DELETE FROM users WHERE id='user123'

What happens:
â€Ē Entire row deleted
â€Ē All columns gone
â€Ē Single tombstone for whole row

Before:
user123: {name: "Alice", email: "alice@email.com", age: 30}

After:
user123: ROW DELETED

Most common tombstone type!

ðŸ“Ķ

3. Partition Tombstone

Deletes entire partition

Query:
DELETE FROM events WHERE user_id='user123'
(Partition key = user_id)

What happens:
â€Ē All rows in partition deleted
â€Ē Could be thousands of rows!
â€Ē Single partition tombstone

Example:
â€Ē user123 has 10,000 events
â€Ē One tombstone deletes all!
â€Ē Very efficient ✓

Efficient for large deletions!

⏰

4. Range Tombstone

Deletes range of clustering keys

Query:
DELETE FROM events
WHERE user_id='user123'
AND timestamp >= '2024-01-01'
AND timestamp < '2024-02-01'

What happens:
â€Ē Deletes date range
â€Ē One tombstone for whole range
â€Ē Marks start→end as deleted

Example:
Range: Jan 1 - Feb 1
Could be 1000s of rows
Single range tombstone!

Perfect for time-series data!

âģ

5. TTL Tombstone

Auto-created when TTL expires

Query:
INSERT INTO sessions (id, data)
VALUES ('session123', 'data')
USING TTL 3600; -- 1 hour

What happens:
â€Ē After 1 hour passes
â€Ē Automatically becomes tombstone
â€Ē No DELETE needed!

Perfect for:
â€Ē Session data
â€Ē Cache entries
â€Ē Temporary records

Set-it-and-forget-it deletion!

📊

Tombstone Sizes

Different types, different overhead:

Cell Tombstone: ~10 bytes
â€Ē Column name + timestamp

Row Tombstone: ~20 bytes
â€Ē Partition key + timestamp

Partition Tombstone: ~15 bytes
â€Ē Partition key only
â€Ē Deletes 1000s of rows!

Range Tombstone: ~30 bytes
â€Ē Start + end range + timestamp

Insight: Tombstones are tiny but powerful!

ðŸŽŪ Interactive Tombstone Simulator - Try It!

See tombstone creation and zombie prevention in action!

ðŸŠĶ Live Tombstone Simulator
ðŸŠĶ Tombstone Simulator Ready!

Simulation scenario:
â€Ē 3 Cassandra servers (A, B, C) with RF=3
â€Ē GC grace seconds: 10 days (simplified to 5 operations)

Try this sequence:
1. Insert Data → Creates user on all servers
2. Server B Goes Down → Simulates offline node
3. Delete Data → Creates tombstone (B misses it!)
4. Server B Comes Back → Watch zombie prevention!
5. Run Compaction → After grace period expires

ðŸŽŊ Goal: See how tombstones prevent zombie data resurrection!
✓
Server A
✓
Server B
✓
Server C
0
Tombstones
0
Operations (Age)

ðŸ’Ą What You're Learning

This simulator demonstrates the critical zombie data problem:

WITHOUT tombstones:
1. Delete data on Server A
2. Server B is offline (misses deletion)
3. B comes back, sees it has data A doesn't
4. B "helpfully" restores data to A
5. Zombie! Deleted data is back! ðŸ‘ŧ

WITH tombstones (what you'll see):
1. Delete creates tombstone on A & C
2. Server B offline (misses tombstone)
3. B comes back with old data
4. Repair: B sees tombstone from A/C
5. Tombstone wins (newer timestamp)
6. B adopts tombstone
7. No resurrection! Deletion preserved! ✓

After grace period:
â€Ē All servers have seen tombstone
â€Ē Compaction removes it safely
â€Ē Data truly gone forever

Key insight: Tombstones turn "absence" into "presence of deletion marker" that can be replicated and wins conflicts!

⚙ïļ How Tombstones Work (Complete Flow)

Let's trace the complete tombstone lifecycle!

Tombstone Complete Lifecycle Step 1: Data Exists Server A: user123 = "Alice" Server B: user123 = "Alice" Step 2: DELETE DELETE FROM users WHERE id='user123' Step 3: Tombstone Server A: ðŸŠĶ TOMBSTONE(ts: 100) Server B: ðŸŠĶ TOMBSTONE(ts: 100) Deletion marker created! Step 4: Query Blocked SELECT * FROM users Result: (empty) Tombstone hides the data Step 5: Replication & Zombie Prevention Server C was offline, comes back with old data: "Alice"(ts: 50) Repair process: Sees tombstone(100) vs data(50) ✓ Tombstone wins! Server C adopts tombstone. No zombie! Step 6: Grace Period (10 days) Tombstone kept for gc_grace_seconds Ensures all servers (even offline) see deletion Default: 864,000 seconds = 10 days Step 7: Compaction (After 10 days) Grace period expired All servers have seen tombstone ✓ Tombstone removed! Data truly gone! T=0 Data exists T=1ms Tombstone T=10 days Grace expires T=10d+ Deleted ðŸŠĶ Key: Tombstones live ~10 days to prevent zombie data, then removed by compaction

⏰ GC Grace Seconds (The Magic Number)

âģ

What It Is

gc_grace_seconds = Tombstone lifespan

Default: 864,000 seconds
â€Ē 864000 ÷ 60 = 14,400 minutes
â€Ē 14400 ÷ 60 = 240 hours
â€Ē 240 ÷ 24 = 10 days

Purpose:
â€Ē Time for offline servers
â€Ē Come back and see deletion
â€Ē Prevents zombie data

After 10 days: safe to remove!

ðŸĪ”

Why 10 Days?

Assumptions about server downtime:

Typical scenarios:
â€Ē Planned maintenance: 1-4 hours
â€Ē Network issues: Minutes to hours
â€Ē Hardware failure: 1-3 days
â€Ē Catastrophic failure: 3-7 days

10 days provides:
â€Ē Plenty of buffer ✓
â€Ē Time for weekend repairs
â€Ē Margin for errors

Conservative but safe!

⚙ïļ

Configuration

Set per table:

View current:
DESC TABLE users;

Change value:
ALTER TABLE users
WITH gc_grace_seconds = 86400;
(1 day = 86400 seconds)

Common values:
â€Ē 10 days: Default (safe)
â€Ē 1 day: Fast-moving data
â€Ē 0: Special cases only!

⚠ïļ

Too Short = Danger!

Setting gc_grace_seconds = 0:

What happens:
â€Ē Tombstone removed immediately
â€Ē Server offline for 1 hour
â€Ē Tombstone already gone!
â€Ē Server comes back with old data
â€Ē Zombie data! ðŸ‘ŧ

Only use 0 when:
â€Ē Single node cluster (testing)
â€Ē Never run repair
â€Ē Understand the risks!

Generally: DON'T DO THIS!

📈

Too Long = Problems!

Setting gc_grace_seconds = 30 days:

What happens:
â€Ē Tombstones live 30 days
â€Ē Accumulate on disk
â€Ē Slow down queries
â€Ē Waste disk space

Problems:
â€Ē Many tombstones in reads
â€Ē Performance degradation
â€Ē Larger SSTables
â€Ē Slower compaction

Balance safety vs performance!

✅

Best Practice

Follow these guidelines:

Default (10 days):
â€Ē Most tables ✓
â€Ē Safe and proven

Shorter (1-3 days):
â€Ē Stable cluster
â€Ē Frequent deletions
â€Ē Monitor closely

Never use 0 unless:
â€Ē Testing only
â€Ē Single node
â€Ē You're an expert

When in doubt: keep default!

⚠ïļ Tombstone Problems (Performance Impact)

ðŸ’Ĩ

The Tombstone Accumulation Problem

Worst-case scenario: Millions of tombstones!

Example: Time-series data with TTL
â€Ē Table: sensor_data (temperature readings)
â€Ē Inserts: 1 million rows/day with TTL=7 days
â€Ē After 7 days: 1M rows become tombstones
â€Ē After 10 days (gc_grace): Tombstones eligible for removal
â€Ē But compaction hasn't run yet!

State after 10 days:
â€Ē Active data: 7M rows (7 days × 1M/day)
â€Ē Tombstones: 3M+ (days 8-10)
â€Ē Ratio: 30% tombstones!

When you query:
â€Ē Must read ALL tombstones
â€Ē Filter them out
â€Ē Return only live data
â€Ē Massive overhead!

Performance impact:
â€Ē Query scans 10M rows (7M live + 3M tombstones)
â€Ē 43% waste (3M ÷ 10M)
â€Ē P99 latency: 50ms → 200ms (4x slower!)
â€Ē Disk I/O: Unnecessary reads
â€Ē Memory: Tombstones in bloom filters
â€Ē Network: Tombstones sent between nodes

Warning threshold:
Cassandra warns at 100,000 tombstones in a single query!

WARN  [ReadStage-2] Read 250000 live rows and 500000 tombstone cells
for query SELECT * FROM sensor_data LIMIT 1000
This means: You have a problem!

📉

1. Read Performance

Tombstones slow queries!

Why:
â€Ē Must read tombstones from disk
â€Ē Process and filter them
â€Ē Then return live data

Example query:
SELECT * FROM users LIMIT 100

Without tombstones:
â€Ē Read 100 rows (5ms)

With 10k tombstones:
â€Ē Read 10,000 tombstones (100ms)
â€Ē Read 100 live rows (5ms)
â€Ē Total: 105ms (21x slower!)

More tombstones = slower queries!

ðŸ’ū

2. Disk Space

Tombstones consume space!

Example table:
â€Ē 1B rows deleted
â€Ē Each tombstone: 20 bytes
â€Ē Total: 20GB of tombstones!

Problem:
â€Ē Disk fills up
â€Ē Can't add new data
â€Ē Expensive storage costs

Solution:
â€Ē Run compaction
â€Ē Remove expired tombstones
â€Ē Free up space ✓

🔧

3. Compaction Overhead

Many tombstones = slow compaction!

Why:
â€Ē Must process all tombstones
â€Ē Check if expired
â€Ē Merge with data

Example:
â€Ē 100GB data
â€Ē 50GB tombstones
â€Ē Compaction processes 150GB
â€Ē Takes 3x longer!

Impact:
â€Ē Higher CPU usage
â€Ē More disk I/O
â€Ē Slower overall system

ðŸĒ Real Company Examples (Production Stories)

📊

A. Discord: TTL Tombstone Disaster (500M Tombstones!)

The Problem:
Discord stores message read status with TTL=30 days. Every time a user reads a message, write with TTL. 150M users × 1000 messages/day = 150 billion operations. After 30 days, massive tombstone accumulation!

What went wrong:
â€Ē Table: user_read_status (user_id, channel_id, message_id, TTL 30 days)
â€Ē Writes: 150 billion/month
â€Ē After TTL expires: All become tombstones
â€Ē gc_grace_seconds: 10 days
â€Ē Total tombstone period: 30 days (TTL) + 10 days (grace) = 40 days
â€Ē Tombstone accumulation: 150B × 40/30 = 200 billion tombstones!

Impact:
â€Ē Query: "Get unread messages for user"
â€Ē Scanned 50M tombstones per query
â€Ē P99 latency: 15ms → 3000ms (200x slower!)
â€Ē Warnings: "Read 100000 tombstones" constantly
â€Ē Disk usage: 2TB of tombstones (50% of cluster!)
â€Ē Compaction: Couldn't keep up

Solution:
1. Reduced TTL: 30 days → 7 days (less accumulation)
2. Reduced gc_grace: 10 days → 1 day (faster removal)
3. Aggressive compaction: Forced major compaction
4. Rewrote approach: Moved to time-bucketed partitions
- New partition each day
- Drop entire partition after 7 days (no tombstones!)

Results after fix:
â€Ē Tombstone count: 200B → 2M (99.999% reduction!)
â€Ē P99 latency: 3000ms → 18ms (167x faster!)
â€Ē Disk usage: 2TB freed
â€Ē Warnings: Gone ✓

Lesson learned:
"TTL creates tombstones! High-volume TTL writes = tombstone explosion. Time-bucketed partitions + partition drops avoid tombstones entirely. This saved our cluster from collapsing."

ðŸ“ą

B. Instagram: Deletion Pattern Optimization

Use case:
User deletes photos/posts. 2B users, average 500 photos each. Deletion rate: 0.1% daily = 1M photo deletions/day. Need efficient deletion without tombstone accumulation.

Initial approach (WRONG):
â€Ē DELETE FROM user_photos WHERE user_id=X AND photo_id=Y
â€Ē Created cell tombstone for each photo
â€Ē 1M deletions/day = 1M tombstones/day
â€Ē After 10 days (gc_grace): 10M tombstones

Problem:
â€Ē Query: "Get user's photos"
â€Ē Some users deleted 1000s of photos
â€Ē Scan 1000s of tombstones per user
â€Ē P99: 50ms → 400ms (8x slower!)

Optimized approach:
1. Soft delete instead of hard delete:
- UPDATE photos SET deleted=true, deleted_at=now()
- No tombstone created! ✓
- Query: WHERE deleted=false
2. Async cleanup job:
- Nightly job processes deleted=true
- Batches deletions into range tombstones
- One range tombstone for 1000 photos!
3. Short gc_grace for cleanup table:
- gc_grace_seconds = 86400 (1 day)
- Rapid tombstone removal

Results:
â€Ē Tombstones per query: 1000s → 0 (live queries see no tombstones!)
â€Ē P99 latency: 400ms → 35ms (11x faster!)
â€Ē Cleanup efficiency: 1000 photos = 1 range tombstone
â€Ē Compaction load: 90% reduction

Instagram engineer quote:
"Never delete in the hot path! Soft delete + async cleanup = best of both worlds. Users see instant deletion (deleted=true), but actual tombstones created in background batches. This keeps read queries fast while still removing data."

🚗

C. Uber: Time-Series gc_grace Tuning

Scenario:
Trip location history - massive time-series writes. Table stores GPS coordinates every 4 seconds during trips. TTL=30 days for GDPR compliance. 20M trips/day × 450 points/trip × 30 days = 270 billion rows!

Challenge:
â€Ē After 30 days: 9B rows/day expire → tombstones
â€Ē gc_grace_seconds: 864000 (10 days default)
â€Ē Tombstone accumulation: 9B/day × 10 days = 90 billion tombstones!
â€Ē Queries scanning expired data regions: massive tombstone overhead

Initial problem:
â€Ē P99 read latency: 25ms → 800ms
â€Ē Disk usage: 80% tombstones
â€Ē Warning logs: 50k/day about tombstones
â€Ē Compaction: Running 24/7, still couldn't catch up

Uber's solution - Aggressive tuning:
1. Reduced gc_grace_seconds: 10 days → 1 day
- Rationale: "Our cluster is stable, nodes rarely down >24h"
- Monitoring: Alert if node down >12 hours
- Risk accepted: Worth it for performance
2. Increased repair frequency:
- Run repair every 12 hours (instead of weekly)
- Ensures deletions propagate quickly
3. Partition strategy:
- Partition by (trip_id, day_bucket)
- Old partitions naturally age out
4. Compaction tuning:
- More frequent minor compactions
- Aggressive tombstone_threshold: 0.1 (vs 0.2 default)

Results:
â€Ē Tombstone lifespan: 10 days → 1 day (90% reduction)
â€Ē Tombstone count: 90B → 9B (10x less!)
â€Ē P99 latency: 800ms → 45ms (18x faster!)
â€Ē Disk usage: 80% → 15% tombstones
â€Ē Compaction: Caught up within 48 hours

Trade-off:
â€Ē Risk: If node down >24 hours, possible zombie data
â€Ē Mitigation: Monitoring + alerts + 12h repair cycles
â€Ē Reality: In 2 years of production, zero zombie incidents

Key insight:
"Default gc_grace_seconds (10 days) is conservative. For stable clusters with good monitoring, 1-3 days is often better. The performance gain from faster tombstone removal outweighs the small zombie risk. But: requires operational excellence!"

✅ Best Practices (Production-Ready Tips!)

✅

1. Monitor Tombstone Warnings

Set up monitoring!

Watch for:
â€Ē "Read X tombstones" warnings
â€Ē Threshold: 100,000 tombstones

Alert when:
â€Ē >10 warnings/hour
â€Ē Consistent growth

Check with nodetool:
nodetool tablestats keyspace.table
(Look for "Tombstoned cells")

Early detection prevents disasters!

ðŸŽŊ

2. Use Time-Bucketed Partitions

For time-series data:

Instead of TTL:
❌ USING TTL 86400
(Creates tombstones)

Use partitions:
✓ Partition key: (sensor_id, day)
â€Ē One partition per day
â€Ē Drop entire partition when old
â€Ē No tombstones created!

Example:
DROP TABLE sensor_data_20241201;
(Deletes whole partition instantly)

⏰

3. Tune gc_grace Carefully

Guidelines by scenario:

Stable cluster (3 data centers):
â€Ē gc_grace: 3-5 days ✓
â€Ē Aggressive repair schedule

Volatile environment:
â€Ē gc_grace: 10 days (default) ✓
â€Ē Conservative, safe

High-delete workload:
â€Ē gc_grace: 1-2 days
â€Ē Fast tombstone removal
â€Ē Requires monitoring!

Never use 0 in production!

🔄

4. Soft Delete When Possible

Instead of DELETE:
UPDATE SET deleted=true

Benefits:
â€Ē No tombstone in hot path!
â€Ē Instant for user
â€Ē Query: WHERE deleted=false

Async cleanup:
â€Ē Background job
â€Ē Batch actual DELETEs
â€Ē Range tombstones

Perfect for:
User-facing deletions!

🔧

5. Regular Compaction

Ensure compaction runs!

Check:
nodetool compactionstats

If tombstones accumulate:
â€Ē Force compaction:
nodetool compact keyspace table

Tune strategy:
â€Ē tombstone_threshold: 0.2→0.1
â€Ē More aggressive removal

Monitor:
â€Ē Pending compactions
â€Ē Should be near 0

📊

6. Avoid Frequent Range Deletes

Range deletes create range tombstones:

Example (BAD):
DELETE FROM events
WHERE user_id=X
AND time < '2024-01-01'

Creates:
â€Ē Large range tombstone
â€Ē Affects all queries in range

Better:
â€Ē Use time-bucketed partitions
â€Ē Drop old partitions
â€Ē Or design for no deletes!

ðŸšĻ Common Mistakes to Avoid

❌ Setting gc_grace_seconds = 0
→ Zombie data guaranteed! Only for testing.
→ Production requires time for offline nodes

❌ Using TTL without understanding tombstones
→ High-volume TTL writes = tombstone explosion
→ Consider time-bucketed partitions instead
→ Monitor tombstone warnings closely

❌ Ignoring tombstone warnings
→ "Read 100000 tombstones" = serious problem
→ Performance degrading, fix immediately!
→ Don't wait until cluster fails

❌ Frequent small range deletes
→ Creates many range tombstones
→ Slows all queries in that range
→ Design around deletions when possible

❌ Not monitoring compaction
→ Compaction removes tombstones
→ If not running: tombstones accumulate forever
→ Check nodetool compactionstats regularly

❌ Deleting in the hot path
→ User clicks "delete", waits for tombstone
→ Use soft delete + async cleanup instead
→ Better user experience + fewer live tombstones

❌ Same gc_grace for all tables
→ Different tables have different needs
→ High-delete tables: shorter gc_grace
→ Rarely-deleted tables: longer is fine

✓ Remember: Tombstones are necessary but need management!

💞 Interview Questions & Answers

1
What are tombstones in Cassandra and why are they needed?
▾

Complete Answer:

Definition:

Tombstones are special deletion markers that Cassandra writes instead of immediately removing data. They are metadata entries with a timestamp indicating when data was deleted, stored alongside regular data in SSTables.

Why tombstones exist (the fundamental problem):

In a distributed system with data replicated across multiple nodes, you cannot simply erase data immediately because:

  • Nodes can be offline: If you delete on Server A but Server B is down for maintenance, B never sees the deletion. When B comes back, it still has the old data and will "restore" it to A through anti-entropy repair, causing zombie data resurrection.
  • Immutable SSTables: Cassandra's SSTables are immutable - once written, they cannot be modified. You can't go back and erase data from an existing file.
  • Eventual consistency: Operations don't happen simultaneously across all replicas. Deletions take time to propagate.

How tombstones solve this:

Instead of erasing data, Cassandra writes a tombstone marker that says "this data was deleted at timestamp X". This marker:

  • Propagates to all replicas just like regular data
  • Wins timestamp conflicts (if tombstone timestamp > data timestamp)
  • Prevents zombie data by explicitly marking deletion
  • Eventually gets removed after gc_grace_seconds (default 10 days)

Example scenario:

  1. Data: user123="Alice" exists on Servers A, B, C (timestamp: 100)
  2. Server B goes offline for maintenance
  3. Client: DELETE user123 → Creates tombstone on A & C (timestamp: 200)
  4. Server B comes back 2 days later, still has user123="Alice"(100)
  5. Repair process: B sees tombstone(200) from A/C vs data(100)
  6. Timestamp comparison: 200 > 100 → Tombstone wins
  7. B adopts tombstone, deletion preserved, no zombie!

When tombstones are removed:

After gc_grace_seconds (default 864,000 seconds = 10 days), compaction removes expired tombstones. This grace period ensures even offline nodes have time to see the deletion before the marker disappears.

Key takeaway for interview:

"Tombstones solve the distributed deletion problem by converting 'absence of data' into 'presence of a deletion marker' that can be replicated and wins timestamp conflicts. They prevent zombie data resurrection in distributed systems where nodes can be temporarily offline."

2
What is gc_grace_seconds and what happens if you set it too low or too high?
▾

Complete Answer:

Definition:

gc_grace_seconds is a table-level setting that specifies how long Cassandra keeps tombstones before they're eligible for removal during compaction. The default is 864,000 seconds (10 days).

Purpose:

This grace period gives offline or partitioned nodes time to come back online and see the deletion tombstone before it's removed. It's insurance against zombie data resurrection.

How it works:

  1. Tombstone created with deletion timestamp T
  2. Current time C is tracked during compaction
  3. If (C - T) > gc_grace_seconds, tombstone is eligible for removal
  4. Compaction removes eligible tombstones

Setting it TOO LOW (dangerous!):

Example: gc_grace_seconds = 0 (worst case)

  • What happens:
    • Tombstones removed immediately during compaction
    • Node goes offline for 1 hour
    • Tombstone already deleted from other nodes
    • Node comes back with old data, no tombstone to stop it
    • Result: Zombie data! Deleted data resurrects
  • Scenarios this can happen:
    • Network partition (nodes can't communicate)
    • Hardware failure requiring replacement
    • Extended maintenance window
    • Data center failure
  • Real impact: GDPR compliance violation (can't guarantee deletion), financial records reappear (audit nightmare), deleted user accounts come back

When you might use shorter gc_grace:

  • Stable, well-monitored cluster
  • Frequent repair runs (every 12-24 hours)
  • High tombstone accumulation issues
  • Example: 1-3 days for production with excellent ops
  • But never 0 in production!

Setting it TOO HIGH (performance issues!):

Example: gc_grace_seconds = 30 days or more

  • What happens:
    • Tombstones accumulate for 30 days
    • High-delete workload: millions of tombstones
    • Queries must read and filter all tombstones
    • Result: Severe performance degradation
  • Specific problems:
    • Read latency: Scanning 100,000+ tombstones per query
    • Disk space: Tombstones consuming GBs of storage
    • Compaction overhead: Processing huge volumes
    • Memory pressure: Bloom filters include tombstones
  • Warning signs: Logs showing "Read 100000 tombstones", P99 latency spikes, disk usage growing despite deletions

The 10-day default rationale:

  • Covers typical failure scenarios (1-7 days)
  • Includes weekends for manual intervention
  • Provides safety margin for unexpected issues
  • Balances safety vs. performance

Tuning guidelines by scenario:

  • Default workload: Keep 10 days (864,000 sec)
  • High-delete, stable cluster: 1-3 days (86,400-259,200 sec)
  • Rarely deleted data: 10-14 days is fine
  • Time-series with TTL: Consider shorter (1-2 days) with aggressive repair

Monitoring requirements:

If using shorter gc_grace_seconds, you MUST:

  • Monitor node downtime (alert if >50% of gc_grace)
  • Run repair frequently (at minimum every gc_grace_seconds)
  • Have operational excellence and rapid response
  • Accept small zombie risk for performance gain

Key takeaway for interview:

"gc_grace_seconds is the tombstone lifespan - too short risks zombie data (if nodes offline longer), too long accumulates tombstones and degrades performance. The 10-day default balances safety and performance. Tuning lower requires operational excellence and frequent repairs. The core trade-off: safety of data deletion vs. performance overhead of keeping tombstones."

3
How do tombstones impact performance and how would you mitigate tombstone-related issues?
▾

Complete Answer:

Performance impacts:

1. Read performance degradation:

  • Problem: Queries must read tombstones from disk, process them, filter them out, then return live data
  • Example: SELECT * FROM users LIMIT 100
    • Without tombstones: Read 100 rows (5ms)
    • With 10,000 tombstones: Read 10,000 tombstones (100ms) + 100 rows (5ms) = 105ms (21x slower!)
  • Warning threshold: Cassandra warns at 100,000 tombstones per query
  • Severe case: Millions of tombstones can make queries timeout entirely

2. Disk space consumption:

  • Each tombstone: ~10-30 bytes
  • 1 billion tombstones = 10-30 GB disk space
  • Can represent 50-80% of total data size in high-delete workloads
  • Disk fills up, preventing new writes

3. Compaction overhead:

  • Compaction must process all tombstones
  • Check expiration, merge with data
  • High tombstone volume → slower compaction → tombstones accumulate further (vicious cycle)
  • CPU and I/O overhead even when tombstones just sitting there

4. Memory pressure:

  • Bloom filters include tombstones
  • Larger bloom filters = more memory
  • Key cache overhead for tombstone partition keys

Mitigation strategies:

Strategy 1: Avoid tombstones entirely (best!):

  • Time-bucketed partitions:
    • Instead of: TTL on rows
    • Use: Partition per time bucket (day/week/month)
    • DROP entire partition when old
    • No tombstones created!
    • Example: sensor_data_20241201, sensor_data_20241202, etc.
  • Design without deletions:
    • Append-only architecture where possible
    • Soft deletes (deleted=true flag) instead of hard deletes
    • Filter at application layer

Strategy 2: Soft delete + async cleanup:

  • User action: UPDATE SET deleted=true (no tombstone!)
  • Queries: WHERE deleted=false (never see deleted rows)
  • Background job: Batch actual DELETEs into range tombstones
  • Benefit: Hot path (user queries) never sees tombstones
  • Used by Instagram for photo deletions

Strategy 3: Optimize gc_grace_seconds:

  • Reduce from 10 days to 1-3 days for stable clusters
  • Tombstones removed 3-10x faster
  • Requires: Frequent repair (every 12-24h), good monitoring, operational discipline
  • Trade-off: Small zombie risk for big performance gain

Strategy 4: Aggressive compaction tuning:

  • tombstone_threshold: Lower from 0.2 to 0.1
    • Triggers compaction when 10% tombstones (vs 20% default)
    • More frequent compaction = faster removal
  • unchecked_tombstone_compaction: Enable to compact tombstone-heavy SSTables even without triggering normal compaction
  • Manual intervention: Force compact when needed: nodetool compact keyspace table

Strategy 5: Monitoring and alerting:

  • Alert on tombstone warnings in logs
  • Track tombstone ratio: nodetool tablestats
  • Monitor compaction pending tasks
  • Query latency correlation with tombstone counts
  • Early detection prevents severe issues

Strategy 6: Query optimization:

  • Avoid full partition scans when many tombstones
  • Add LIMIT to bound tombstone scanning
  • Use partition-restricted queries
  • Filter at application when possible

Real-world example (Discord):

  • Problem: 150B TTL writes/month → 200B tombstones, queries reading 50M tombstones, P99: 15ms → 3000ms (200x slower!)
  • Solution:
    • Reduced TTL: 30 → 7 days
    • Reduced gc_grace: 10 → 1 day
    • Rewrote to time-bucketed partitions (drop partitions instead of TTL)
  • Results: Tombstones: 200B → 2M (99.999% reduction!), P99: 3000ms → 18ms (167x faster!)

When to use each strategy:

  • Time-series data: Time-bucketed partitions (avoid tombstones entirely)
  • User-facing deletes: Soft delete + async cleanup
  • High-volume deletes: Reduce gc_grace + aggressive compaction
  • Any workload: Monitoring is mandatory

Key takeaway for interview:

"Tombstones impact performance by requiring reads, processing, and storage while providing no value. Best mitigation: avoid them entirely through time-bucketed partitions or soft deletes. When unavoidable: reduce gc_grace_seconds (with operational excellence), tune compaction aggressively, and monitor closely. The pattern: prevention > optimization > monitoring."

Advertisement

Responsive Ad