GC GRACE SECONDS - The Grace Period
โฐ Master Cassandra's grace period with interactive timeline, animations, and real-world scenarios. Learn from absolute basics!
๐ Fundamentals (Start Here!)
Understanding ALL basic concepts before diving into GC Grace Seconds!
Tombstone Recap
Tombstone = Deletion marker
When you DELETE data:
โข Cassandra writes tombstone
โข Not actual deletion!
โข Marker says "deleted"
โข Has timestamp
Purpose:
Prevents zombie data resurrection
in distributed systems
Tombstones spread to all servers!
Compaction
Compaction = Cleanup process
What it does:
โข Merges multiple SSTables
โข Removes old data versions
โข Deletes expired tombstones
โข Frees disk space
When it runs:
โข Automatically in background
โข Or manually triggered
Compaction is the janitor!
Grace Period
Grace Period = Extra time given
Like a payment grace period
Example:
โข Bill due: Jan 1
โข Grace period: 10 days
โข Can pay until: Jan 11
โข No penalty during grace!
For tombstones:
Extra time before removal
Ensures all servers see it!
Zombie Data
Zombie = Deleted data comes back!
How it happens:
1. Delete data on Server A
2. Server B offline (missed it)
3. Tombstone removed too early
4. Server B comes back
5. Still has old data
6. Copies back to A!
7. Zombie! Data resurrected!
Grace period prevents this!
Seconds
Second = Unit of time
Conversions:
โข 60 seconds = 1 minute
โข 3,600 seconds = 1 hour
โข 86,400 seconds = 1 day
โข 864,000 seconds = 10 days
Why seconds:
Computer precision
Standard time unit
gc_grace_seconds uses seconds!
Default Setting
Default = Pre-configured value
What it means:
โข Value set automatically
โข You don't need to configure
โข Works out-of-the-box
โข Can be changed if needed
gc_grace_seconds default:
864,000 seconds (10 days)
Most users keep the default!
๐ Quick Reference
You now know:
โข Tombstone: Deletion marker to prevent zombie data
โข Compaction: Cleanup process that removes tombstones
โข Grace Period: Extra time before removal
โข Zombie Data: Deleted data coming back (bad!)
โข Seconds: Time unit (86,400 = 1 day)
โข Default: 864,000 seconds = 10 days
Ready to learn GC Grace Seconds! โฐ
๐ค What is gc_grace_seconds?
โฐ The Library Book Return Analogy
Imagine a library with a book checkout system:
The Scenario:
โข You borrow a book on January 1st
โข Due date: January 10th
โข But library has 10-day grace period
โข No late fees until: January 20th
Why grace period exists:
โข Maybe you're sick
โข Maybe traveling
โข Maybe forgot
โข Library gives you extra time โ
This is EXACTLY gc_grace_seconds!
In Cassandra:
โข Tombstone created: January 1st
โข Should be removed: Immediately (efficiency)
โข But Cassandra waits: 10 days grace
โข Actually removed: January 11th (or later)
Why wait 10 days?
โข Server A creates tombstone
โข Server B is offline (maintenance)
โข Server B needs to see tombstone!
โข 10 days = enough time for B to come back
โข B sees tombstone, deletes its copy
โข Then tombstone can be removed safely โ
Without grace period:
โข Tombstone removed immediately
โข Server B comes back
โข Still has old data
โข No tombstone to stop it
โข Zombie! Data resurrects! ๐ป
With grace period:
โข Tombstone kept for 10 days
โข Server B comes back (day 5)
โข Sees tombstone!
โข Deletes its copy
โข Everyone in sync
โข Day 11: Tombstone removed safely โ
Perfect analogy! Just like library gives you grace to return books, Cassandra gives servers grace to see deletions!
Definition
gc_grace_seconds = Tombstone lifespan
Full name:
"Garbage Collection Grace Seconds"
What it controls:
How long tombstones live
before compaction removes them
Configured per table:
Each table can have different value
Default: 864,000 seconds = 10 days
Purpose
Ensures safe tombstone removal
Protection for:
โข Offline servers
โข Network partitions
โข Maintenance windows
โข Failed nodes recovering
Guarantees:
All servers see deletion
before tombstone removed
Result: No zombie data!
How It Works
Step-by-step:
1. Tombstone created
Timestamp: T (deletion time)
2. Compaction runs
Current time: C
3. Check age
Age = C - T
4. Decision
If Age > gc_grace_seconds:
โ Remove tombstone โ
Otherwise: Keep it!
โ Why 10 Days Default?
The Reasoning Behind 10 Days
10 days (864,000 seconds) was chosen based on real-world failure scenarios:
โ
Typical downtime scenarios it covers:
1. Planned Maintenance (1-4 hours):
โข OS updates: 1-2 hours
โข Hardware upgrades: 2-4 hours
โข 10 days: Massive overkill โโโ
2. Network Issues (minutes to hours):
โข Router failures: 30 minutes
โข Switch replacement: 2 hours
โข Data center connectivity: 4-8 hours
โข 10 days: Very safe โโ
3. Hardware Failures (1-3 days):
โข Disk failure: 1 day (replace + restore)
โข Server failure: 2 days (ship new hardware)
โข Power supply: 1-2 days
โข 10 days: Comfortable margin โ
4. Catastrophic Failures (3-7 days):
โข Data center outage: 3-5 days
โข Multiple simultaneous failures: 5-7 days
โข Emergency procurement: 7 days
โข 10 days: Just enough โ
Additional factors:
โข Weekend buffer: Friday failure + weekend = Monday response (3 days)
โข Holidays: Extended weekend, slower response
โข Human error margin: Mistakes in diagnosis/repair
โข Multi-region: Cross-region failures take longer
The philosophy: "Better safe than sorry"
10 days provides enough buffer for 99% of real-world failures while being short enough to not cause severe tombstone accumulation.
Trade-off balance:
โข Too short (1 day): Risk zombie data
โข Too long (30 days): Tombstone accumulation problems
โข 10 days: Sweet spot! โ
๐ฎ Interactive Grace Period Timeline
Day 0 of 10
Ready to start...
๐งฎ The Math (Simple Calculation!)
Converting Seconds to Days (Easy!)
Default value: 864,000 seconds
Let's break it down step-by-step:
Step 1: Seconds โ Minutes
864,000 seconds รท 60 = 14,400 minutes
(60 seconds in 1 minute)
Step 2: Minutes โ Hours
14,400 minutes รท 60 = 240 hours
(60 minutes in 1 hour)
Step 3: Hours โ Days
240 hours รท 24 = 10 days โ
(24 hours in 1 day)
Quick formula:
Days = Seconds รท 86,400
(86,400 = seconds in 1 day)
Examples:
โข 86,400 seconds = 1 day
โข 259,200 seconds = 3 days
โข 864,000 seconds = 10 days (default)
โข 2,592,000 seconds = 30 days
Remember: 86,400 is your magic number!
Multiply days by 86,400 to get seconds.
โ ๏ธ Setting gc_grace Too Short (Danger!)
๐ The Zombie Apocalypse Scenario
Real disaster: gc_grace_seconds = 0 (NO grace period!)
Timeline of disaster:
Monday 9:00 AM:
โข User deletes account (user_id=12345)
โข Tombstone created on Servers A, B, C
โข gc_grace_seconds = 0 (someone thought "faster is better!")
Monday 9:05 AM:
โข Server B goes down for emergency maintenance
โข Kernel panic, needs reboot
Monday 9:30 AM:
โข Compaction runs on Servers A & C
โข Tombstone age: 30 minutes (1,800 seconds)
โข Grace period: 0 seconds
โข 1,800 > 0 โ Tombstone REMOVED!
โข Servers A & C: No trace of deletion now
Monday 10:00 AM:
โข Server B back online
โข Still has: user_id=12345 (full account data)
โข Anti-entropy repair starts...
Monday 10:05 AM:
โข Repair: "Server B has user 12345, A & C don't!"
โข B thinks: "They lost the data, I'll restore it!"
โข B copies data to A & C
โข ZOMBIE! Account resurrected! ๐ป
Consequences:
โข User deleted account โ Still exists!
โข GDPR violation (right to be forgotten)
โข Legal liability
โข User trust destroyed
โข Potential lawsuit
With gc_grace = 10 days:
โข Tombstone kept for 10 days
โข Server B comes back after 1 hour
โข Sees tombstone on A & C
โข Adopts tombstone
โข Deletes its copy
โข Day 11: Tombstone removed safely โ
โข No zombie!
gc_grace = 0 seconds
Worst case! NEVER use in production!
What happens:
โข Tombstone removed immediately
โข Any server offline = risk
โข Even 5 minutes down = zombie!
When to use (rare):
โข Single-node testing only
โข Never run repair
โข Throwaway data
Production: DON'T DO THIS!
gc_grace = 1 day (86,400)
Risky but sometimes acceptable
Risk:
โข Server down >24 hours = zombie
โข Weekend failures problematic
โข Hardware delivery takes days
Only if:
โข Extremely stable cluster
โข 24/7 monitoring
โข Rapid response team
โข Frequent repairs (every 12h)
Requires operational excellence!
Real Failure Statistics
Actual downtime data:
Planned maintenance:
โข 95%: < 4 hours โ
โข 5%: 4-24 hours
Hardware failures:
โข 70%: < 1 day
โข 25%: 1-3 days
โข 5%: 3-7 days โ ๏ธ
If gc_grace = 1 day:
30% of hardware failures = zombie!
10 days covers 95% of failures!
๐ Setting gc_grace Too Long (Performance!)
๐ The Slow Death by Tombstones
Disaster scenario: gc_grace_seconds = 90 days (someone thought "safer is better!")
The setup:
โข Time-series data table (sensor readings)
โข TTL = 30 days (auto-delete old readings)
โข gc_grace_seconds = 7,776,000 (90 days!)
โข 1 million rows/day inserted
What happens over time:
Day 30:
โข First batch of rows expire (TTL)
โข 1M rows โ tombstones
โข Total tombstones: 1M
Day 60:
โข Another 30M rows expired
โข Total tombstones: 31M
โข Queries starting to slow...
Day 90:
โข Total tombstones: 61M
โข Active data: 30M rows
โข Tombstone ratio: 67%!
โข P99 latency: 50ms โ 400ms (8x slower)
Day 120:
โข Grace period FINALLY expires for day 30 tombstones
โข But day 31-90 tombstones still accumulating!
โข Total tombstones: 91M
โข Active data: 30M
โข Ratio: 75% tombstones!
โข P99: 800ms (16x slower!)
โข Warnings: "Read 500000 tombstones" constantly
Crisis point:
โข Queries timing out
โข Disk 90% full (tombstones!)
โข Compaction can't keep up
โข Cluster degraded
With gc_grace = 10 days:
โข TTL expires: 30 days
โข Grace period: +10 days
โข Total: 40 days max
โข Tombstones: 10M (vs 91M!)
โข Performance: Normal โ
Read Performance
Tombstones slow every query!
Example query:
SELECT * FROM sensors LIMIT 1000
With 10-day grace:
โข Scan 5,000 tombstones
โข Find 1,000 live rows
โข Time: 50ms
With 90-day grace:
โข Scan 500,000 tombstones!
โข Find 1,000 live rows
โข Time: 800ms (16x slower!)
More tombstones = linear slowdown!
Disk Space
Tombstones waste disk space!
Calculation:
โข 100M tombstones
โข ~20 bytes each
โข Total: 2GB wasted!
Impact:
โข Disk fills up
โข Can't add new data
โข Emergency expansion needed
โข Expensive!
Shorter grace = more free space!
Compaction Load
More tombstones = slower compaction
With 10-day grace:
โข Process 50GB data
โข Compaction: 30 minutes
With 90-day grace:
โข Process 50GB data + 30GB tombstones
โข Compaction: 1.5 hours (3x!)
Vicious cycle:
Slow compaction โ tombstones accumulate faster โ even slower!
โ๏ธ Tuning Guide (How to Configure)
Configuration Commands
-- View current gc_grace_seconds DESC TABLE users; -- Set gc_grace to 1 day (86,400 seconds) ALTER TABLE users WITH gc_grace_seconds = 86400; -- Set gc_grace to 3 days ALTER TABLE events WITH gc_grace_seconds = 259200; -- Set gc_grace to 7 days ALTER TABLE logs WITH gc_grace_seconds = 604800; -- Keep default 10 days (no action needed) -- Default: 864000 seconds
Decision Matrix
Keep default (10 days) if:
โข Standard workload โ
โข Unsure what to set
โข Low delete volume
โข Value stability > performance
Reduce to 1-3 days if:
โข High-delete workload
โข Stable cluster (99.9% uptime)
โข 24/7 monitoring
โข Frequent repairs
Increase to 14-30 days if:
โข Volatile environment
โข Slow response times
โข Critical data (no zombie risk)
When in doubt: keep default!
By Workload Type
Time-series with TTL:
โข gc_grace: 1-2 days
โข High tombstone volume
โข Fast removal critical
User data (CRUD):
โข gc_grace: 10 days (default)
โข Occasional deletes
โข Safety important
Append-only logs:
โข gc_grace: 10 days
โข Rarely delete
โข Default fine
Session data:
โข gc_grace: 3-5 days
โข Moderate deletes
โข Balanced approach
Monitoring Required
If reducing gc_grace, monitor:
1. Node downtime:
Alert if down > 50% of gc_grace
2. Repair frequency:
Must run at least every gc_grace period
3. Tombstone warnings:
Watch logs for zombie data
4. Query latency:
Improved after tuning? โ
No monitoring = keep default!
๐ข Real Company Examples
Uber: Aggressive 1-Day Tuning
Challenge: 270 billion time-series rows with 30-day TTL creating massive tombstone accumulation with default 10-day grace period.
Decision: Reduce to 1 day (86,400 seconds)
Configuration:
ALTER TABLE trip_locations WITH gc_grace_seconds = 86400;
Results:
โข Tombstone lifespan: 40 days โ 31 days (30-day TTL + 1-day grace)
โข Tombstone count: 90B โ 9B (10x reduction!)
โข P99 latency: 800ms โ 45ms (18x faster!)
โข Disk usage: 80% tombstones โ 15%
Requirements met:
โ 24/7 monitoring with <12h downtime alerts
โ Repair every 12 hours (vs weekly)
โ Highly stable cluster (99.95% uptime)
โ 2 years production: zero zombie incidents
Key lesson: "For stable, well-monitored clusters, 1-day grace is viable and delivers massive performance gains."
Netflix: Conservative 10-Day Default
Philosophy: "Safety first - we can always optimize later."
Approach: Keep 10-day default for all tables
Rationale:
โข Multi-region deployment (cross-region delays)
โข Occasional multi-day outages happen
โข GDPR compliance critical (no zombie resurrections)
โข Performance managed through other optimizations
Alternative optimizations:
โข Time-bucketed partitions (avoid tombstones entirely)
โข Soft delete + async cleanup
โข Aggressive compaction tuning
โข Partition drops instead of TTL
Results:
โข Zero zombie incidents in 5 years โ
โข Performance good through design, not gc_grace tuning
โข Peace of mind for operations team
Key lesson: "Don't tune gc_grace as first optimization - design around deletions instead."
Discord: The 90-Day Disaster
Mistake: Engineer set gc_grace to 90 days "to be extra safe."
Table: user_read_status with 30-day TTL
Configuration: gc_grace_seconds = 7776000 (90 days)
Timeline of pain:
โข Month 1: No issues yet
โข Month 2: Queries slowing (tombstones accumulating)
โข Month 3: P99 latency 15ms โ 800ms
โข Month 4: Crisis - queries timing out, disk 95% full
Peak problem:
โข 200 billion tombstones
โข 75% of cluster storage = tombstones
โข Scanning 50M tombstones per query
โข Compaction: 24/7, still can't keep up
Fix applied:
1. Reduced gc_grace: 90 days โ 1 day
2. Forced major compaction
3. Redesigned to time-bucketed partitions
Recovery results:
โข After compaction: Tombstones 200B โ 2M (99.999% reduction!)
โข P99 latency: 800ms โ 18ms (44x faster!)
โข Disk freed: 2TB
Lesson learned: "Longer gc_grace is NOT always safer - it causes different disasters. Understand the trade-offs!"
โ Best Practices
1. Start with Default
10 days is proven and safe
Don't tune until you:
โข Measure actual problem
โข Understand trade-offs
โข Have monitoring in place
Premature optimization = danger!
Default works for 90% of use cases โ
2. Monitor First
Before tuning, measure:
โข Tombstone warnings in logs
โข Query latency trends
โข Disk usage growth
โข Compaction frequency
Data-driven decisions only!
No monitoring = keep default
3. Tune Per Table
Different tables, different needs
Time-series: 1-2 days
User data: 10 days
Logs: 10 days
Sessions: 3-5 days
One-size does NOT fit all!
Analyze each table separately
4. Pair with Repair
gc_grace and repair go together
Rule: Repair every gc_grace period
If gc_grace = 3 days:
Run repair every 3 days!
This ensures deletions spread
before tombstones removed โ
5. Never Use 0
gc_grace_seconds = 0 is DANGEROUS
Only acceptable for:
โข Single-node testing
โข Throwaway data
โข Never repair
Production: NEVER!
Even 1 day is better than 0
6. Design > Tuning
Better: Avoid tombstones entirely!
Instead of tuning gc_grace:
โข Time-bucketed partitions
โข Soft deletes
โข Partition drops
โข Design without deletes
Architecture > configuration!
๐จ Common Mistakes
โ Setting to 0 for "performance"
โ Zombie data guaranteed! Keep at least 1 day.
โ Setting to 90 days for "safety"
โ Tombstone explosion! Slower queries, disk full.
โ Tuning without monitoring
โ Can't measure impact. Keep default instead.
โ Same value for all tables
โ Different workloads need different settings.
โ Reducing without frequent repairs
โ Repair must run within grace period!
โ When in doubt: Keep the default 10 days!