JVM Tuning in Cassandra
Give Cassandra a performance makeover! JVM tuning optimizes memory and garbage collection so your database runs faster, avoids slowdowns, and handles large-scale data like a champ.
📖 The Story: Tom's 10-Second GC Pause Nightmare
Tom's Cassandra cluster handled 1 million requests per second perfectly... until one day, everything froze for 10 seconds. Then it happened again. And again. Users complained. Monitoring showed: "Stop-the-world GC pause: 10,247ms". Here's what went wrong...
💥 The Disaster: 32GB Heap Size
Tom's Original Configuration:
What Happened:
- 🕐 First Week: Everything fine, performance great
- 💥 Week 2: First 10-second freeze during traffic spike
- 😱 Week 3: Freezes happening hourly
- ⏱️ GC Logs: "Full GC (Allocation Failure) 10.2sec"
- 💀 User Impact: Timeouts, 500 errors, angry customers
- 📉 Monitoring: 99.9% → 90% availability!
Why 32GB Heap Failed:
- Huge Heap: 32GB of objects to scan during GC
- Old Generation Full: 28GB of old gen objects accumulated
- Full GC Triggered: Must scan ALL 32GB objects
- Stop-the-World: Entire JVM pauses (no reads, no writes!)
- 10 Second Pause: Scanning 32GB takes 10+ seconds
- All Threads Frozen: Cassandra completely unresponsive
✅ The Fix: 12GB Heap + G1GC!
Tom's New Configuration:
What Changed:
- Smaller Heap: 12GB instead of 32GB
- Faster GC: Scanning 12GB takes 200ms (not 10 seconds!)
- G1GC Algorithm: Concurrent, incremental collection
- Predictable Pauses: Target 200ms, never exceed 500ms
- More RAM for OS Cache: 52GB for OS page cache (was 32GB)
- Better Performance: OS cache improves read speeds!
The Results:
- ⚡ GC Pauses: 10 seconds → 150ms (66x faster!)
- ✅ P99 GC Time: < 200ms (predictable!)
- 🚀 Zero Freezes: No more 10-second hangs!
- 📈 Availability: 90% → 99.99% (back to normal!)
- 💚 Read Performance: Actually improved! (more OS cache)
- 😊 Customers Happy: No more timeouts!
Tom learned: More heap ≠ better! Keep it 8-16GB max! 🎉
☕ JVM Tuning Fundamentals
Understand how the JVM affects Cassandra performance!
🎯 Why JVM Tuning Matters
Cassandra runs on the JVM. During garbage collection (GC), the JVM pauses ALL threads to clean up memory. If GC takes 10 seconds, Cassandra is completely frozen for 10 seconds - no reads, no writes, no repairs. Proper JVM tuning keeps GC pauses under 200ms!
The Garbage Collection Problem
The Heap Size Paradox
Too Large (32GB+)
- GC Pauses: 5-10+ seconds
- Stop-the-World: Complete freeze
- Timeouts: All requests fail
- Memory Scan: Must check 32GB
- Result: Production outage!
Never use > 16GB heap!
Optimal (8-16GB)
- GC Pauses: 50-200ms
- Predictable: Consistent performance
- No Timeouts: Requests succeed
- Fast Scan: Only 8-16GB to check
- Result: Production stable!
Sweet spot: 8-16GB!
Too Small (< 4GB)
- GC Pauses: 10-50ms (good!)
- But... Frequent GC cycles
- High CPU: 20% CPU on GC
- OOM Risk: Out of memory errors
- Result: Unstable!
Too small = frequent GC
The Golden Rule
Heap Size: 8-16GB regardless of total server RAM!
64GB server? Use 12GB heap.
256GB server? Still use 12-16GB heap!
Give the rest to OS page cache!
📏 Heap Sizing Guide
How to choose the right heap size for your server!
Heap Sizing by Server RAM
| Total RAM | Heap Size | OS Cache | Notes |
|---|---|---|---|
| 16GB | 8GB | ~7GB | Minimum production |
| 32GB | 8-12GB | ~20-23GB | Recommended |
| 64GB | 12-16GB ⭐ | ~47-51GB | Ideal balance |
| 128GB | 16GB | ~110GB | Huge OS cache! |
| 256GB | 16-24GB | ~230GB | Max heap 24GB |
Why Not Exceed 32GB Heap?
Problem 1: Long GC Pauses
GC time proportional to heap size. 32GB heap = 10+ second pauses!
Problem 2: Compressed Oops Lost
JVM uses 64-bit pointers if heap > 32GB, increasing memory overhead by 20-50%!
Problem 3: OS Cache Starvation
Large heap leaves little RAM for OS page cache, hurting read performance!
Never Exceed These Limits
- ❌ Never > 32GB: Loses compressed oops, massive GC pauses
- ❌ Never > 16GB: For most workloads (8-16GB is optimal)
- ❌ Never > 50% RAM: Starves OS cache
- ✅ Sweet Spot: 8-16GB regardless of total RAM
Configuration Example
🔄 G1GC - The Right Garbage Collector
Why G1GC is best for Cassandra!
⚡ G1GC: Garbage-First Garbage Collector
G1GC divides heap into regions and collects garbage incrementally, focusing on regions with most garbage first. It can meet pause time targets (200ms) and runs mostly concurrently. Perfect for Cassandra!
G1GC vs CMS (Old Default)
CMS (Old)
- Fragmentation: Over time, heap fragments
- Full GC: Falls back to stop-the-world
- Unpredictable: Pauses vary wildly
- Deprecated: Removed in Java 14
- Result: Occasional 5-10s pauses
Don't use CMS anymore!
G1GC (Current)
- Compacting: Eliminates fragmentation
- Concurrent: Most work happens in background
- Predictable: Meets pause time targets
- Default: Since Java 9
- Result: Consistent 50-200ms pauses
Use G1GC! (default on Java 9+)
ZGC / Shenandoah
- Low Latency: < 10ms pauses
- Huge Heaps: Can handle TB+ heaps
- Still Experimental: For Cassandra
- More CPU: Higher overhead
- Result: Future option
For future (not yet recommended)
G1GC Configuration
How G1GC Works
1. Young Generation Collection (Fast)
Happens frequently, collects short-lived objects
- Pause time: 10-50ms
- Frequency: Every few seconds
- Impact: Minimal
2. Concurrent Marking (Background)
Runs in background, marks live objects
- Mostly concurrent (no pauses)
- Identifies garbage regions
- Prepares for mixed collections
3. Mixed Collection (Incremental)
Collects old + young regions with most garbage
- Pause time: 50-200ms
- Frequency: As needed
- Collects regions with most garbage first
4. Full GC (Rare, Last Resort)
Only if heap is exhausted (should be rare!)
- Pause time: 500ms-2s
- Should happen < 1/day
- If frequent: increase heap or reduce load
G1GC Benefits for Cassandra
- ✅ Predictable Pauses: Meets 200ms target consistently
- ✅ No Fragmentation: Compacts heap incrementally
- ✅ Concurrent: Most work in background
- ✅ Handles Large Heaps: Better than CMS for 8-16GB
- ✅ Self-Tuning: Adapts to workload automatically
📊 GC Logging & Monitoring
Essential for troubleshooting GC issues!
Enable GC Logging
Check GC Statistics
Reading GC Logs
GC Monitoring Checklist
| Metric | Healthy | Warning | Critical |
|---|---|---|---|
| Max GC Pause | < 200ms ✅ | 200-500ms ⚠️ | > 1000ms ❌ |
| P99 GC Pause | < 150ms ✅ | 150-300ms ⚠️ | > 500ms ❌ |
| Time in GC | < 5% ✅ | 5-10% ⚠️ | > 10% ❌ |
| Full GC Frequency | < 1/day ✅ | 1/hour ⚠️ | > 1/minute ❌ |
| Heap Usage | 50-70% ✅ | 70-85% ⚠️ | > 90% ❌ |
🎛️ Complete JVM Tuning Guide
Production-ready jvm.options configuration!
Recommended jvm.options Template
Configuration by Server Size
Small Server (16GB RAM)
Medium Server (32GB RAM) ⭐ Recommended
Large Server (64GB+ RAM)
Extra-Large Server (128GB+ RAM)
🔧 Troubleshooting GC Issues
How to diagnose and fix JVM problems!
Common GC Problems & Solutions
Problem 1: Long GC Pauses (> 1s)
Symptoms:
- GC pauses > 1000ms
- Clients timeout
- Nodes marked DOWN temporarily
Diagnosis:
Solution:
Problem 2: Frequent GC (> 10% time in GC)
Symptoms:
- GC happening constantly
- High CPU usage
- Degraded performance
Diagnosis:
Solution:
Problem 3: OutOfMemoryError
Symptoms:
- Cassandra crashes with OOM
- Heap dump generated
- Node won't start
Diagnosis:
Solution:
- Increase heap to 14-16GB (if < 12GB)
- Fix data model (partition sizes)
- Reduce row cache size
- Check for application memory leaks
Problem 4: Unpredictable GC Pauses
Symptoms:
- GC pauses vary from 50ms to 2000ms
- Sometimes fast, sometimes slow
- Using CMS collector
Solution:
💼 Interview Questions & Expert Answers
Master JVM tuning for your interview!
Answer: GC pause time is proportional to heap size. A 32GB+ heap causes 5-10 second stop-the-world pauses that freeze Cassandra completely, while 8-16GB keeps pauses under 200ms. Extra RAM goes to OS page cache which improves performance without GC penalties.
The Problem with Large Heaps:
- GC Must Scan All Objects: During GC, JVM must examine every object in heap
- Time Proportional to Size: 32GB heap takes 4x longer than 8GB heap
- Stop-the-World: ALL Cassandra threads freeze during GC
- User Impact: 10-second pause = 10 seconds of complete unresponsiveness
- Compressed Oops Lost: Heaps > 32GB lose pointer compression (20-50% overhead)
Performance Comparison:
| Heap Size | GC Pause | User Experience |
|---|---|---|
| 8GB | 50-150ms ✅ | Imperceptible |
| 12GB | 100-200ms ✅ | Barely noticeable |
| 16GB | 150-300ms ⚠️ | Slight delay |
| 32GB | 5-10s ❌ | Complete freeze! |
| 64GB | 10-20s ❌ | Catastrophic! |
Why Not Use All RAM for Heap?
Cassandra benefits more from OS page cache than from large heap!
Key Takeaway: Cassandra's architecture is designed for large OS caches, not large heaps. Keep heap 8-16GB and let OS cache do the heavy lifting!
Answer: G1GC (Garbage-First) is a concurrent, compacting collector with predictable pause times (< 200ms), while CMS (Concurrent Mark-Sweep) is a non-compacting collector prone to fragmentation and unpredictable Full GC pauses. G1GC is better for Cassandra's workload.
Key Differences:
| Feature | CMS | G1GC |
|---|---|---|
| Compacting | No ❌ | Yes ✅ |
| Fragmentation | Severe over time | Minimal |
| Pause Times | Unpredictable | Predictable (target) |
| Full GC | Frequent (when fragmented) | Rare |
| Heap Support | < 8GB optimal | 8-16GB optimal |
| Status | Deprecated (Java 14) | Default (Java 9+) |
Why CMS Fails for Cassandra:
Problem 1: Fragmentation
- CMS doesn't compact memory (no defragmentation)
- Over days/weeks, heap becomes fragmented
- Eventually can't allocate large objects
- Triggers Full GC (stop-the-world)
- Full GC can take 5-10+ seconds!
Problem 2: Concurrent Mode Failure
If heap fills up before CMS completes, it falls back to Full GC:
Why G1GC is Better:
Advantage 1: Predictable Pauses
- Set target: -XX:MaxGCPauseMillis=200
- G1GC tries to meet target (usually does!)
- Pauses stay consistent: 50-200ms
- No surprise 10-second pauses
Advantage 2: Automatic Compaction
- G1GC compacts memory during mixed collections
- No fragmentation buildup
- Full GC extremely rare (< 1/day)
- Stable performance over time
Advantage 3: Region-Based
- Heap divided into regions (1-32MB each)
- Collects regions with most garbage first
- Can meet pause time targets incrementally
- Better than CMS's generational approach
Configuration Example:
Key Takeaway: G1GC is the standard for Cassandra since Java 9. CMS is deprecated and causes fragmentation issues. Always use G1GC!
Answer: Check GC logs for Full GC frequency, analyze heap usage patterns, verify heap size is 8-16GB, ensure G1GC is enabled, check for memory leaks (large partitions, too many SSTables), and monitor if heap is consistently > 85% full.
Step-by-Step Troubleshooting:
Step 1: Check GC Statistics
Step 2: Analyze GC Logs
Step 3: Common Causes & Solutions
Cause 1: Heap Too Large
Cause 2: Using CMS Instead of G1GC
Cause 3: Memory Leak (Large Partitions)
Cause 4: Too Many SSTables
Cause 5: Row Cache Too Large
Step 4: Monitor After Changes
Decision Tree:
- Heap > 16GB? → Reduce to 12GB
- Using CMS? → Switch to G1GC
- Heap > 85% full? → Increase slightly or reduce load
- Large partitions? → Fix data model
- Many SSTables? → Run compaction
- Large row cache? → Reduce size
Answer: Setting -Xms = -Xmx prevents expensive heap resizing operations, provides predictable GC behavior, and ensures the OS doesn't reclaim committed memory. Dynamic sizing causes performance variability and additional GC overhead.
Why -Xms = -Xmx?
Advantage 1: No Heap Resizing
When -Xms < -Xmx, JVM can resize heap dynamically:
- Growing heap: Expensive operation (can take 100s of ms)
- Shrinking heap: Requires Full GC (stop-the-world)
- Unpredictable pauses during resize
- With -Xms = -Xmx: Heap size fixed, no resizing ever!
Advantage 2: Predictable Performance
Advantage 3: OS Memory Commitment
When -Xms < -Xmx:
- OS may not commit all memory upfront
- Pages allocated on-demand (can be slow)
- OS might swap unused heap pages
- Risk of OOM if system RAM exhausted
When -Xms = -Xmx:
- OS commits all memory at startup
- All pages allocated upfront (use -XX:+AlwaysPreTouch)
- No on-demand allocation delays
- Guaranteed memory availability
Advantage 4: Faster Startup
Real-World Impact:
| Configuration | Behavior | Production Use |
|---|---|---|
| -Xms4G -Xmx12G | Dynamic, unpredictable | ❌ Never use |
| -Xms12G -Xmx12G | Fixed, predictable | ✅ Always use |
Recommended Configuration:
Key Takeaway: -Xms = -Xmx eliminates heap resizing overhead and provides predictable, consistent performance. Always use equal values in production!
Answer: Still use 16GB heap (not proportional to RAM!), enable G1GC with 200ms pause target, enable GC logging, heap dump on OOM, and AlwaysPreTouch. The key is keeping heap small (16GB) regardless of total RAM, leaving 110GB+ for OS page cache.
Complete Configuration for 128GB Server:
Memory Allocation Breakdown:
Total: 128GB RAM
- JVM Heap: 16GB (12.5%)
- Row Cache: 2-4GB (optional, 3%)
- Key Cache: 1GB (1%)
- OS Page Cache: 107-109GB (84%) ⭐
Why This Configuration?
Heap Size Rationale:
- 16GB keeps GC pauses < 200ms consistently
- 32GB would cause 5-10 second pauses!
- 64GB would cause 20+ second pauses!
- 16GB is optimal regardless of total RAM
OS Cache Strategy:
- 110GB OS cache can cache HUGE portion of SSTables
- 100GB dataset → 100% cached in RAM!
- 500GB dataset → 22% cached (hot data)
- OS cache has zero GC overhead
- Benefits ALL tables, not just one
G1GC Configuration:
- MaxGCPauseMillis=200: Target 200ms pauses
- InitiatingHeapOccupancyPercent=70: Start concurrent GC at 70% full
- G1ReservePercent=25: Reserve 25% to avoid evacuation failures
- Result: Consistent, predictable GC behavior
Common Mistake to Avoid:
Key Takeaway: More RAM doesn't mean bigger heap! Keep heap 8-16GB regardless of total RAM. Give extra RAM to OS page cache for massive performance gains!
🎓 Chapter Summary: JVM Tuning Mastery
You now understand JVM tuning at a production level!
The Golden Rules:
- ☕ Heap Size: 8-16GB - Regardless of total RAM!
- ♻️ Use G1GC - Predictable 50-200ms pauses
- ⚖️ -Xms = -Xmx - No dynamic resizing
- ❌ Never > 32GB heap - Causes 10+ second pauses
- 💾 Maximize OS cache - Give extra RAM to OS
Quick Configuration:
Target Metrics:
- ✅ Max GC pause < 200ms
- ✅ P99 GC pause < 150ms
- ✅ Time in GC < 5%
- ✅ Full GC < 1/day
Remember Tom: 32GB heap = disaster, 12GB heap = success! 🚀
Responsive Ad