Advanced Performance Topics

JVM Tuning in Cassandra

Give Cassandra a performance makeover! JVM tuning optimizes memory and garbage collection so your database runs faster, avoids slowdowns, and handles large-scale data like a champ.

📖 The Story: Tom's 10-Second GC Pause Nightmare

Tom's Cassandra cluster handled 1 million requests per second perfectly... until one day, everything froze for 10 seconds. Then it happened again. And again. Users complained. Monitoring showed: "Stop-the-world GC pause: 10,247ms". Here's what went wrong...

💥 The Disaster: 32GB Heap Size

Tom's Original Configuration:

-- Tom's jvm.options (WRONG!) -Xms32G -Xmx32G # "More heap = better performance, right?" ❌

What Happened:

  • 🕐 First Week: Everything fine, performance great
  • 💥 Week 2: First 10-second freeze during traffic spike
  • 😱 Week 3: Freezes happening hourly
  • ⏱️ GC Logs: "Full GC (Allocation Failure) 10.2sec"
  • 💀 User Impact: Timeouts, 500 errors, angry customers
  • 📉 Monitoring: 99.9% → 90% availability!

Why 32GB Heap Failed:

  1. Huge Heap: 32GB of objects to scan during GC
  2. Old Generation Full: 28GB of old gen objects accumulated
  3. Full GC Triggered: Must scan ALL 32GB objects
  4. Stop-the-World: Entire JVM pauses (no reads, no writes!)
  5. 10 Second Pause: Scanning 32GB takes 10+ seconds
  6. All Threads Frozen: Cassandra completely unresponsive
-- GC log showing disaster: /* [Full GC (Allocation Failure) [PSYoungGen: 1024K->0K(2048K)] [ParOldGen: 31457280K->28934567K(32768000K)] 31458304K->28934567K(32770048K), 10.2471893 secs] [Times: user=42.18 sys=0.89, real=10.25 secs] ENTIRE CLUSTER FROZEN FOR 10.25 SECONDS! 💥 */

✅ The Fix: 12GB Heap + G1GC!

Tom's New Configuration:

-- Optimized jvm.options -Xms12G -Xmx12G -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:G1RSetUpdatingPauseTimePercent=5 # Small enough to GC quickly! # G1GC = predictable pauses!

What Changed:

  1. Smaller Heap: 12GB instead of 32GB
  2. Faster GC: Scanning 12GB takes 200ms (not 10 seconds!)
  3. G1GC Algorithm: Concurrent, incremental collection
  4. Predictable Pauses: Target 200ms, never exceed 500ms
  5. More RAM for OS Cache: 52GB for OS page cache (was 32GB)
  6. Better Performance: OS cache improves read speeds!

The Results:

  • ⚡ GC Pauses: 10 seconds → 150ms (66x faster!)
  • ✅ P99 GC Time: < 200ms (predictable!)
  • 🚀 Zero Freezes: No more 10-second hangs!
  • 📈 Availability: 90% → 99.99% (back to normal!)
  • 💚 Read Performance: Actually improved! (more OS cache)
  • 😊 Customers Happy: No more timeouts!

Tom learned: More heap ≠ better! Keep it 8-16GB max! 🎉

☕ JVM Tuning Fundamentals

Understand how the JVM affects Cassandra performance!

🎯 Why JVM Tuning Matters

Cassandra runs on the JVM. During garbage collection (GC), the JVM pauses ALL threads to clean up memory. If GC takes 10 seconds, Cassandra is completely frozen for 10 seconds - no reads, no writes, no repairs. Proper JVM tuning keeps GC pauses under 200ms!

The Garbage Collection Problem

GC Pause Impact Processing GC PAUSE 10 seconds! ❄️ FROZEN ❄️ Resume What Happens During GC Pause: ❌ All read requests timeout ❌ All write requests timeout ❌ Coordinator marked as DOWN ❌ Clients retry on other nodes ❌ Cluster appears unhealthy ❌ Users see 500 errors ✅ Target: GC < 200ms ✅ Users don't notice ✅ No timeouts ✅ Cluster stays healthy

The Heap Size Paradox

💥

Too Large (32GB+)

  • GC Pauses: 5-10+ seconds
  • Stop-the-World: Complete freeze
  • Timeouts: All requests fail
  • Memory Scan: Must check 32GB
  • Result: Production outage!

Never use > 16GB heap!

✅

Optimal (8-16GB)

  • GC Pauses: 50-200ms
  • Predictable: Consistent performance
  • No Timeouts: Requests succeed
  • Fast Scan: Only 8-16GB to check
  • Result: Production stable!

Sweet spot: 8-16GB!

⚠️

Too Small (< 4GB)

  • GC Pauses: 10-50ms (good!)
  • But... Frequent GC cycles
  • High CPU: 20% CPU on GC
  • OOM Risk: Out of memory errors
  • Result: Unstable!

Too small = frequent GC

The Golden Rule

Heap Size: 8-16GB regardless of total server RAM!
64GB server? Use 12GB heap.
256GB server? Still use 12-16GB heap!
Give the rest to OS page cache!

📏 Heap Sizing Guide

How to choose the right heap size for your server!

Heap Sizing by Server RAM

Total RAM Heap Size OS Cache Notes
16GB 8GB ~7GB Minimum production
32GB 8-12GB ~20-23GB Recommended
64GB 12-16GB ⭐ ~47-51GB Ideal balance
128GB 16GB ~110GB Huge OS cache!
256GB 16-24GB ~230GB Max heap 24GB

Why Not Exceed 32GB Heap?

Problem 1: Long GC Pauses

GC time proportional to heap size. 32GB heap = 10+ second pauses!

-- Real GC log from 32GB heap: /* [Full GC 32G->28G(32G), 10.247 secs] 10 seconds where: - No reads processed - No writes processed - Coordinator appears down - All clients timeout */

Problem 2: Compressed Oops Lost

JVM uses 64-bit pointers if heap > 32GB, increasing memory overhead by 20-50%!

-- Heap ≤ 32GB: Compressed pointers (4 bytes) -XX:+UseCompressedOops # Automatic if ≤ 32GB -- Heap > 32GB: Full pointers (8 bytes) # 2x pointer size = 20-50% more memory per object! # Effectively get LESS usable memory!

Problem 3: OS Cache Starvation

Large heap leaves little RAM for OS page cache, hurting read performance!

-- 64GB server with 48GB heap (BAD!): Heap: 48GB OS Cache: 14GB ← Tiny! Slow reads! -- 64GB server with 12GB heap (GOOD!): Heap: 12GB OS Cache: 50GB ← Huge! Fast reads!

Never Exceed These Limits

  • ❌ Never > 32GB: Loses compressed oops, massive GC pauses
  • ❌ Never > 16GB: For most workloads (8-16GB is optimal)
  • ❌ Never > 50% RAM: Starves OS cache
  • ✅ Sweet Spot: 8-16GB regardless of total RAM

Configuration Example

-- jvm.options (64GB server) # Heap size: 12GB -Xms12G -Xmx12G # Why Xms = Xmx? # - Prevents heap resizing (expensive!) # - More predictable GC behavior # - Faster startup # Result: # - JVM heap: 12GB # - Row cache: 2GB (optional) # - Key cache: 1GB # - OS page cache: 49GB (automatic!)

🔄 G1GC - The Right Garbage Collector

Why G1GC is best for Cassandra!

⚡ G1GC: Garbage-First Garbage Collector

G1GC divides heap into regions and collects garbage incrementally, focusing on regions with most garbage first. It can meet pause time targets (200ms) and runs mostly concurrently. Perfect for Cassandra!

G1GC vs CMS (Old Default)

⚠️

CMS (Old)

  • Fragmentation: Over time, heap fragments
  • Full GC: Falls back to stop-the-world
  • Unpredictable: Pauses vary wildly
  • Deprecated: Removed in Java 14
  • Result: Occasional 5-10s pauses

Don't use CMS anymore!

✅

G1GC (Current)

  • Compacting: Eliminates fragmentation
  • Concurrent: Most work happens in background
  • Predictable: Meets pause time targets
  • Default: Since Java 9
  • Result: Consistent 50-200ms pauses

Use G1GC! (default on Java 9+)

🚀

ZGC / Shenandoah

  • Low Latency: < 10ms pauses
  • Huge Heaps: Can handle TB+ heaps
  • Still Experimental: For Cassandra
  • More CPU: Higher overhead
  • Result: Future option

For future (not yet recommended)

G1GC Configuration

-- jvm.options - G1GC configuration # Enable G1GC (default on Java 9+) -XX:+UseG1GC # Target max pause time (200ms recommended) -XX:MaxGCPauseMillis=200 # Max % of time spent in GC (default: 10) -XX:GCTimeRatio=9 # G1 region size (auto-calculated, usually 8-32MB) # -XX:G1HeapRegionSize=16M # How much time to spend updating remembered sets -XX:G1RSetUpdatingPauseTimePercent=5 # Initiating heap occupancy percent (default: 45) -XX:InitiatingHeapOccupancyPercent=70 # Reserve % of heap to avoid to-space exhaustion -XX:G1ReservePercent=25

How G1GC Works

1. Young Generation Collection (Fast)

Happens frequently, collects short-lived objects

  • Pause time: 10-50ms
  • Frequency: Every few seconds
  • Impact: Minimal

2. Concurrent Marking (Background)

Runs in background, marks live objects

  • Mostly concurrent (no pauses)
  • Identifies garbage regions
  • Prepares for mixed collections

3. Mixed Collection (Incremental)

Collects old + young regions with most garbage

  • Pause time: 50-200ms
  • Frequency: As needed
  • Collects regions with most garbage first

4. Full GC (Rare, Last Resort)

Only if heap is exhausted (should be rare!)

  • Pause time: 500ms-2s
  • Should happen < 1/day
  • If frequent: increase heap or reduce load

G1GC Benefits for Cassandra

  • ✅ Predictable Pauses: Meets 200ms target consistently
  • ✅ No Fragmentation: Compacts heap incrementally
  • ✅ Concurrent: Most work in background
  • ✅ Handles Large Heaps: Better than CMS for 8-16GB
  • ✅ Self-Tuning: Adapts to workload automatically

📊 GC Logging & Monitoring

Essential for troubleshooting GC issues!

Enable GC Logging

-- jvm.options - GC logging configuration # Enable GC logging (Java 9+) -Xlog:gc*:file=/var/log/cassandra/gc.log:time,uptime,level,tags:filecount=10,filesize=10M # Legacy format (Java 8): -Xloggc:/var/log/cassandra/gc.log -XX:+PrintGCDetails -XX:+PrintGCDateStamps -XX:+PrintGCApplicationStoppedTime -XX:+PrintPromotionFailure -XX:+UseGCLogFileRotation -XX:NumberOfGCLogFiles=10 -XX:GCLogFileSize=10M # Heap dump on OOM (for debugging) -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/log/cassandra/heapdump.hprof

Check GC Statistics

-- Check current GC stats nodetool gcstats -- Output: /* Interval (ms) Max (ms) Total (ms) Stdev (ms) GC # 1000 42 150 8.2 3 Max (ms): Longest GC pause in interval Total (ms): Total GC time in interval Stdev: Standard deviation GC #: Number of GC cycles */ -- Good GC stats: // Max < 200ms (pause time) // Total < 500ms per second (< 50% time in GC) // Stdev low (predictable) -- Bad GC stats: // Max > 1000ms (❌ freezing!) // Total > 500ms per second (❌ too much GC!)

Reading GC Logs

-- Good GC log (G1GC, 12GB heap): /* [2025-01-06T10:15:23.456+0000] GC(45) Pause Young (Normal) 2048M->512M(12288M) 42.123ms GOOD! 42ms pause, young collection only */ -- Warning GC log: /* [2025-01-06T10:20:15.789+0000] GC(78) Pause Full (Allocation Failure) 11264M->9876M(12288M) 1247.456ms ⚠️ WARNING! Full GC took 1.2 seconds If frequent → increase heap or reduce load */ -- Disaster GC log (32GB heap): /* [2025-01-06T10:25:42.123+0000] GC(102) Pause Full (Ergonomics) 31457M->28934M(32768M) 10247.893ms 💥 DISASTER! 10 second pause! Reduce heap to 12-16GB immediately! */

GC Monitoring Checklist

Metric Healthy Warning Critical
Max GC Pause < 200ms ✅ 200-500ms ⚠️ > 1000ms ❌
P99 GC Pause < 150ms ✅ 150-300ms ⚠️ > 500ms ❌
Time in GC < 5% ✅ 5-10% ⚠️ > 10% ❌
Full GC Frequency < 1/day ✅ 1/hour ⚠️ > 1/minute ❌
Heap Usage 50-70% ✅ 70-85% ⚠️ > 90% ❌

🎛️ Complete JVM Tuning Guide

Production-ready jvm.options configuration!

Recommended jvm.options Template

#################################### # CASSANDRA JVM OPTIONS - PRODUCTION # For 64GB server with G1GC #################################### # ========== HEAP SIZE ========== # Rule: 8-16GB regardless of total RAM -Xms12G -Xmx12G # ========== GC ALGORITHM ========== -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:G1RSetUpdatingPauseTimePercent=5 -XX:InitiatingHeapOccupancyPercent=70 -XX:G1ReservePercent=25 # ========== GC LOGGING ========== -Xlog:gc*:file=/var/log/cassandra/gc.log:time,uptime,level,tags:filecount=10,filesize=10M # ========== HEAP DUMP ON OOM ========== -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/log/cassandra/heapdump.hprof # ========== STRING DEDUPLICATION ========== -XX:+UseStringDeduplication # ========== PERFORMANCE ========== -XX:+AlwaysPreTouch -XX:-UseBiasedLocking # ========== LARGE PAGES (optional) ========== # -XX:+UseLargePages # ========== THREAD STACK SIZE ========== -Xss256k # ========== JMX MONITORING ========== -Dcom.sun.management.jmxremote -Dcom.sun.management.jmxremote.port=7199 -Dcom.sun.management.jmxremote.ssl=false -Dcom.sun.management.jmxremote.authenticate=true

Configuration by Server Size

Small Server (16GB RAM)

-Xms8G -Xmx8G -XX:+UseG1GC -XX:MaxGCPauseMillis=200 # Allocation: # Heap: 8GB # Caches: 1GB # OS: 7GB

Medium Server (32GB RAM) ⭐ Recommended

-Xms12G -Xmx12G -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:G1RSetUpdatingPauseTimePercent=5 # Allocation: # Heap: 12GB # Caches: 3GB (row + key) # OS: 17GB

Large Server (64GB+ RAM)

-Xms16G -Xmx16G -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:G1RSetUpdatingPauseTimePercent=5 -XX:InitiatingHeapOccupancyPercent=70 # Allocation (64GB server): # Heap: 16GB # Caches: 4GB (row + key) # OS: 44GB (huge cache!)

Extra-Large Server (128GB+ RAM)

-Xms16G -Xmx16G # Still 16GB! Don't increase heap! -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:G1RSetUpdatingPauseTimePercent=5 -XX:InitiatingHeapOccupancyPercent=70 -XX:G1ReservePercent=25 # Allocation (128GB server): # Heap: 16GB # Caches: 4GB # OS: 108GB (massive cache!)

🔧 Troubleshooting GC Issues

How to diagnose and fix JVM problems!

Common GC Problems & Solutions

Problem 1: Long GC Pauses (> 1s)

Symptoms:

  • GC pauses > 1000ms
  • Clients timeout
  • Nodes marked DOWN temporarily

Diagnosis:

nodetool gcstats // Max > 1000ms → Heap too large! grep "Full GC" /var/log/cassandra/gc.log // Frequent Full GC → Heap too large or memory leak

Solution:

# Reduce heap size -Xms12G # was 32G -Xmx12G # Ensure G1GC -XX:+UseG1GC -XX:MaxGCPauseMillis=200

Problem 2: Frequent GC (> 10% time in GC)

Symptoms:

  • GC happening constantly
  • High CPU usage
  • Degraded performance

Diagnosis:

nodetool gcstats // Total GC time > 10% of interval → Too much GC! nodetool info | grep Heap // Heap Usage: 11.8GB / 12GB → Too full!

Solution:

# Option 1: Increase heap slightly -Xms14G # was 12G -Xmx14G # Option 2: Reduce load or add nodes # Option 3: Check for memory leaks # - Large partitions? # - Too many SSTables? # - Row cache too large?

Problem 3: OutOfMemoryError

Symptoms:

  • Cassandra crashes with OOM
  • Heap dump generated
  • Node won't start

Diagnosis:

# Analyze heap dump jhat /var/log/cassandra/heapdump.hprof # Or use Eclipse MAT (better) # Look for: // - Large objects (partitions > 100MB?) // - Too many objects of same type // - Memory leaks

Solution:

  • Increase heap to 14-16GB (if < 12GB)
  • Fix data model (partition sizes)
  • Reduce row cache size
  • Check for application memory leaks

Problem 4: Unpredictable GC Pauses

Symptoms:

  • GC pauses vary from 50ms to 2000ms
  • Sometimes fast, sometimes slow
  • Using CMS collector

Solution:

# Switch from CMS to G1GC # Remove CMS flags: # -XX:+UseConcMarkSweepGC # -XX:+CMSParallelRemarkEnabled # etc... # Add G1GC: -XX:+UseG1GC -XX:MaxGCPauseMillis=200

💼 Interview Questions & Expert Answers

Master JVM tuning for your interview!

1 Why is JVM heap size limited to 8-16GB for Cassandra, even on servers with 256GB RAM? ▼

Answer: GC pause time is proportional to heap size. A 32GB+ heap causes 5-10 second stop-the-world pauses that freeze Cassandra completely, while 8-16GB keeps pauses under 200ms. Extra RAM goes to OS page cache which improves performance without GC penalties.

The Problem with Large Heaps:

  1. GC Must Scan All Objects: During GC, JVM must examine every object in heap
  2. Time Proportional to Size: 32GB heap takes 4x longer than 8GB heap
  3. Stop-the-World: ALL Cassandra threads freeze during GC
  4. User Impact: 10-second pause = 10 seconds of complete unresponsiveness
  5. Compressed Oops Lost: Heaps > 32GB lose pointer compression (20-50% overhead)

Performance Comparison:

Heap Size GC Pause User Experience
8GB 50-150ms ✅ Imperceptible
12GB 100-200ms ✅ Barely noticeable
16GB 150-300ms ⚠️ Slight delay
32GB 5-10s ❌ Complete freeze!
64GB 10-20s ❌ Catastrophic!

Why Not Use All RAM for Heap?

Cassandra benefits more from OS page cache than from large heap!

-- 256GB server comparison: // Bad: 128GB heap Heap: 128GB → GC pauses: 20+ seconds! 💥 OS Cache: 120GB // Good: 16GB heap Heap: 16GB → GC pauses: 200ms ✅ OS Cache: 235GB → Caches way more SSTables!

Key Takeaway: Cassandra's architecture is designed for large OS caches, not large heaps. Keep heap 8-16GB and let OS cache do the heavy lifting!

2 Explain the difference between G1GC and CMS. Why is G1GC recommended for Cassandra? ▼

Answer: G1GC (Garbage-First) is a concurrent, compacting collector with predictable pause times (< 200ms), while CMS (Concurrent Mark-Sweep) is a non-compacting collector prone to fragmentation and unpredictable Full GC pauses. G1GC is better for Cassandra's workload.

Key Differences:

Feature CMS G1GC
Compacting No ❌ Yes ✅
Fragmentation Severe over time Minimal
Pause Times Unpredictable Predictable (target)
Full GC Frequent (when fragmented) Rare
Heap Support < 8GB optimal 8-16GB optimal
Status Deprecated (Java 14) Default (Java 9+)

Why CMS Fails for Cassandra:

Problem 1: Fragmentation

  • CMS doesn't compact memory (no defragmentation)
  • Over days/weeks, heap becomes fragmented
  • Eventually can't allocate large objects
  • Triggers Full GC (stop-the-world)
  • Full GC can take 5-10+ seconds!

Problem 2: Concurrent Mode Failure

If heap fills up before CMS completes, it falls back to Full GC:

-- CMS log showing failure: /* [CMS-concurrent-mark: 2.157/2.443 secs] [CMS-concurrent-mode-failure: 8G->7G(12G), 8.234 secs] ↑ Fell back to Full GC! 💥 */

Why G1GC is Better:

Advantage 1: Predictable Pauses

  • Set target: -XX:MaxGCPauseMillis=200
  • G1GC tries to meet target (usually does!)
  • Pauses stay consistent: 50-200ms
  • No surprise 10-second pauses

Advantage 2: Automatic Compaction

  • G1GC compacts memory during mixed collections
  • No fragmentation buildup
  • Full GC extremely rare (< 1/day)
  • Stable performance over time

Advantage 3: Region-Based

  • Heap divided into regions (1-32MB each)
  • Collects regions with most garbage first
  • Can meet pause time targets incrementally
  • Better than CMS's generational approach

Configuration Example:

-- G1GC configuration (recommended): -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:G1RSetUpdatingPauseTimePercent=5 -XX:InitiatingHeapOccupancyPercent=70 -- Result: Consistent 50-200ms pauses ✅

Key Takeaway: G1GC is the standard for Cassandra since Java 9. CMS is deprecated and causes fragmentation issues. Always use G1GC!

3 How would you troubleshoot a Cassandra node experiencing frequent Full GC cycles? ▼

Answer: Check GC logs for Full GC frequency, analyze heap usage patterns, verify heap size is 8-16GB, ensure G1GC is enabled, check for memory leaks (large partitions, too many SSTables), and monitor if heap is consistently > 85% full.

Step-by-Step Troubleshooting:

Step 1: Check GC Statistics

-- Check current GC stats nodetool gcstats -- Output: /* Interval (ms) Max (ms) Total (ms) GC # 1000 2147 3456 12 Max: 2147ms → PROBLEM! Full GC taking 2+ seconds GC #: 12 → PROBLEM! 12 GCs in 1 second is too many! */ -- Check heap usage nodetool info | grep Heap /* Heap Memory (MB): 11534.21 / 12288.00 → 94% full! 💥 */

Step 2: Analyze GC Logs

-- Find Full GC occurrences grep "Full GC" /var/log/cassandra/gc.log | tail -20 -- Look for patterns: /* [Full GC (Allocation Failure) 11.8G->10.2G(12G), 2.147 secs] [Full GC (Allocation Failure) 11.7G->10.5G(12G), 2.234 secs] [Full GC (Allocation Failure) 11.9G->10.1G(12G), 2.089 secs] Pattern: Frequent Full GCs, only reclaiming ~1.5GB each time Problem: Heap too full, not enough garbage to collect! */

Step 3: Common Causes & Solutions

Cause 1: Heap Too Large

-- Check current heap grep -E "Xms|Xmx" /etc/cassandra/jvm.options // -Xms32G // -Xmx32G → PROBLEM! Way too large! -- Solution: Reduce heap -Xms12G -Xmx12G

Cause 2: Using CMS Instead of G1GC

-- Check GC algorithm grep "UseG1GC\|UseConcMarkSweepGC" /etc/cassandra/jvm.options -- Solution: Switch to G1GC # Remove: -XX:+UseConcMarkSweepGC # Add: -XX:+UseG1GC -XX:MaxGCPauseMillis=200

Cause 3: Memory Leak (Large Partitions)

-- Check partition sizes nodetool tablestats keyspace.table -- Output: /* Compacted partition maximum bytes: 524288000 (500MB!) Compacted partition mean bytes: 104857600 (100MB!) PROBLEM! Partitions way too large (should be < 100MB) */ -- Solution: Fix data model (add bucketing)

Cause 4: Too Many SSTables

-- Check SSTable count nodetool tablestats keyspace.table | grep "SSTable count" // SSTable count: 547 → PROBLEM! Way too many! -- Solution: Force compaction nodetool compact keyspace table

Cause 5: Row Cache Too Large

-- Check row cache nodetool info | grep "Row Cache" // Row Cache: size 8.5 GB → Too large! -- Solution: Reduce row cache # In cassandra.yaml: row_cache_size_in_mb: 2048 # was 8192

Step 4: Monitor After Changes

-- Watch GC in real-time watch -n 5 "nodetool gcstats" -- Good indicators: // Max < 200ms // GC # < 5 per second // Heap usage 50-70%

Decision Tree:

  • Heap > 16GB? → Reduce to 12GB
  • Using CMS? → Switch to G1GC
  • Heap > 85% full? → Increase slightly or reduce load
  • Large partitions? → Fix data model
  • Many SSTables? → Run compaction
  • Large row cache? → Reduce size
4 What is the significance of -Xms and -Xmx being equal? Why not let the heap grow dynamically? ▼

Answer: Setting -Xms = -Xmx prevents expensive heap resizing operations, provides predictable GC behavior, and ensures the OS doesn't reclaim committed memory. Dynamic sizing causes performance variability and additional GC overhead.

Why -Xms = -Xmx?

Advantage 1: No Heap Resizing

When -Xms < -Xmx, JVM can resize heap dynamically:

  • Growing heap: Expensive operation (can take 100s of ms)
  • Shrinking heap: Requires Full GC (stop-the-world)
  • Unpredictable pauses during resize
  • With -Xms = -Xmx: Heap size fixed, no resizing ever!

Advantage 2: Predictable Performance

-- Bad: Dynamic sizing -Xms4G -Xmx12G /* Start: 4GB heap Hour 1: Grows to 6GB (pause!) Hour 2: Grows to 8GB (pause!) Hour 3: Grows to 10GB (pause!) Low traffic: Shrinks to 6GB (Full GC!) Result: Unpredictable pauses throughout day! */ -- Good: Fixed sizing -Xms12G -Xmx12G /* Always: 12GB heap No growing, no shrinking Consistent GC behavior Predictable pauses! */

Advantage 3: OS Memory Commitment

When -Xms < -Xmx:

  • OS may not commit all memory upfront
  • Pages allocated on-demand (can be slow)
  • OS might swap unused heap pages
  • Risk of OOM if system RAM exhausted

When -Xms = -Xmx:

  • OS commits all memory at startup
  • All pages allocated upfront (use -XX:+AlwaysPreTouch)
  • No on-demand allocation delays
  • Guaranteed memory availability

Advantage 4: Faster Startup

-- With -XX:+AlwaysPreTouch: -Xms12G -Xmx12G -XX:+AlwaysPreTouch /* At startup: 1. Allocate 12GB from OS 2. Touch every page (pre-fault) 3. All memory ready immediately 4. No page faults during operation Result: Predictable startup, consistent performance */

Real-World Impact:

Configuration Behavior Production Use
-Xms4G -Xmx12G Dynamic, unpredictable ❌ Never use
-Xms12G -Xmx12G Fixed, predictable ✅ Always use

Recommended Configuration:

-- Always set equal + AlwaysPreTouch -Xms12G -Xmx12G -XX:+AlwaysPreTouch # Pre-fault all pages

Key Takeaway: -Xms = -Xmx eliminates heap resizing overhead and provides predictable, consistent performance. Always use equal values in production!

5 What JVM flags would you use for a 128GB Cassandra server and why? ▼

Answer: Still use 16GB heap (not proportional to RAM!), enable G1GC with 200ms pause target, enable GC logging, heap dump on OOM, and AlwaysPreTouch. The key is keeping heap small (16GB) regardless of total RAM, leaving 110GB+ for OS page cache.

Complete Configuration for 128GB Server:

#################################### # CASSANDRA JVM OPTIONS # Server: 128GB RAM #################################### # ========== HEAP SIZE ========== # CRITICAL: Still only 16GB! # NOT 64GB or 32GB! -Xms16G -Xmx16G # Why only 16GB on 128GB server? # - Larger heap = longer GC pauses # - 16GB is sweet spot for GC performance # - Rest (110GB) goes to OS page cache! # - OS cache benefitsALL reads # ========== GC ALGORITHM ========== -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:G1RSetUpdatingPauseTimePercent=5 -XX:InitiatingHeapOccupancyPercent=70 -XX:G1ReservePercent=25 # ========== MEMORY PRE-TOUCH ========== -XX:+AlwaysPreTouch # Pre-fault all 16GB at startup # Avoids page faults during operation # ========== GC LOGGING ========== -Xlog:gc*:file=/var/log/cassandra/gc.log:time,uptime,level,tags:filecount=10,filesize=10M # ========== HEAP DUMP ON OOM ========== -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/log/cassandra/heapdump.hprof -XX:+ExitOnOutOfMemoryError # ========== OPTIMIZATION ========== -XX:+UseStringDeduplication -XX:-UseBiasedLocking -XX:+UseTLAB # ========== THREAD STACK ========== -Xss256k # ========== JMX ========== -Dcom.sun.management.jmxremote -Dcom.sun.management.jmxremote.port=7199 -Dcom.sun.management.jmxremote.ssl=false -Dcom.sun.management.jmxremote.authenticate=true # ========== LARGE PAGES (optional) ========== # Requires OS configuration # -XX:+UseLargePages # -XX:LargePageSizeInBytes=2M

Memory Allocation Breakdown:

Total: 128GB RAM

  • JVM Heap: 16GB (12.5%)
  • Row Cache: 2-4GB (optional, 3%)
  • Key Cache: 1GB (1%)
  • OS Page Cache: 107-109GB (84%) ⭐

Why This Configuration?

Heap Size Rationale:

  • 16GB keeps GC pauses < 200ms consistently
  • 32GB would cause 5-10 second pauses!
  • 64GB would cause 20+ second pauses!
  • 16GB is optimal regardless of total RAM

OS Cache Strategy:

  • 110GB OS cache can cache HUGE portion of SSTables
  • 100GB dataset → 100% cached in RAM!
  • 500GB dataset → 22% cached (hot data)
  • OS cache has zero GC overhead
  • Benefits ALL tables, not just one

G1GC Configuration:

  • MaxGCPauseMillis=200: Target 200ms pauses
  • InitiatingHeapOccupancyPercent=70: Start concurrent GC at 70% full
  • G1ReservePercent=25: Reserve 25% to avoid evacuation failures
  • Result: Consistent, predictable GC behavior

Common Mistake to Avoid:

-- ❌ WRONG: Proportional to RAM -Xms64G # 50% of 128GB -Xmx64G /* Result: GC pauses: 10-20 seconds! 💥 Complete cluster freezes Timeouts everywhere Production disaster! */ -- ✅ CORRECT: Fixed optimal size -Xms16G # Same as 64GB server! -Xmx16G /* Result: GC pauses: 150-200ms ✅ Predictable performance 110GB for OS cache Production stable! */

Key Takeaway: More RAM doesn't mean bigger heap! Keep heap 8-16GB regardless of total RAM. Give extra RAM to OS page cache for massive performance gains!

🎓 Chapter Summary: JVM Tuning Mastery

You now understand JVM tuning at a production level!

The Golden Rules:

  • ☕ Heap Size: 8-16GB - Regardless of total RAM!
  • ♻️ Use G1GC - Predictable 50-200ms pauses
  • ⚖️ -Xms = -Xmx - No dynamic resizing
  • ❌ Never > 32GB heap - Causes 10+ second pauses
  • 💾 Maximize OS cache - Give extra RAM to OS

Quick Configuration:

-Xms12G -Xmx12G -XX:+UseG1GC -XX:MaxGCPauseMillis=200 -XX:+AlwaysPreTouch

Target Metrics:

  • ✅ Max GC pause < 200ms
  • ✅ P99 GC pause < 150ms
  • ✅ Time in GC < 5%
  • ✅ Full GC < 1/day

Remember Tom: 32GB heap = disaster, 12GB heap = success! 🚀

Advertisement

Responsive Ad