Advanced Performance Topics

Memory Management in Cassandra

Make the most of every byte! Memory management controls how Cassandra uses RAM so data stays quick to access, workloads stay stable, and performance never misses a beat.

๐Ÿ“– The Story: Emma's Memory Crisis

Emma's 128GB Cassandra cluster was supposed to be high-performance. But after 2 weeks, nodes started running out of memory. OOM errors appeared. Nodes crashed randomly. She had no idea where all the memory went. Here's what she discovered...

๐Ÿ˜ฑ The Crisis: Memory Disappeared!

Emma's Setup (128GB Server):

-- Emma's configuration JVM Heap: 16GB # Correct! Row Cache: 8GB # Seemed reasonable? Key Cache: 2GB # More is better? # Expected: 26GB used, 102GB free for OS

What Actually Happened:

  1. Week 1: Everything perfect, memory at 40GB used
  2. Week 2: Memory creeping up - 60GB, then 80GB
  3. Day 14: Memory at 95GB! Swap being used!
  4. Day 15: First OOM error - node crashed
  5. Day 16: Multiple nodes OOM - partial outage!
  6. Panic: No idea where 70GB of memory went!

The Hidden Memory Consumers:

-- Emma ran 'free -h' and found: /* total used free buffers cache Mem: 128G 120G 8G 2G 25G โ†‘ Where did 120GB go?! */ -- She discovered the culprits: JVM Heap: 16GB โœ“ Expected Row Cache: 8GB โœ“ Expected Key Cache: 2GB โœ“ Expected Memtables: 12GB โŒ SURPRISE! Bloom Filters: 8GB โŒ SURPRISE! Compression Buffers: 6GB โŒ SURPRISE! SSTable Index: 15GB โŒ SURPRISE! Network Buffers: 4GB โŒ SURPRISE! OS Overhead: 3GB โŒ SURPRISE! Native Memory: 18GB โŒ SURPRISE! โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Total: 92GB ๐Ÿ’ฅ Way over budget!

Why It Happened:

  • ๐Ÿ’ฅ Forgot Off-Heap Memory: Memtables, bloom filters, indexes all off-heap!
  • ๐Ÿ’ฅ Data Growth: More SSTables = more bloom filters + indexes
  • ๐Ÿ’ฅ High Write Rate: Large memtables consuming memory
  • ๐Ÿ’ฅ No Monitoring: Didn't track native memory usage
  • ๐Ÿ’ฅ Swap Enabled: System swapping = death spiral!

โœ… The Fix: Complete Memory Budget!

Emma's New Memory Budget (128GB):

Heap & Caches (26GB):

  • JVM Heap: 16GB
  • Row Cache: 4GB (reduced from 8GB)
  • Key Cache: 1GB (reduced from 2GB)
  • Chunk Cache: 5GB (new!)

Off-Heap Estimates (35GB):

  • Memtables: 12GB (25% of heap)
  • Bloom Filters: 8GB
  • Compression Buffers: 6GB
  • SSTable Index: 5GB
  • Network Buffers: 4GB

OS & Overhead (12GB):

  • OS Kernel: 3GB
  • Native Memory: 6GB
  • Safety Buffer: 3GB

OS Page Cache (55GB):

  • Available: 55GB
  • Caches hot SSTables automatically

Total: 26 + 35 + 12 + 55 = 128GB โœ…

Actions Taken:

  1. Disabled Swap: swapoff -a (critical!)
  2. Reduced Caches: Row cache 8GB โ†’ 4GB
  3. Tuned Memtables: Reduced flush threshold
  4. Added Monitoring: Track native memory usage
  5. Set vm.swappiness=1: Prevent kernel swapping
  6. Added Alerts: Notify if memory > 80%

The Results:

  • ๐Ÿ’š Memory Stable: 73GB used, 55GB free
  • โœ… No More OOMs: Zero crashes in 3 months
  • โšก Better Performance: 55GB OS cache improves reads
  • ๐Ÿ“Š Predictable: Memory usage flat over time
  • ๐Ÿ˜Š Peace of Mind: Can predict memory needs

Emma learned: Account for ALL memory, not just heap! ๐ŸŽ‰

๐Ÿ’พ Memory Architecture

Understanding where Cassandra uses memory!

๐ŸŽฏ Memory Categories

Cassandra uses memory in 4 main categories: (1) JVM Heap (GC-managed), (2) Off-Heap (native, no GC), (3) OS Page Cache (kernel-managed), (4) OS Overhead (kernel + buffers). Understanding each is critical for capacity planning!

Complete Memory Map

Cassandra Memory Architecture (128GB Server) JVM Heap 16GB Objects, GC-managed Memtable refs, cache Off-Heap Native 35GB Memtables, bloom filters Indexes, compression Explicit Caches 10GB Row cache, key cache Chunk cache OS Overhead 12GB Kernel, buffers, native OS Page Cache (Automatic) 55GB Caches SSTable blocks from disk Memory Allocation Breakdown JVM Heap (16GB - 12.5%) โ€ข Managed objects, references โ€ข Memtable headers, cache refs โ€ข GC overhead included Off-Heap (35GB - 27%) โ€ข Memtable data: ~12GB โ€ข Bloom filters: ~8GB โ€ข Compression, index: ~15GB Caches (10GB - 8%) โ€ข Row cache: 4GB โ€ข Key cache: 1GB โ€ข Chunk cache: 5GB OS Overhead (12GB - 9%) โ€ข Kernel structures โ€ข Network/disk buffers โ€ข Native libraries OS Page Cache (55GB - 43%) โ€ข Automatically caches frequently read SSTables โ€ข No configuration needed! โ€ข Largest performance contributor

Memory Category Details

โ˜•

JVM Heap

  • Size: 8-16GB fixed
  • Contains: Java objects, references
  • Managed: Garbage collector
  • GC Impact: Pause times increase with size
  • Typical: 12-15% of total RAM
๐Ÿ’Ž

Off-Heap Native

  • Size: 25-40% of total RAM
  • Contains: Memtables, bloom filters, indexes
  • Managed: Cassandra directly
  • GC Impact: Zero! No GC pauses
  • Growth: Scales with data + writes
๐Ÿ”ฅ

Explicit Caches

  • Size: 5-10GB configured
  • Contains: Row/key/chunk cache
  • Managed: Configured by admin
  • GC Impact: Zero (off-heap)
  • Optional: Can be disabled
โš™๏ธ

OS Overhead

  • Size: 8-12GB typical
  • Contains: Kernel, drivers, buffers
  • Managed: Operating system
  • GC Impact: N/A
  • Reserve: Always budget 10-15%
๐Ÿ’พ

OS Page Cache

  • Size: All remaining RAM
  • Contains: SSTable blocks
  • Managed: Kernel (LRU)
  • GC Impact: Zero
  • Goal: Maximize this!

๐Ÿ’Ž Off-Heap Memory Deep Dive

The hidden memory that causes OOM errors!

Critical: Off-Heap Memory is Often Underestimated!

Off-heap memory can easily consume 30-50GB on a busy node. It's not managed by GC, grows with data/writes, and is the #1 cause of unexpected OOM errors!

Off-Heap Components

1. Memtables (~25% of Heap)

In-memory write buffers before flushing to SSTables

-- Memtable memory estimation: memtable_memory = heap_size ร— 0.25 -- Example (16GB heap): 16GB ร— 0.25 = 4GB per memtable pool -- With 3 pools (typical): Total memtable memory = 12GB -- Configuration (cassandra.yaml): memtable_heap_space_in_mb: 4096 # 25% of 16GB memtable_offheap_space_in_mb: 4096

2. Bloom Filters (Scales with SSTables)

Probabilistic data structures to avoid unnecessary disk reads

-- Bloom filter memory estimation: bloom_memory = num_partitions ร— 10 bytes -- Example scenarios: 100M partitions = 1GB bloom filters 500M partitions = 5GB bloom filters 1B partitions = 10GB bloom filters -- With 100 SSTables ร— 10M partitions each: 1B total partitions = 10GB bloom filters -- Check actual usage: nodetool tablestats | grep "Bloom filter"

3. Compression Buffers

Buffers for compressing/decompressing data blocks

-- Compression buffer estimation: compression_memory = concurrent_reads ร— chunk_size ร— 2 -- Typical configuration: concurrent_reads: 32 chunk_size: 64KB buffers_per_thread: 2 -- Calculation: 32 threads ร— 64KB ร— 2 buffers = 4MB per operation 1000 concurrent operations = 4GB -- High write load can use 5-8GB!

4. SSTable Index & Summary

Partition index and sampling structures

-- Index memory estimation: index_memory = num_partitions ร— 50 bytes -- Example: 100M partitions ร— 50 bytes = 5GB 500M partitions ร— 50 bytes = 25GB! -- Grows with: // - Number of SSTables // - Number of partitions // - Index sampling rate -- Check usage: nodetool tablestats | grep "Index"

5. Network & Thread Buffers

Buffers for network I/O and thread pools

-- Network buffer estimation: native_transport_threads: 128 buffer_per_connection: 256KB -- With 1000 active connections: 1000 ร— 256KB = 256MB -- Streaming (repairs, bootstrap): streaming_buffers = stream_throughput ร— 2 200MB/s ร— 2 = 400MB -- Total network: 2-4GB typical

Off-Heap Memory Budget Calculator

-- Calculate total off-heap memory needed: # 1. Memtables heap_size = 16GB memtable_memory = heap_size ร— 0.25 ร— 3 = 12GB # 2. Bloom Filters partitions = 500M bloom_memory = partitions ร— 10 bytes = 5GB # 3. Compression Buffers compression_memory = 6GB # 4. SSTable Index index_memory = partitions ร— 50 bytes = 25GB # 5. Network Buffers network_memory = 3GB # 6. Misc Native native_overhead = 4GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Total Off-Heap: 55GB!

OOM Risk: Underestimating Off-Heap

Common Mistake:

-- Planning for 64GB server: Heap: 16GB Caches: 5GB Expected free: 43GB for OS cache -- Reality after 1 month: Heap: 16GB Caches: 5GB Off-Heap: 35GB โ† SURPRISE! OS: 6GB Total: 62GB Swap: 2GB โ† DISASTER! Result: OOM errors, node crashes! ๐Ÿ’ฅ

๐Ÿ“Š Memory Budget Planning

How to budget memory for different server sizes!

Memory Budget Templates

32GB Server (Minimum Production)

-- 32GB Total RAM JVM Heap: 8GB (25%) Row Cache: 1GB (3%) Key Cache: 512MB (1.5%) Chunk Cache: 2GB (6%) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Subtotal Heap: 11.5GB Memtables: 6GB (3 ร— 25% heap) Bloom Filters: 2GB Compression: 2GB Index/Summary: 2GB Network: 1GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Subtotal Off-Heap: 13GB OS Overhead: 3GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ OS Page Cache: 4.5GB โš ๏ธ Limited! Note: Tight! Monitor closely.

64GB Server (Recommended) โญ

-- 64GB Total RAM JVM Heap: 12GB (19%) Row Cache: 2GB (3%) Key Cache: 1GB (1.5%) Chunk Cache: 4GB (6%) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Subtotal Heap: 19GB Memtables: 9GB (3 ร— 25% heap) Bloom Filters: 4GB Compression: 4GB Index/Summary: 5GB Network: 2GB Native: 3GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Subtotal Off-Heap: 27GB OS Overhead: 4GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ OS Page Cache: 14GB โœ… Good! Note: Balanced, good for most workloads

128GB Server (High Performance)

-- 128GB Total RAM JVM Heap: 16GB (12.5%) Row Cache: 4GB (3%) Key Cache: 1GB (1%) Chunk Cache: 5GB (4%) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Subtotal Heap: 26GB Memtables: 12GB (3 ร— 25% heap) Bloom Filters: 8GB Compression: 6GB Index/Summary: 10GB Network: 4GB Native: 5GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Subtotal Off-Heap: 45GB OS Overhead: 6GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ OS Page Cache: 51GB โญ Excellent! Note: Ideal for read-heavy workloads

256GB Server (Premium)

-- 256GB Total RAM JVM Heap: 16GB (6%) Row Cache: 8GB (3%) Key Cache: 2GB (1%) Chunk Cache: 10GB (4%) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Subtotal Heap: 36GB Memtables: 12GB (3 ร— 25% heap) Bloom Filters: 15GB Compression: 10GB Index/Summary: 20GB Network: 6GB Native: 8GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Subtotal Off-Heap: 71GB OS Overhead: 10GB โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ OS Page Cache: 139GB ๐Ÿš€ Massive! Note: Can cache 100GB+ datasets entirely!

Memory Budget Validation

-- Check actual memory usage (Linux): # 1. Overall memory free -h /* total used free buffers cache Mem: 128G 73G 55G 2G 25G */ # 2. JVM heap usage nodetool info | grep Heap // Heap Memory (MB): 11245 / 12288 (91% - OK) # 3. Off-heap (estimate) pmap -x | tail -1 // Look at "total" RSS # 4. Detailed breakdown nodetool tablestats // Shows bloom filter size, index size # 5. Check for swap usage swapon -s // Should be EMPTY! Swap = disaster!

Quick Memory Budget Formula

Conservative Estimate:

Heap + Caches: 15-30% of RAM Off-Heap: 25-40% of RAM OS Overhead: 10-15% of RAM โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ OS Page Cache: 30-50% of RAM (remaining) Rule of Thumb: If OS cache < 30%, need more RAM!

๐Ÿ“ˆ Memory Monitoring

Essential commands to track memory usage!

Key Metrics to Monitor

-- 1. System-wide memory free -h /* Watch for: - "used" approaching "total" (> 90% = problem!) - Any swap usage (swap = death spiral!) - "cache" shrinking (means other memory growing) */ -- 2. JVM heap usage nodetool info | grep "Heap Memory" // Target: 60-80% utilization // > 90% = frequent GC // < 40% = heap too large -- 3. Cache statistics nodetool info | grep Cache /* Row Cache: entries X, size Y GB, capacity Z GB Key Cache: entries X, size Y GB, capacity Z GB Check: size approaching capacity? */ -- 4. Table statistics nodetool tablestats keyspace.table /* Look for: - Bloom filter space used - Index summary size - Compression metadata */ -- 5. Process memory map pmap -x $(pgrep -f cassandra) | tail -20 // Shows detailed memory breakdown -- 6. Swap usage (CRITICAL!) vmstat 1 // si/so columns (swap in/out) // ANY activity = PROBLEM!

Monitoring Checklist

Metric Healthy Warning Critical
Total RAM Used < 80% โœ… 80-90% โš ๏ธ > 90% โŒ
Heap Usage 60-80% โœ… 80-90% โš ๏ธ > 90% โŒ
OS Page Cache > 30% RAM โœ… 20-30% โš ๏ธ < 20% โŒ
Swap Usage 0 bytes โœ… Any usage โš ๏ธ Active I/O โŒ
Cache Hit Rate > 90% โœ… 70-90% โš ๏ธ < 70% โŒ

Alert Thresholds

-- Recommended alerts: # WARNING alerts: - Total RAM usage > 80% - Heap usage > 85% - OS page cache < 30% of RAM - Any swap usage detected # CRITICAL alerts: - Total RAM usage > 90% - Heap usage > 95% - Active swap I/O (si/so > 0) - OOM errors in logs # Example monitoring script: #!/bin/bash MEM_PCT=$(free | grep Mem | awk '{print ($3/$2) * 100.0}') if (( $(echo "$MEM_PCT > 80" | bc -l) )); then echo "WARNING: Memory at ${MEM_PCT}%" # Send alert fi

๐Ÿ”ง Memory Troubleshooting

How to diagnose and fix memory issues!

Common Memory Problems

Problem 1: Out of Memory (OOM) Errors

Symptoms:

  • Node crashes with OOM
  • Memory grows over time
  • Logs show "java.lang.OutOfMemoryError"

Diagnosis:

-- Check system memory free -h // Total used > 95%? -- Check heap usage nodetool info | grep Heap // Heap at 100%? -- Check for memory leak indicators grep -i "OutOfMemoryError" /var/log/cassandra/system.log -- Analyze heap dump (if generated) jhat /var/log/cassandra/heapdump.hprof

Solutions:

  • โœ… Increase server RAM (if undersized)
  • โœ… Reduce row cache if too large
  • โœ… Fix large partitions (data model issue)
  • โœ… Reduce write load or add nodes
  • โœ… Force compaction to reduce bloom filters

Problem 2: System Using Swap

Symptoms:

  • Extremely slow performance
  • High disk I/O wait times
  • Swap space being used

Diagnosis:

-- Check swap status swapon -s // ANY output = PROBLEM! -- Check swap activity vmstat 1 10 // si/so columns > 0 = swapping actively! free -h // Look at "Swap" row

Solutions:

# IMMEDIATE: Disable swap sudo swapoff -a # PERMANENT: Edit /etc/fstab # Comment out swap line # Set swappiness to 1 (emergency only) sudo sysctl vm.swappiness=1 echo "vm.swappiness=1" | sudo tee -a /etc/sysctl.conf # Restart Cassandra sudo systemctl restart cassandra

Problem 3: Memory Leak (Growing Over Time)

Symptoms:

  • Memory usage grows steadily
  • Eventually leads to OOM
  • Restart temporarily fixes

Common Causes:

  • ๐Ÿ” Too many SSTables: Bloom filters + indexes growing
  • ๐Ÿ” Large partitions: Memtables can't flush efficiently
  • ๐Ÿ” Data growth: Off-heap structures scale with data
  • ๐Ÿ” Bug: Actual memory leak in Cassandra or driver

Solutions:

-- Check SSTable count nodetool tablestats | grep "SSTable count" // > 100 per table = problem! -- Force compaction nodetool compact keyspace table -- Check partition sizes nodetool tablestats | grep "partition" // > 100MB = fix data model! -- Monitor bloom filter growth nodetool tablestats | grep "Bloom filter"

Problem 4: OS Page Cache Too Small

Symptoms:

  • High read latency
  • Excessive disk I/O
  • OS cache < 30% of RAM

Diagnosis:

free -h /* total used free cache Mem: 64G 60G 4G 8G โ†‘ Too small! */ -- Target: cache > 30% of total RAM // 64GB ร— 0.30 = 19GB minimum

Solutions:

  • โœ… Reduce JVM heap (if > 16GB)
  • โœ… Reduce row cache size
  • โœ… Add more RAM to server
  • โœ… Reduce off-heap usage (tune memtables)

๐Ÿ’ผ Interview Questions & Expert Answers

Master memory management for your interview!

1 Explain the difference between heap and off-heap memory in Cassandra. Why is off-heap memory often underestimated? โ–ผ

Answer: Heap memory is GC-managed Java objects (8-16GB fixed), while off-heap is native memory outside JVM (memtables, bloom filters, indexes - can be 30-50GB). Off-heap is underestimated because it's invisible to 'nodetool info', grows with data/writes, and is the #1 cause of OOM errors.

Heap Memory:

  • Size: 8-16GB (fixed, configured)
  • Contains: Java objects, references, metadata
  • Management: Garbage collector
  • Visibility: 'nodetool info' shows it
  • Growth: Fixed size, doesn't grow
  • Performance: GC pauses affect it

Off-Heap Memory:

  • Size: 25-50GB (dynamic, grows!)
  • Contains: Memtables, bloom filters, indexes, compression buffers, network buffers
  • Management: Cassandra directly (no GC)
  • Visibility: Not shown in nodetool info!
  • Growth: Scales with data, writes, SSTables
  • Performance: Zero GC overhead

Why Off-Heap is Underestimated:

-- Planning scenario (64GB server): // Visible memory (what people plan for): Heap: 16GB โ† nodetool info shows this Row cache: 2GB โ† nodetool info shows this Key cache: 1GB โ† nodetool info shows this Expected: 19GB // Reality (what actually happens): Heap: 16GB Caches: 3GB Memtables: 12GB โ† Hidden! Bloom filters: 8GB โ† Hidden! Compression: 4GB โ† Hidden! Index: 6GB โ† Hidden! Network: 2GB โ† Hidden! Native: 4GB โ† Hidden! Actual: 55GB! Result: OOM crash! ๐Ÿ’ฅ

Key Takeaway: Always budget for 25-40% of RAM as off-heap memory. It's not optional - it WILL be used!

2 How would you size a Cassandra server? Walk through the complete memory budget for a 128GB server. โ–ผ

Answer: For 128GB server: 16GB heap + 10GB caches + 35-45GB off-heap + 12GB OS overhead = 73-83GB used, leaving 45-55GB for OS page cache. Always budget all categories including hidden off-heap memory.

Complete 128GB Memory Budget:

======================================== CATEGORY 1: JVM HEAP (16GB - 12.5%) ======================================== JVM Heap: 16GB Why 16GB not 32GB or 64GB? - Keeps GC pauses < 200ms - 32GB+ causes 5-10 second pauses! - Sweet spot regardless of total RAM ======================================== CATEGORY 2: EXPLICIT CACHES (10GB - 8%) ======================================== Row Cache: 4GB # Hot data Key Cache: 1GB # Partition locations Chunk Cache: 5GB # Compressed blocks ======================================== CATEGORY 3: OFF-HEAP (40GB - 31%) ======================================== Memtables: 12GB # 3 ร— 25% heap Bloom Filters: 8GB # ~10 bytes ร— partitions Compression Buffers: 6GB # Concurrent operations SSTable Index/Summary: 10GB # ~50 bytes ร— partitions Network Buffers: 3GB # Client connections Native Memory: 1GB # Misc native libs ======================================== CATEGORY 4: OS OVERHEAD (12GB - 9%) ======================================== OS Kernel: 3GB Disk/Network Buffers: 2GB System Libraries: 3GB Safety Buffer: 4GB ======================================== CATEGORY 5: OS PAGE CACHE (50GB - 39%) ======================================== Available for Caching: 50GB This is THE MOST IMPORTANT part! - Automatically caches hot SSTables - Zero configuration needed - Benefits ALL reads - Maximize this by keeping heap small! ======================================== TOTAL: 128GB ======================================== 16 + 10 + 40 + 12 + 50 = 128GB โœ…

Validation Steps:

  1. Check Total: Sum should equal server RAM
  2. OS Cache Check: Should be 30-50% of RAM minimum
  3. Off-Heap Budget: 25-40% is realistic estimate
  4. Monitor After Deploy: Watch actual usage vs budget

Key Principle: If OS page cache < 30% of RAM, you need more RAM or must reduce heap/caches!

3 Why should swap always be disabled on Cassandra nodes? What happens if swap is enabled? โ–ผ

Answer: Swap causes disk I/O for memory access (1000x slower), creating a death spiral where slow operations cause more memory pressure causing more swapping causing more slowness. Result: 10+ second queries, timeouts, node appears down. Always disable swap completely.

What Happens With Swap Enabled:

-- The Death Spiral: 1. Memory fills up (> 95%) 2. Kernel swaps pages to disk 3. Memory access now requires disk I/O RAM: 0.1ยตs access time Disk: 10ms access time โ†‘ 100,000x slower! 4. Queries slow down dramatically Normal: 5ms Swapping: 10,000ms (10 seconds!) 5. Slow queries hold resources longer โ†’ More threads blocked โ†’ More memory pressure โ†’ MORE swapping! 6. Complete system collapse โ†’ All queries timeout โ†’ Clients mark node DOWN โ†’ Other nodes overwhelmed โ†’ Cascading failure! ๐Ÿ’ฅ

Real-World Example:

-- Before swap disabled: Query latency P99: 12s โ† Disaster! Disk I/O wait: 85% Swap I/O: 500MB/s โ† Death spiral! Client timeouts: 95% -- After 'swapoff -a': Query latency P99: 15ms โœ… Disk I/O wait: 15% Swap I/O: 0MB/s โœ… Client timeouts: 0% โœ…

How to Disable Swap:

# 1. Immediate (temporary) sudo swapoff -a # 2. Permanent (edit /etc/fstab) sudo vi /etc/fstab # Comment out swap line: # /dev/sda2 swap swap defaults 0 0 # 3. Set vm.swappiness (emergency-only swap) sudo sysctl vm.swappiness=1 echo "vm.swappiness=1" | sudo tee -a /etc/sysctl.conf # 4. Verify swapon -s # Should show NOTHING free -h /* Swap: 0B 0B 0B โœ… Perfect! */

Why vm.swappiness=1 (not 0)?

  • swappiness=0: Kernel NEVER swaps (can cause OOM kills)
  • swappiness=1: Only swap in absolute emergency
  • Better to have emergency valve than sudden OOM kill
  • But really: If swapping, need more RAM!

Key Takeaway: Swap = death sentence for Cassandra. Always disable completely. If node needs swap, it needs more RAM!

4 A Cassandra node is showing 95% memory usage but nodetool info shows heap at only 70%. Where is the memory going? โ–ผ

Answer: Memory is in off-heap structures invisible to nodetool info: memtables (~12GB), bloom filters (scales with partitions), compression buffers, SSTable indexes, and network buffers. Check with pmap, tablestats, and system monitoring tools.

Investigation Steps:

Step 1: Check System Memory

free -h /* total used free cache Mem: 64G 61G 3G 8G 61GB used but heap only shows 11GB โ†’ 50GB unaccounted! */

Step 2: Check Heap Usage

nodetool info | grep Heap // Heap Memory (MB): 11245 / 16384 (69%) Only accounts for 11GB of 61GB used!

Step 3: Check Caches

nodetool info | grep Cache /* Row Cache: size 2.1 GB Key Cache: size 0.9 GB */ Caches account for 3GB. Running total: 11 + 3 = 14GB of 61GB

Step 4: Check Table Statistics

nodetool tablestats /* Bloom filter space used: 8.5 GB โ† Found 8.5GB! Index summary size: 6.2 GB โ† Found 6.2GB! Compression metadata: 2.1 GB โ† Found 2.1GB! */ Running total: 14 + 8.5 + 6.2 + 2.1 = 30.8GB Still missing: 61 - 30.8 = 30.2GB!

Step 5: Estimate Memtables

-- Memtable memory (not shown directly): Memtable pools: 3 Size per pool: 25% of heap = 4GB Total memtables: 12GB โ† Found 12GB! Running total: 30.8 + 12 = 42.8GB Still missing: 61 - 42.8 = 18.2GB

Step 6: Check Process Memory Map

pmap -x $(pgrep -f cassandra) | tail -20 /* Shows detailed memory breakdown: - Native libraries: 3GB - Thread stacks: 2GB - Network buffers: 4GB - Compression buffers: 5GB - Misc native: 4GB Total: ~18GB โ† Found the rest! */

Complete Breakdown:

-- 61GB total used explained: Heap: 11GB โœ“ nodetool info Row + Key Cache: 3GB โœ“ nodetool info Bloom Filters: 8.5GB โœ“ tablestats Index Summary: 6.2GB โœ“ tablestats Compression Meta: 2.1GB โœ“ tablestats Memtables: 12GB โœ“ estimated Native/Buffers: 18GB โœ“ pmap โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Total: 60.8GB โ‰ˆ 61GB โœ…

Key Takeaway: Most memory is off-heap and invisible to nodetool info. Always use system tools (free, pmap) + tablestats to see complete picture!

5 What are the warning signs that a Cassandra cluster is running out of memory? How would you prevent it? โ–ผ

Answer: Warning signs: memory usage > 80%, OS page cache < 30%, heap > 85%, any swap usage, GC pauses increasing, slow queries. Prevention: proper capacity planning with complete memory budget, monitoring all memory types, disabling swap, and adding RAM before reaching 80% usage.

Early Warning Signs:

Warning Sign Threshold Action
Total RAM usage > 80% Plan for more RAM/nodes
OS page cache < 30% of RAM Reduce heap/caches
Heap usage > 85% Investigate memory leak
Swap usage Any usage CRITICAL - disable immediately
GC pause times Increasing trend Heap pressure building
Query latency P99 Increasing trend Cache eviction, disk I/O up

Prevention Strategy:

1. Proper Capacity Planning

-- Create complete memory budget BEFORE deploying: Heap: 16GB (12.5%) Caches: 10GB (8%) Off-Heap Est: 40GB (31%) OS Overhead: 12GB (9%) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Used: 78GB OS Page Cache: 50GB (39%) โœ… Rule: If OS cache < 30%, need more RAM!

2. Comprehensive Monitoring

# Monitor ALL memory categories: - Total system memory (free -h) - Heap usage (nodetool info) - Cache sizes (nodetool info) - Off-heap estimates (tablestats) - OS page cache size - Swap usage (MUST be zero!) - GC pause trends - Query latency P99

3. Proactive Actions

  • โœ… Set alerts at 80% memory usage
  • โœ… Disable swap completely
  • โœ… Review memory budget monthly
  • โœ… Plan capacity 6 months ahead
  • โœ… Add RAM/nodes before 85% usage
  • โœ… Test memory limits in staging first

4. Emergency Response Plan

# If memory reaches 90%: 1. Check for swap - disable immediately sudo swapoff -a 2. Reduce row cache if large nodetool setcachecapacity 0 0 3. Force compaction (reduce bloom filters) nodetool compact 4. Emergency: Restart node (clears memtables) 5. Long-term: Add RAM or scale out

Key Takeaway: Don't wait for OOM! Monitor continuously, act at 80% usage, and always have complete memory budget including off-heap estimates.

๐ŸŽ“ Chapter Summary: Memory Management Mastery

You now understand Cassandra memory management at a production level!

The Five Memory Categories:

  • โ˜• JVM Heap: 8-16GB (12-15% of RAM)
  • ๐Ÿ’Ž Off-Heap: 25-40% (memtables, bloom, indexes)
  • ๐Ÿ”ฅ Caches: 5-10% (row, key, chunk)
  • โš™๏ธ OS Overhead: 10-15% (kernel, buffers)
  • ๐Ÿ’พ OS Page Cache: 30-50% (maximize this!)

Critical Rules:

  • โŒ Disable swap always - death spiral guaranteed
  • ๐Ÿ“Š Budget ALL memory - off-heap is 25-40%!
  • โš ๏ธ Alert at 80% RAM - act before 90%
  • ๐ŸŽฏ OS cache > 30% - or need more RAM

128GB Server Example:

16GB heap + 10GB caches + 40GB off-heap + 12GB OS = 78GB used, 50GB for OS cache โœ…

Remember Emma: Account for ALL memory, not just heap! ๐Ÿš€

Advertisement

Responsive Ad