Memory Management in Cassandra
Make the most of every byte! Memory management controls how Cassandra uses RAM so data stays quick to access, workloads stay stable, and performance never misses a beat.
๐ The Story: Emma's Memory Crisis
Emma's 128GB Cassandra cluster was supposed to be high-performance. But after 2 weeks, nodes started running out of memory. OOM errors appeared. Nodes crashed randomly. She had no idea where all the memory went. Here's what she discovered...
๐ฑ The Crisis: Memory Disappeared!
Emma's Setup (128GB Server):
What Actually Happened:
- Week 1: Everything perfect, memory at 40GB used
- Week 2: Memory creeping up - 60GB, then 80GB
- Day 14: Memory at 95GB! Swap being used!
- Day 15: First OOM error - node crashed
- Day 16: Multiple nodes OOM - partial outage!
- Panic: No idea where 70GB of memory went!
The Hidden Memory Consumers:
Why It Happened:
- ๐ฅ Forgot Off-Heap Memory: Memtables, bloom filters, indexes all off-heap!
- ๐ฅ Data Growth: More SSTables = more bloom filters + indexes
- ๐ฅ High Write Rate: Large memtables consuming memory
- ๐ฅ No Monitoring: Didn't track native memory usage
- ๐ฅ Swap Enabled: System swapping = death spiral!
โ The Fix: Complete Memory Budget!
Emma's New Memory Budget (128GB):
Heap & Caches (26GB):
- JVM Heap: 16GB
- Row Cache: 4GB (reduced from 8GB)
- Key Cache: 1GB (reduced from 2GB)
- Chunk Cache: 5GB (new!)
Off-Heap Estimates (35GB):
- Memtables: 12GB (25% of heap)
- Bloom Filters: 8GB
- Compression Buffers: 6GB
- SSTable Index: 5GB
- Network Buffers: 4GB
OS & Overhead (12GB):
- OS Kernel: 3GB
- Native Memory: 6GB
- Safety Buffer: 3GB
OS Page Cache (55GB):
- Available: 55GB
- Caches hot SSTables automatically
Total: 26 + 35 + 12 + 55 = 128GB โ
Actions Taken:
- Disabled Swap: swapoff -a (critical!)
- Reduced Caches: Row cache 8GB โ 4GB
- Tuned Memtables: Reduced flush threshold
- Added Monitoring: Track native memory usage
- Set vm.swappiness=1: Prevent kernel swapping
- Added Alerts: Notify if memory > 80%
The Results:
- ๐ Memory Stable: 73GB used, 55GB free
- โ No More OOMs: Zero crashes in 3 months
- โก Better Performance: 55GB OS cache improves reads
- ๐ Predictable: Memory usage flat over time
- ๐ Peace of Mind: Can predict memory needs
Emma learned: Account for ALL memory, not just heap! ๐
๐พ Memory Architecture
Understanding where Cassandra uses memory!
๐ฏ Memory Categories
Cassandra uses memory in 4 main categories: (1) JVM Heap (GC-managed), (2) Off-Heap (native, no GC), (3) OS Page Cache (kernel-managed), (4) OS Overhead (kernel + buffers). Understanding each is critical for capacity planning!
Complete Memory Map
Memory Category Details
JVM Heap
- Size: 8-16GB fixed
- Contains: Java objects, references
- Managed: Garbage collector
- GC Impact: Pause times increase with size
- Typical: 12-15% of total RAM
Off-Heap Native
- Size: 25-40% of total RAM
- Contains: Memtables, bloom filters, indexes
- Managed: Cassandra directly
- GC Impact: Zero! No GC pauses
- Growth: Scales with data + writes
Explicit Caches
- Size: 5-10GB configured
- Contains: Row/key/chunk cache
- Managed: Configured by admin
- GC Impact: Zero (off-heap)
- Optional: Can be disabled
OS Overhead
- Size: 8-12GB typical
- Contains: Kernel, drivers, buffers
- Managed: Operating system
- GC Impact: N/A
- Reserve: Always budget 10-15%
OS Page Cache
- Size: All remaining RAM
- Contains: SSTable blocks
- Managed: Kernel (LRU)
- GC Impact: Zero
- Goal: Maximize this!
๐ Off-Heap Memory Deep Dive
The hidden memory that causes OOM errors!
Critical: Off-Heap Memory is Often Underestimated!
Off-heap memory can easily consume 30-50GB on a busy node. It's not managed by GC, grows with data/writes, and is the #1 cause of unexpected OOM errors!
Off-Heap Components
1. Memtables (~25% of Heap)
In-memory write buffers before flushing to SSTables
2. Bloom Filters (Scales with SSTables)
Probabilistic data structures to avoid unnecessary disk reads
3. Compression Buffers
Buffers for compressing/decompressing data blocks
4. SSTable Index & Summary
Partition index and sampling structures
5. Network & Thread Buffers
Buffers for network I/O and thread pools
Off-Heap Memory Budget Calculator
OOM Risk: Underestimating Off-Heap
Common Mistake:
๐ Memory Budget Planning
How to budget memory for different server sizes!
Memory Budget Templates
32GB Server (Minimum Production)
64GB Server (Recommended) โญ
128GB Server (High Performance)
256GB Server (Premium)
Memory Budget Validation
Quick Memory Budget Formula
Conservative Estimate:
๐ Memory Monitoring
Essential commands to track memory usage!
Key Metrics to Monitor
Monitoring Checklist
| Metric | Healthy | Warning | Critical |
|---|---|---|---|
| Total RAM Used | < 80% โ | 80-90% โ ๏ธ | > 90% โ |
| Heap Usage | 60-80% โ | 80-90% โ ๏ธ | > 90% โ |
| OS Page Cache | > 30% RAM โ | 20-30% โ ๏ธ | < 20% โ |
| Swap Usage | 0 bytes โ | Any usage โ ๏ธ | Active I/O โ |
| Cache Hit Rate | > 90% โ | 70-90% โ ๏ธ | < 70% โ |
Alert Thresholds
๐ง Memory Troubleshooting
How to diagnose and fix memory issues!
Common Memory Problems
Problem 1: Out of Memory (OOM) Errors
Symptoms:
- Node crashes with OOM
- Memory grows over time
- Logs show "java.lang.OutOfMemoryError"
Diagnosis:
Solutions:
- โ Increase server RAM (if undersized)
- โ Reduce row cache if too large
- โ Fix large partitions (data model issue)
- โ Reduce write load or add nodes
- โ Force compaction to reduce bloom filters
Problem 2: System Using Swap
Symptoms:
- Extremely slow performance
- High disk I/O wait times
- Swap space being used
Diagnosis:
Solutions:
Problem 3: Memory Leak (Growing Over Time)
Symptoms:
- Memory usage grows steadily
- Eventually leads to OOM
- Restart temporarily fixes
Common Causes:
- ๐ Too many SSTables: Bloom filters + indexes growing
- ๐ Large partitions: Memtables can't flush efficiently
- ๐ Data growth: Off-heap structures scale with data
- ๐ Bug: Actual memory leak in Cassandra or driver
Solutions:
Problem 4: OS Page Cache Too Small
Symptoms:
- High read latency
- Excessive disk I/O
- OS cache < 30% of RAM
Diagnosis:
Solutions:
- โ Reduce JVM heap (if > 16GB)
- โ Reduce row cache size
- โ Add more RAM to server
- โ Reduce off-heap usage (tune memtables)
๐ผ Interview Questions & Expert Answers
Master memory management for your interview!
Answer: Heap memory is GC-managed Java objects (8-16GB fixed), while off-heap is native memory outside JVM (memtables, bloom filters, indexes - can be 30-50GB). Off-heap is underestimated because it's invisible to 'nodetool info', grows with data/writes, and is the #1 cause of OOM errors.
Heap Memory:
- Size: 8-16GB (fixed, configured)
- Contains: Java objects, references, metadata
- Management: Garbage collector
- Visibility: 'nodetool info' shows it
- Growth: Fixed size, doesn't grow
- Performance: GC pauses affect it
Off-Heap Memory:
- Size: 25-50GB (dynamic, grows!)
- Contains: Memtables, bloom filters, indexes, compression buffers, network buffers
- Management: Cassandra directly (no GC)
- Visibility: Not shown in nodetool info!
- Growth: Scales with data, writes, SSTables
- Performance: Zero GC overhead
Why Off-Heap is Underestimated:
Key Takeaway: Always budget for 25-40% of RAM as off-heap memory. It's not optional - it WILL be used!
Answer: For 128GB server: 16GB heap + 10GB caches + 35-45GB off-heap + 12GB OS overhead = 73-83GB used, leaving 45-55GB for OS page cache. Always budget all categories including hidden off-heap memory.
Complete 128GB Memory Budget:
Validation Steps:
- Check Total: Sum should equal server RAM
- OS Cache Check: Should be 30-50% of RAM minimum
- Off-Heap Budget: 25-40% is realistic estimate
- Monitor After Deploy: Watch actual usage vs budget
Key Principle: If OS page cache < 30% of RAM, you need more RAM or must reduce heap/caches!
Answer: Swap causes disk I/O for memory access (1000x slower), creating a death spiral where slow operations cause more memory pressure causing more swapping causing more slowness. Result: 10+ second queries, timeouts, node appears down. Always disable swap completely.
What Happens With Swap Enabled:
Real-World Example:
How to Disable Swap:
Why vm.swappiness=1 (not 0)?
- swappiness=0: Kernel NEVER swaps (can cause OOM kills)
- swappiness=1: Only swap in absolute emergency
- Better to have emergency valve than sudden OOM kill
- But really: If swapping, need more RAM!
Key Takeaway: Swap = death sentence for Cassandra. Always disable completely. If node needs swap, it needs more RAM!
Answer: Memory is in off-heap structures invisible to nodetool info: memtables (~12GB), bloom filters (scales with partitions), compression buffers, SSTable indexes, and network buffers. Check with pmap, tablestats, and system monitoring tools.
Investigation Steps:
Step 1: Check System Memory
Step 2: Check Heap Usage
Step 3: Check Caches
Step 4: Check Table Statistics
Step 5: Estimate Memtables
Step 6: Check Process Memory Map
Complete Breakdown:
Key Takeaway: Most memory is off-heap and invisible to nodetool info. Always use system tools (free, pmap) + tablestats to see complete picture!
Answer: Warning signs: memory usage > 80%, OS page cache < 30%, heap > 85%, any swap usage, GC pauses increasing, slow queries. Prevention: proper capacity planning with complete memory budget, monitoring all memory types, disabling swap, and adding RAM before reaching 80% usage.
Early Warning Signs:
| Warning Sign | Threshold | Action |
|---|---|---|
| Total RAM usage | > 80% | Plan for more RAM/nodes |
| OS page cache | < 30% of RAM | Reduce heap/caches |
| Heap usage | > 85% | Investigate memory leak |
| Swap usage | Any usage | CRITICAL - disable immediately |
| GC pause times | Increasing trend | Heap pressure building |
| Query latency P99 | Increasing trend | Cache eviction, disk I/O up |
Prevention Strategy:
1. Proper Capacity Planning
2. Comprehensive Monitoring
3. Proactive Actions
- โ Set alerts at 80% memory usage
- โ Disable swap completely
- โ Review memory budget monthly
- โ Plan capacity 6 months ahead
- โ Add RAM/nodes before 85% usage
- โ Test memory limits in staging first
4. Emergency Response Plan
Key Takeaway: Don't wait for OOM! Monitor continuously, act at 80% usage, and always have complete memory budget including off-heap estimates.
๐ Chapter Summary: Memory Management Mastery
You now understand Cassandra memory management at a production level!
The Five Memory Categories:
- โ JVM Heap: 8-16GB (12-15% of RAM)
- ๐ Off-Heap: 25-40% (memtables, bloom, indexes)
- ๐ฅ Caches: 5-10% (row, key, chunk)
- โ๏ธ OS Overhead: 10-15% (kernel, buffers)
- ๐พ OS Page Cache: 30-50% (maximize this!)
Critical Rules:
- โ Disable swap always - death spiral guaranteed
- ๐ Budget ALL memory - off-heap is 25-40%!
- โ ๏ธ Alert at 80% RAM - act before 90%
- ๐ฏ OS cache > 30% - or need more RAM
128GB Server Example:
16GB heap + 10GB caches + 40GB off-heap + 12GB OS = 78GB used, 50GB for OS cache โ
Remember Emma: Account for ALL memory, not just heap! ๐
Responsive Ad