ROW CACHE - Lightning-Fast Memory
💾 Master Cassandra's row cache with interactive examples, animations, and real-world scenarios. Learn caching from absolute basics!
📖 Fundamentals (Start Here If You're Brand New!)
Before we dive into row cache, let's make sure we understand ALL the basic terms. No prior knowledge needed!
What is Memory (RAM)?
RAM = Random Access Memory
Think of it as: Your computer's short-term memory
Like your brain:
• Remembers things RIGHT NOW
• Super fast to access
• Forgets when power off
• Limited space
Example: 16GB RAM means 16 gigabytes of "working memory"
Speed: 0.0001 milliseconds (0.1 microseconds)
Nickname: "Memory" = RAM (same thing!)
What is a Disk?
Disk = Storage Drive
Think of it as: Your computer's long-term memory
Two types:
• HDD: Spinning disk (like a CD player)
• SSD: Solid state (like a USB stick)
Properties:
• Remembers forever (even when power off)
• SLOW to access (100,000x slower than RAM!)
• Huge capacity (1TB = 1000GB typical)
Why slow: Physical parts must move to find data
Time Units Explained
Millisecond (ms) = 1/1000 of a second
• 1 second = 1,000 milliseconds
• Blink of an eye = 300-400ms
Examples:
• 1ms = Very fast (you can't feel it)
• 10ms = Fast (barely noticeable)
• 100ms = Slow (you notice the delay)
• 1000ms = 1 second (very slow for computer)
Why it matters: Users notice anything over 100ms as "lag"
Microsecond (μs) = 1/1,000,000 of a second (even smaller!)
Size Units (GB, MB, KB)
Byte = Smallest unit (1 letter = 1 byte)
Kilobyte (KB) = 1,000 bytes
• Text message: ~1KB
• Short email: ~5KB
Megabyte (MB) = 1,000 KB
• Photo: ~2-5MB
• Song: ~3-5MB
Gigabyte (GB) = 1,000 MB
• Movie: ~2-4GB
• Typical RAM: 8-32GB
1 GB = 1 billion bytes!
What is JVM and Heap?
JVM = Java Virtual Machine
• Software that runs Java programs
• Cassandra is written in Java
• Think: JVM = the engine running Cassandra
Heap = Memory pool for JVM
• Part of RAM set aside for Java
• Where Java stores objects
• Example: "8GB heap" = 8GB of RAM for Java
Why it matters: Row cache lives IN the heap!
What is Garbage Collection (GC)?
GC = Automatic memory cleanup
Like a Roomba vacuum: Runs automatically to clean up
What it does:
• Java creates objects in heap
• Old unused objects pile up
• GC finds and deletes unused objects
• Frees up space
THE PROBLEM: While GC runs, application PAUSES!
• Small heap (8GB): Pause 100ms
• Large heap (32GB): Pause 2-5 seconds!
Why it matters: Larger row cache = more GC pauses!
Latency vs Throughput
Latency = How long ONE thing takes
• "This query took 5 milliseconds"
• Lower is better!
• Like: How long to drive to store
Throughput = How many things per second
• "We handle 10,000 queries per second"
• Higher is better!
• Like: How many cars can use highway per hour
Both matter: Fast individual queries + handle many queries
Percentiles (P50, P99)
Simple explanation: Performance for different users
Imagine 100 people query:
• Sort their speeds: fastest → slowest
• P50 (median): 50th person's speed
• P99: 99th person's speed (only 1 slower)
Example:
• P50 = 2ms: Typical user experience
• P99 = 20ms: Worst case for 1% of users
Why P99 matters: Don't want even unlucky users to suffer!
What is a Node/Server?
Server = A computer
• Physical machine in data center
• Or virtual machine in cloud
• Runs 24/7
Node = Server running Cassandra
• Same thing in Cassandra world
• One Cassandra node = one server
Example: "50 nodes" = 50 servers running Cassandra
Typical server: 32GB RAM, 1TB disk, 16 CPU cores
What is a Table and Query?
Table = Spreadsheet of data
• Rows and columns
• Example: "users" table
• Columns: id, name, email
Query = Request for data
• Written in CQL (Cassandra Query Language)
• Example: "Get user with id=123"
• Like asking database a question
SQL-like: SELECT * FROM users WHERE id='123'
What is Disk I/O?
I/O = Input/Output
Disk I/O = Reading from or writing to disk
Input (Read):
• Get data FROM disk TO memory
• Example: Load user profile from disk
Output (Write):
• Save data FROM memory TO disk
• Example: Store new user profile
Why we care: Disk I/O is SLOW (30ms+)
Goal: Minimize disk I/O with caching!
LRU (Eviction Policy)
LRU = Least Recently Used
Problem: Cache fills up - which item to remove?
LRU Strategy:
• Track when each item last used
• When cache full, remove oldest
• Keep frequently-used items
Example:
Cache has: A(2min ago), B(5min ago), C(1min ago)
Cache full! Remove B (oldest = 5min ago)
Smart: Hot data stays, cold data goes!
Cache Invalidation
Invalidation = Removing stale data from cache
Why needed:
• User's email cached as "old@email.com"
• User updates to "new@email.com"
• Cache still has old value!
• Must remove (invalidate) cached entry
When it happens:
• Every write operation
• Deletes cached row
• Next read will reload from disk
Why writes hurt cache: Constant invalidation!
SSTables & Memtable
Memtable = Recent writes in memory
• Buffer for new writes
• Lives in RAM
• Fast to check (1ms)
SSTable = Sorted String Table (on disk)
• Old data stored on disk
• Immutable (never changes)
• Slow to read (20ms+)
Read path:
1. Check row cache (0.5ms)
2. Check memtable (1ms)
3. Check SSTables (20ms)
YAML Configuration
YAML = Configuration file format
• Human-readable
• Uses indentation (like Python)
• Key: value pairs
cassandra.yaml = Cassandra's config file
• Where you set cache size
• And other settings
Example:
row_cache_size_in_mb: 2048
(means: 2GB row cache)
Location: Usually in /etc/cassandra/
ROI (Return on Investment)
ROI = Money saved or earned
Investment → Return
Example:
• Spend $100 on extra RAM
• Save $12,500/month on servers
• ROI = Pays for itself in 1 day!
Good ROI: Small investment, big savings
Caching has GREAT ROI: Cheap RAM → Huge savings!
🎓 Quick Reference: All Terms
Now you know ALL the vocabulary!
Memory & Storage:
• RAM/Memory = Fast short-term memory (0.1ms)
• Disk = Slow long-term storage (30ms)
• Heap = RAM pool for Java/Cassandra
• Off-heap = RAM outside JVM (no GC impact)
Time & Size:
• Millisecond (ms) = 1/1000 second
• KB/MB/GB = Thousand/Million/Billion bytes
Performance:
• Latency = How long ONE request takes
• Throughput = Requests per second
• P50/P99 = Typical/worst-case performance
Technical:
• JVM = Java runtime (runs Cassandra)
• GC = Automatic memory cleanup (causes pauses)
• I/O = Input/Output (disk read/write)
• LRU = Remove least recently used items
• Invalidation = Remove stale cache entries
Database:
• Node/Server = Computer running Cassandra
• Table = Spreadsheet of data
• Query = Request for data
• Memtable = Recent writes in memory
• SSTable = Old data on disk
Configuration:
• YAML = Config file format
• cassandra.yaml = Cassandra settings
Business:
• ROI = Return on Investment (money saved/earned)
Ready to learn row cache! 🚀
🤔 What is a Cache? (Absolute Basics)
Let's start from the very beginning - imagine you're a student...
📚 The Library Analogy (Everyone Can Understand This!)
Imagine you're studying for exams:
❌ Without a cache (slow way):
• Need a book? Walk to library (5 minutes)
• Find book on shelf (2 minutes)
• Walk back to dorm (5 minutes)
• Total: 12 minutes per book!
• Need 10 books? 120 minutes = 2 hours wasted!
✅ With a cache (smart way):
• Keep frequently-used books on your desk!
• Need that book? Grab it instantly (5 seconds)
• No walking, no searching
• Total: 5 seconds!
• 144x faster! (12 minutes vs 5 seconds)
💡 That's exactly what a cache is:
Your desk = Cache (fast, small, nearby)
Library = Database (slow, huge, far away)
The trade-off:
• Desk is small (maybe 10 books) = Limited cache size
• Library is huge (1 million books) = Full database
• You keep ONLY frequently-used books on desk
• Rarely-used books stay in library
This is caching! Keep hot data close, cold data far.
WITHOUT Cache
Every request:
1. Query arrives: "Get user 123"
2. Check memtable (1ms)
3. Check SSTables on disk (20ms)
4. Read disk, decompress (10ms)
5. Return data
Total: 31ms
1000 requests = 31 seconds!
Every single request hits slow disk.
WITH Cache
First request (miss):
1. Check cache: Not found (0.1ms)
2. Read from disk (31ms)
3. Store in cache for next time
Total: 31.1ms
Next 999 requests (hit):
1. Check cache: Found! (0.5ms)
2. Return immediately
Total: 0.5ms each
1000 requests = 0.53 seconds!
58x faster than without cache!
Key Concept
Cache = Fast but Small
• Lives in RAM (memory)
• 100,000x faster than disk
• But limited size (2GB typical)
Database = Slow but Huge
• Lives on disk
• Very slow access
• But huge capacity (1TB+)
Strategy: Keep hot data in cache!
🎮 Interactive Cache Simulator - Try It Yourself!
Type user IDs below to see cache hits and misses in real-time!
📝 Instructions: Type a user ID and click "Query Cache"
✨ First query will be a MISS (load from disk)
⚡ Subsequent queries will be a HIT (instant from cache)
🎯 Try querying the same user multiple times to see the speed difference!
💡 What You're Learning
First query (MISS): Takes ~30ms because data loaded from slow disk
Subsequent queries (HIT): Takes ~0.5ms because data already in fast cache
60x speed improvement! This is why caching matters so much.
Try these experiments:
1. Query "user123" multiple times → See how fast hits are!
2. Query different users → See misses loading from disk
3. Re-query previous users → They're now cached (hits!)
4. Watch your hit rate improve as you re-query users
⚙️ How Row Cache Works (Step-by-Step)
Let's see the complete journey of a query with row cache!
🎯 Cache Hit vs Cache Miss (The Critical Difference)
Understanding hits and misses is key to cache optimization!
Cache HIT (Good!)
What: Data found in cache
Speed: 0.5ms (instant!)
Example: Query user123, already cached
Path: Check cache → Found → Return
No disk access needed!
Why it happens:
• User queried recently
• Popular "hot" data
• Cache has space for this row
Cache MISS (Slow)
What: Data NOT in cache
Speed: 30ms (60x slower!)
Example: First time querying user456
Path: Cache miss → Check disk → Load → Cache it
Must read from slow disk
Why it happens:
• First time accessing data
• Cache too small (evicted)
• Cold/rarely-used data
Hit Rate (Most Important Metric!)
Formula: Hits / (Hits + Misses) × 100
Example: 90 hits, 10 misses = 90% hit rate
What's good:
• 90%+ = Excellent ✓
• 70-90% = Good
• 50-70% = Needs tuning
• <50% = Cache too small!
Why it matters:
Every 10% improvement = ~3ms faster average latency!
📈 Real Impact: Hit Rate Math
Scenario: 1000 queries per second
With 90% hit rate:
• 900 queries hit cache @ 0.5ms = 450ms total
• 100 queries miss cache @ 30ms = 3000ms total
• Average latency: 3.45ms per query ✓
With 50% hit rate (cache too small!):
• 500 queries hit cache @ 0.5ms = 250ms total
• 500 queries miss cache @ 30ms = 15,000ms total
• Average latency: 15.25ms per query ❌
Result: 90% vs 50% hit rate = 4.4x faster!
For a high-traffic site:
• 15ms latency = Users notice lag
• 3.5ms latency = Feels instant
• This is why cache size and hit rate matter so much!
⚙️ Configuring Row Cache
Cache Size
Parameter: row_cache_size_in_mb
Default: 0 (disabled)
Typical: 100MB - 2GB
Per-table setting!
Example:
32GB RAM server:
• JVM heap: 8GB
• Row cache: 2GB
• OS cache: 22GB
Rule: Row cache should be 10-25% of heap
Row Size Limit
Parameter: row_cache_size_in_mb per table
Best for: Small rows (< 1KB)
Avoid for: Large rows (> 10KB)
Example:
• User profiles: 500 bytes → Great! ✓
• Images/blobs: 100KB → Don't cache ❌
Why: Large rows waste cache space. Cache fills with few rows!
When to Enable
Enable if:
• Read-heavy workload (reads >> writes)
• Small rows (< 1KB typical)
• Hot data (same rows accessed often)
• User profiles, sessions, metadata
Don't enable if:
• Write-heavy (cache invalidated constantly)
• Large rows (wastes space)
• Uniform access (no hot data)
• Time-series (old data never re-read)
💻 Configuration Example (Copy-Paste Ready!)
-- Enable row cache for users table CREATE TABLE users ( id text PRIMARY KEY, name text, email text ) WITH caching = { 'keys': 'ALL', 'rows_per_partition': 'ALL' }; -- In cassandra.yaml row_cache_size_in_mb: 2048 # 2GB cache row_cache_save_period: 14400 # Save every 4 hours
🚀 Performance Impact (Real Numbers)
Latency Reduction
Without row cache:
• P50: 15ms
• P99: 50ms
• Every query hits disk
With row cache (90% hit rate):
• P50: 2ms (7.5x faster!) ✓
• P99: 8ms (6x faster!) ✓
• 90% skip disk entirely
User experience:
15ms = Noticeable lag
2ms = Feels instant!
Throughput Increase
Without cache:
• 30ms per query average
• 33 queries/sec per thread
• Disk is bottleneck
With cache (90% hit rate):
• 3.5ms per query average
• 285 queries/sec per thread
• 8.6x higher throughput! ✓
Result: Same hardware handles 8x more users!
Cost Savings
Scenario: 10,000 req/sec needed
Without cache:
• Need 30 servers @ $500/mo
• Total: $15,000/month
With cache (90% hit):
• Need 5 servers @ $500/mo
• Total: $2,500/month
Savings: $12,500/month = $150k/year!
RAM investment: $100/server for extra RAM. Pays for itself in days!
🏢 Real Company Row Cache Stories
💾 Instagram: 2GB Row Cache for 1B Users
Challenge: User profiles queried millions of times per second. Each profile: 400 bytes (name, bio, follower count, etc.). Without caching = disk overwhelmed!
Solution: Row Cache Configuration
• Cache size: 2GB per node (50 nodes total)
• Row size: 400 bytes average
• Cached users: 5 million most active per node
• Total cached: 250 million users (25% of user base)
Results:
• Hit rate: 92% (most queries hit cache!)
• P50 latency: 0.8ms (was 18ms without cache)
• P99 latency: 3.2ms (was 45ms)
• Throughput: 10x improvement
• Cost: Avoided buying 200 more servers!
Why it worked:
• Hot users (celebrities, influencers) queried constantly
• Small row size (400 bytes = many rows fit)
• 2GB cache holds 5 million profiles
• Read-heavy workload (profile views >> updates)
Instagram Engineering: "Row cache transformed our infrastructure. 92% hit rate means 92% of queries never touch disk. This is 22.5x faster. $100k investment in RAM saved us $2M in server costs!"
🎮 Discord: Session Cache for 150M Users
Use case: User sessions (online status, current channel, voice state)
Pattern: Same users check online status every 30 seconds!
Configuration:
• Row cache: 4GB per node
• Session data: 800 bytes per user
• Cached sessions: 5 million active users per node
• Covers 95% of online users at any time
Implementation:
• Separate table just for sessions (small rows)
• TTL: 1 hour (auto-expire inactive sessions)
• Row cache enabled: 'rows_per_partition': 'ALL'
• Cache invalidation on logout
Results:
• Hit rate: 97% (!)
• Latency: 0.5ms average
• Queries: 500k/sec per node (was 50k without cache)
• User experience: Instant online status updates
Key insight: Sessions are perfect for row cache - small, frequently accessed, read-heavy. 97% hit rate means disk is barely touched!
📱 Spotify: Tiered Caching Strategy
Challenge: 400M users, varying access patterns
Smart Strategy: Different Cache Sizes per Table
1. User Profiles Table:
• Row cache: 2GB
• Hit rate: 90%
• Why: Frequently accessed, small (600 bytes)
2. Listening History Table:
• Row cache: 500MB
• Hit rate: 60%
• Why: Recent history accessed, but large rows (5KB)
3. Playlist Table:
• Row cache: DISABLED
• Why: Large rows (50KB+), not worth caching
• Use key cache instead
Results:
• Overall P99 latency: 12ms
• Profile queries: 1.5ms (cached)
• Playlist queries: 25ms (not cached, but OK)
• Optimal RAM usage: Each table configured for its pattern
Spotify Engineering: "One size doesn't fit all. Analyze each table's access pattern. Cache hot, small data. Skip cold, large data. This strategy saves RAM while maximizing hit rate where it matters!"
✅ Best Practices: Row Cache Optimization
1. Enable for Right Tables
Perfect candidates:
• User profiles, sessions
• Metadata tables
• Reference data
• Small rows (< 1KB)
• Read-heavy (reads:writes > 10:1)
Avoid for:
• Time-series (old data never re-read)
• Large blobs (> 10KB)
• Write-heavy tables
2. Monitor Hit Rate Constantly
Commands:
• nodetool info | grep -A 3 "Row Cache"
• Check hit rate, size, entries
Target: 80%+ hit rate
If hit rate drops:
• Cache too small → increase size
• Access pattern changed → re-evaluate
• Too much churn → check invalidations
3. Size Appropriately
Start conservative:
• Begin with 100-500MB
• Monitor hit rate
• Increase if hit rate < 80%
Rule of thumb:
• Row cache should be 10-25% of JVM heap
• 8GB heap → 1-2GB row cache max
Don't exceed: Larger cache = longer GC pauses!
4. Understand GC Impact
Row cache lives in JVM heap
• Larger cache = more GC pressure
• GC pauses impact all queries!
Symptoms of too large:
• GC pauses > 500ms
• Latency spikes every few minutes
Solution:
• Reduce cache size
• Or use off-heap key cache instead
5. Test Before Production
Don't guess - measure!
• Enable cache in staging first
• Run production-like load tests
• Measure hit rate, latency, GC
Key metrics:
• Hit rate (aim for 80%+)
• P99 latency improvement
• GC pause frequency/duration
If not helping: Don't use it!
6. Warm Up Cache
Problem: After restart, cache empty!
• First 1000 queries slow (all misses)
• Takes time to warm up
Solutions:
• row_cache_save_period: Save cache to disk
• Pre-warm: Query hot keys at startup
• Rolling restart: One node at a time
Result: Minimize cold start impact
🚨 Common Mistakes
- ❌ Enabling for ALL tables: Wastes RAM on cold data
- ❌ Too large cache: GC pauses hurt more than cache helps
- ❌ Caching large rows: Few rows fit, low hit rate
- ❌ Not monitoring hit rate: Don't know if it's helping!
- ❌ Ignoring GC metrics: Cache might be causing pauses
- ❌ Same size for all tables: One size doesn't fit all
- ❌ Enabling on write-heavy tables: Constant invalidation, low hit rate
💼 Interview Questions & Answers
Complete Answer:
Row cache is an in-memory cache in Cassandra that stores complete rows in the JVM heap to dramatically speed up read queries by avoiding disk I/O.
What It Is:
- Location: JVM heap memory (same memory space as Cassandra process)
- Stores: Complete deserialized rows (not just keys)
- Speed: 0.5ms access (vs 30ms disk read = 60x faster)
- Size: Configurable per-table (typically 100MB-2GB)
- Eviction: LRU (Least Recently Used) - oldest unused rows removed when full
- Persistence: Can be saved to disk between restarts
How It Works:
- Query arrives: SELECT * FROM users WHERE id='user123'
- Check cache: Is this row already cached? (0.5ms lookup)
- Cache HIT: Return immediately, skip all disk I/O
- Cache MISS: Read from memtable/SSTables (30ms), then store in cache for next time
When to Use Row Cache (Perfect Scenarios):
1. Read-Heavy Tables:
- Read:write ratio > 10:1
- Example: User profiles, product catalogs
- Why: Writes invalidate cache entries, so write-heavy defeats purpose
2. Small Rows:
- Row size < 1KB ideal
- Example: User sessions (800 bytes), metadata
- Why: 2GB cache holds 2.5 million 800-byte rows vs only 20,000 100KB rows
- More rows cached = higher hit rate
3. Hot Data Access Pattern:
- Same rows accessed repeatedly
- Example: Celebrity profiles on Instagram, popular products
- Why: High hit rate (80%+) is essential for cache to be worthwhile
4. Latency-Sensitive Applications:
- Need sub-5ms P99 latency
- Example: Real-time dashboards, gaming leaderboards
- Why: Cache makes 95%+ of queries sub-millisecond
When NOT to Use Row Cache (Avoid These):
1. Large Rows:
- Row size > 10KB
- Example: Images, documents, large JSON blobs
- Problem: Cache fills with few rows, low hit rate
2. Write-Heavy Tables:
- Write:read ratio > 1:5
- Example: Time-series data, logs
- Problem: Every write invalidates cache entry, constant churn
3. Uniform Access Pattern:
- Every row accessed equally (no "hot" data)
- Example: Analytical scans, full table queries
- Problem: Low hit rate, cache constantly thrashing
4. Time-Series Data:
- Data accessed once then never again
- Example: Logs, metrics with TTL
- Problem: Cache misses on every new data point
Real-World Example (Instagram):
- Scenario: User profiles table, 1 billion users
- Row size: 400 bytes per profile
- Access pattern: Celebrity profiles accessed millions of times/day
- Configuration: 2GB row cache per node
- Capacity: 5 million profiles cached per node
- Hit rate: 92% (hot users always cached)
- Impact: P99 latency 3ms (was 45ms without cache)
- Benefit: 15x faster, avoided 200 additional servers
Trade-offs to Consider:
- Memory cost: Uses JVM heap (increases GC pressure)
- GC impact: Larger cache = longer GC pauses
- Sweet spot: 10-25% of heap size
- Monitor: Hit rate must be 80%+ to justify memory cost
Key Insight: Row cache is NOT a universal optimization. It's a specialized tool for specific workloads: read-heavy, small rows, hot data. When used correctly (90%+ hit rate), it's transformative - 60x faster reads. When misused (low hit rate, large rows), it wastes RAM and increases GC pauses. Always test and measure before enabling in production!
Complete Answer:
Cache hit and miss represent the two possible outcomes when checking if data exists in cache. Hit rate is the most critical cache metric because it directly determines your read performance and cost efficiency.
Cache HIT - The Good Path:
- Definition: Requested data found in cache
- Path: Query → Check cache → Found! → Return immediately
- Latency: ~0.5ms (memory access)
- Disk I/O: Zero (huge win!)
- Example: Query user123, it's cached from previous query
Cache MISS - The Slow Path:
- Definition: Requested data NOT in cache
- Path: Query → Check cache → Not found → Check memtable → Check SSTables on disk → Load & decompress → Store in cache → Return
- Latency: ~30ms (disk access = 60x slower)
- Disk I/O: Full read path executed
- Example: First time querying user456
Hit Rate Formula:
Hit Rate = (Cache Hits / Total Queries) × 100
Example: 900 hits, 100 misses = 900/(900+100) × 100 = 90%
Why Hit Rate Matters - The Math:
Scenario: 1000 queries/second
With 90% hit rate (GOOD):
- 900 queries hit cache @ 0.5ms = 450ms total
- 100 queries miss cache @ 30ms = 3,000ms total
- Total time: 3,450ms
- Average latency: 3.45ms per query ✓
With 50% hit rate (BAD - cache too small):
- 500 queries hit cache @ 0.5ms = 250ms total
- 500 queries miss cache @ 30ms = 15,000ms total
- Total time: 15,250ms
- Average latency: 15.25ms per query ❌
Impact: 90% vs 50% hit rate = 4.4x performance difference!
Real-World Impact Example (Instagram):
Before row cache (0% hit rate):
- Every query hits disk
- P50: 18ms, P99: 45ms
- Disk IOPS maxed out at 50k queries/sec
- Needed 50 database servers
After row cache (92% hit rate):
- 92% skip disk entirely
- P50: 0.8ms, P99: 3.2ms
- Throughput: 500k queries/sec (10x!)
- Needed only 5 servers
Why Each Percentage Point Matters:
- 80% → 85%: 5% more queries skip 30ms disk read = ~1.5ms average latency improvement
- 85% → 90%: Another ~1.5ms improvement
- 90% → 95%: Another ~1.5ms improvement
- Pattern: Linear improvement, every % counts!
Hit Rate Thresholds:
- 95%+: Excellent - cache sized perfectly
- 85-95%: Very good - operating efficiently
- 75-85%: Good - room for improvement
- 60-75%: Mediocre - consider increasing cache size
- <60%: Poor - cache too small OR wrong data cached
Factors Affecting Hit Rate:
1. Cache Size:
- Larger cache = more rows fit = higher hit rate
- But: Larger cache = more GC pressure
- Sweet spot: 10-25% of heap
2. Working Set Size:
- Working set = frequently accessed rows
- If working set fits in cache → high hit rate
- If working set larger than cache → low hit rate
3. Access Pattern:
- Skewed (hot data): High hit rate possible
- Uniform (all data equal): Low hit rate inevitable
- Example: 80/20 rule - 80% of queries hit 20% of data
4. Write Frequency:
- Every write invalidates cached row
- High write rate = constant cache invalidation = low hit rate
Monitoring Hit Rate:
nodetool info | grep -A 3 "Row Cache" # Output example: Row Cache : entries 1250000, size 1.8 GiB, capacity 2 GiB 45000000 hits, 5000000 requests, 0.900 recent hit rate # 90% hit rate = excellent!
Improving Low Hit Rate:
- If <80%: Increase cache size
- Still low: Wrong table for caching (uniform access or large rows)
- Frequent invalidation: Table too write-heavy
- Decision: If can't achieve 80%+, disable cache (wasting RAM)
Cost Implications:
Example: 10,000 queries/sec needed
50% hit rate (inefficient):
- Average 15ms/query
- Need 30 servers @ $500/mo = $15,000/month
90% hit rate (efficient):
- Average 3.5ms/query
- Need 5 servers @ $500/mo = $2,500/month
- Savings: $12,500/month = $150k/year!
Key Insight: Hit rate is THE most important cache metric. The difference between 50% and 90% is the difference between acceptable and terrible performance - literally 4-5x. Monitor it obsessively. If you can't achieve 80%+ hit rate, the cache is wasting resources and should be disabled. When hit rate is high (90%+), cache provides 10-50x performance improvement and massive cost savings!
Responsive Ad