Row Cache

ROW CACHE - Lightning-Fast Memory

💾 Master Cassandra's row cache with interactive examples, animations, and real-world scenarios. Learn caching from absolute basics!

📖 Fundamentals (Start Here If You're Brand New!)

Before we dive into row cache, let's make sure we understand ALL the basic terms. No prior knowledge needed!

💾

What is Memory (RAM)?

RAM = Random Access Memory
Think of it as: Your computer's short-term memory

Like your brain:
• Remembers things RIGHT NOW
• Super fast to access
• Forgets when power off
• Limited space

Example: 16GB RAM means 16 gigabytes of "working memory"

Speed: 0.0001 milliseconds (0.1 microseconds)
Nickname: "Memory" = RAM (same thing!)

💿

What is a Disk?

Disk = Storage Drive
Think of it as: Your computer's long-term memory

Two types:
• HDD: Spinning disk (like a CD player)
• SSD: Solid state (like a USB stick)

Properties:
• Remembers forever (even when power off)
• SLOW to access (100,000x slower than RAM!)
• Huge capacity (1TB = 1000GB typical)

Why slow: Physical parts must move to find data

⏱️

Time Units Explained

Millisecond (ms) = 1/1000 of a second
• 1 second = 1,000 milliseconds
• Blink of an eye = 300-400ms

Examples:
• 1ms = Very fast (you can't feel it)
• 10ms = Fast (barely noticeable)
• 100ms = Slow (you notice the delay)
• 1000ms = 1 second (very slow for computer)

Why it matters: Users notice anything over 100ms as "lag"

Microsecond (μs) = 1/1,000,000 of a second (even smaller!)

📏

Size Units (GB, MB, KB)

Byte = Smallest unit (1 letter = 1 byte)

Kilobyte (KB) = 1,000 bytes
• Text message: ~1KB
• Short email: ~5KB

Megabyte (MB) = 1,000 KB
• Photo: ~2-5MB
• Song: ~3-5MB

Gigabyte (GB) = 1,000 MB
• Movie: ~2-4GB
• Typical RAM: 8-32GB

1 GB = 1 billion bytes!

☕

What is JVM and Heap?

JVM = Java Virtual Machine
• Software that runs Java programs
• Cassandra is written in Java
• Think: JVM = the engine running Cassandra

Heap = Memory pool for JVM
• Part of RAM set aside for Java
• Where Java stores objects
• Example: "8GB heap" = 8GB of RAM for Java

Why it matters: Row cache lives IN the heap!

🗑️

What is Garbage Collection (GC)?

GC = Automatic memory cleanup
Like a Roomba vacuum: Runs automatically to clean up

What it does:
• Java creates objects in heap
• Old unused objects pile up
• GC finds and deletes unused objects
• Frees up space

THE PROBLEM: While GC runs, application PAUSES!
• Small heap (8GB): Pause 100ms
• Large heap (32GB): Pause 2-5 seconds!

Why it matters: Larger row cache = more GC pauses!

📊

Latency vs Throughput

Latency = How long ONE thing takes
• "This query took 5 milliseconds"
• Lower is better!
• Like: How long to drive to store

Throughput = How many things per second
• "We handle 10,000 queries per second"
• Higher is better!
• Like: How many cars can use highway per hour

Both matter: Fast individual queries + handle many queries

🎯

Percentiles (P50, P99)

Simple explanation: Performance for different users

Imagine 100 people query:
• Sort their speeds: fastest → slowest
• P50 (median): 50th person's speed
• P99: 99th person's speed (only 1 slower)

Example:
• P50 = 2ms: Typical user experience
• P99 = 20ms: Worst case for 1% of users

Why P99 matters: Don't want even unlucky users to suffer!

🖥️

What is a Node/Server?

Server = A computer
• Physical machine in data center
• Or virtual machine in cloud
• Runs 24/7

Node = Server running Cassandra
• Same thing in Cassandra world
• One Cassandra node = one server

Example: "50 nodes" = 50 servers running Cassandra

Typical server: 32GB RAM, 1TB disk, 16 CPU cores

📋

What is a Table and Query?

Table = Spreadsheet of data
• Rows and columns
• Example: "users" table
• Columns: id, name, email

Query = Request for data
• Written in CQL (Cassandra Query Language)
• Example: "Get user with id=123"
• Like asking database a question

SQL-like: SELECT * FROM users WHERE id='123'

💿

What is Disk I/O?

I/O = Input/Output
Disk I/O = Reading from or writing to disk

Input (Read):
• Get data FROM disk TO memory
• Example: Load user profile from disk

Output (Write):
• Save data FROM memory TO disk
• Example: Store new user profile

Why we care: Disk I/O is SLOW (30ms+)
Goal: Minimize disk I/O with caching!

🔄

LRU (Eviction Policy)

LRU = Least Recently Used
Problem: Cache fills up - which item to remove?

LRU Strategy:
• Track when each item last used
• When cache full, remove oldest
• Keep frequently-used items

Example:
Cache has: A(2min ago), B(5min ago), C(1min ago)
Cache full! Remove B (oldest = 5min ago)

Smart: Hot data stays, cold data goes!

🔥

Cache Invalidation

Invalidation = Removing stale data from cache

Why needed:
• User's email cached as "old@email.com"
• User updates to "new@email.com"
• Cache still has old value!
• Must remove (invalidate) cached entry

When it happens:
• Every write operation
• Deletes cached row
• Next read will reload from disk

Why writes hurt cache: Constant invalidation!

🗄️

SSTables & Memtable

Memtable = Recent writes in memory
• Buffer for new writes
• Lives in RAM
• Fast to check (1ms)

SSTable = Sorted String Table (on disk)
• Old data stored on disk
• Immutable (never changes)
• Slow to read (20ms+)

Read path:
1. Check row cache (0.5ms)
2. Check memtable (1ms)
3. Check SSTables (20ms)

⚙️

YAML Configuration

YAML = Configuration file format
• Human-readable
• Uses indentation (like Python)
• Key: value pairs

cassandra.yaml = Cassandra's config file
• Where you set cache size
• And other settings

Example:
row_cache_size_in_mb: 2048
(means: 2GB row cache)

Location: Usually in /etc/cassandra/

💰

ROI (Return on Investment)

ROI = Money saved or earned
Investment → Return

Example:
• Spend $100 on extra RAM
• Save $12,500/month on servers
• ROI = Pays for itself in 1 day!

Good ROI: Small investment, big savings

Caching has GREAT ROI: Cheap RAM → Huge savings!

🎓 Quick Reference: All Terms

Now you know ALL the vocabulary!

Memory & Storage:
• RAM/Memory = Fast short-term memory (0.1ms)
• Disk = Slow long-term storage (30ms)
• Heap = RAM pool for Java/Cassandra
• Off-heap = RAM outside JVM (no GC impact)

Time & Size:
• Millisecond (ms) = 1/1000 second
• KB/MB/GB = Thousand/Million/Billion bytes

Performance:
• Latency = How long ONE request takes
• Throughput = Requests per second
• P50/P99 = Typical/worst-case performance

Technical:
• JVM = Java runtime (runs Cassandra)
• GC = Automatic memory cleanup (causes pauses)
• I/O = Input/Output (disk read/write)
• LRU = Remove least recently used items
• Invalidation = Remove stale cache entries

Database:
• Node/Server = Computer running Cassandra
• Table = Spreadsheet of data
• Query = Request for data
• Memtable = Recent writes in memory
• SSTable = Old data on disk

Configuration:
• YAML = Config file format
• cassandra.yaml = Cassandra settings

Business:
• ROI = Return on Investment (money saved/earned)

Ready to learn row cache! 🚀

🤔 What is a Cache? (Absolute Basics)

Let's start from the very beginning - imagine you're a student...

📚 The Library Analogy (Everyone Can Understand This!)

Imagine you're studying for exams:

❌ Without a cache (slow way):
• Need a book? Walk to library (5 minutes)
• Find book on shelf (2 minutes)
• Walk back to dorm (5 minutes)
• Total: 12 minutes per book!
• Need 10 books? 120 minutes = 2 hours wasted!

✅ With a cache (smart way):
• Keep frequently-used books on your desk!
• Need that book? Grab it instantly (5 seconds)
• No walking, no searching
• Total: 5 seconds!
• 144x faster! (12 minutes vs 5 seconds)

💡 That's exactly what a cache is:
Your desk = Cache (fast, small, nearby)
Library = Database (slow, huge, far away)

The trade-off:
• Desk is small (maybe 10 books) = Limited cache size
• Library is huge (1 million books) = Full database
• You keep ONLY frequently-used books on desk
• Rarely-used books stay in library

This is caching! Keep hot data close, cold data far.

🐌

WITHOUT Cache

Every request:
1. Query arrives: "Get user 123"
2. Check memtable (1ms)
3. Check SSTables on disk (20ms)
4. Read disk, decompress (10ms)
5. Return data
Total: 31ms

1000 requests = 31 seconds!
Every single request hits slow disk.

⚡

WITH Cache

First request (miss):
1. Check cache: Not found (0.1ms)
2. Read from disk (31ms)
3. Store in cache for next time
Total: 31.1ms

Next 999 requests (hit):
1. Check cache: Found! (0.5ms)
2. Return immediately
Total: 0.5ms each

1000 requests = 0.53 seconds!
58x faster than without cache!

💡

Key Concept

Cache = Fast but Small
• Lives in RAM (memory)
• 100,000x faster than disk
• But limited size (2GB typical)

Database = Slow but Huge
• Lives on disk
• Very slow access
• But huge capacity (1TB+)

Strategy: Keep hot data in cache!

Memory vs Disk Speed Comparison 💾 ROW CACHE (Memory/RAM) Speed: 0.5 milliseconds ⚡ Lightning Fast! 0.5ms 💿 DISK (SSD/HDD) Speed: 30 milliseconds 🐌 60x Slower! 30ms If Memory is 1 second, Disk is 1 MINUTE!

🎮 Interactive Cache Simulator - Try It Yourself!

Type user IDs below to see cache hits and misses in real-time!

🎮 Live Cache Simulator
$
💾 Cache Simulator Ready!
📝 Instructions: Type a user ID and click "Query Cache"
✨ First query will be a MISS (load from disk)
⚡ Subsequent queries will be a HIT (instant from cache)
🎯 Try querying the same user multiple times to see the speed difference!
0
Cache Hits
0
Cache Misses
0%
Hit Rate
0ms
Avg Latency

💡 What You're Learning

First query (MISS): Takes ~30ms because data loaded from slow disk
Subsequent queries (HIT): Takes ~0.5ms because data already in fast cache
60x speed improvement! This is why caching matters so much.

Try these experiments:
1. Query "user123" multiple times → See how fast hits are!
2. Query different users → See misses loading from disk
3. Re-query previous users → They're now cached (hits!)
4. Watch your hit rate improve as you re-query users

⚙️ How Row Cache Works (Step-by-Step)

Let's see the complete journey of a query with row cache!

Row Cache Flow: Hit vs Miss Query Arrives SELECT * FROM users WHERE id='user123' Step 1 💾 ROW CACHE Check: Is user123 cached? Speed: 0.5ms (In-memory lookup) ✅ CACHE HIT! ⚡ Return Data Immediately Total Time: 0.5ms 🎯 Data: {id: "user123", name: "Alice", email: "[email protected]"} ❌ CACHE MISS Step 2: Memtable Check recent writes ~1ms Step 3: SSTables Read from disk ~20ms (slow!) 💿 DISK READ Physical disk access Decompression + Parse Data loaded Step 4: Cache It! Store in row cache For next time ⚡ Return Data Total: ~31ms (this time) 📊 The Magic of Caching First Request (MISS): • Row cache check: 0.5ms • Memtable check: 1ms • Disk read: 20ms + Decompress: 10ms = Total: 31.5ms Subsequent Requests (HIT): • Row cache check: 0.5ms ✓ • Found immediately! Return data. • Total: 0.5ms (63x FASTER!)

🎯 Cache Hit vs Cache Miss (The Critical Difference)

Understanding hits and misses is key to cache optimization!

✅

Cache HIT (Good!)

What: Data found in cache
Speed: 0.5ms (instant!)
Example: Query user123, already cached
Path: Check cache → Found → Return
No disk access needed!

Why it happens:
• User queried recently
• Popular "hot" data
• Cache has space for this row

❌

Cache MISS (Slow)

What: Data NOT in cache
Speed: 30ms (60x slower!)
Example: First time querying user456
Path: Cache miss → Check disk → Load → Cache it
Must read from slow disk

Why it happens:
• First time accessing data
• Cache too small (evicted)
• Cold/rarely-used data

📊

Hit Rate (Most Important Metric!)

Formula: Hits / (Hits + Misses) × 100
Example: 90 hits, 10 misses = 90% hit rate

What's good:
• 90%+ = Excellent ✓
• 70-90% = Good
• 50-70% = Needs tuning
• <50% = Cache too small!

Why it matters:
Every 10% improvement = ~3ms faster average latency!

📈 Real Impact: Hit Rate Math

Scenario: 1000 queries per second

With 90% hit rate:
• 900 queries hit cache @ 0.5ms = 450ms total
• 100 queries miss cache @ 30ms = 3000ms total
• Average latency: 3.45ms per query ✓

With 50% hit rate (cache too small!):
• 500 queries hit cache @ 0.5ms = 250ms total
• 500 queries miss cache @ 30ms = 15,000ms total
• Average latency: 15.25ms per query ❌

Result: 90% vs 50% hit rate = 4.4x faster!

For a high-traffic site:
• 15ms latency = Users notice lag
• 3.5ms latency = Feels instant
• This is why cache size and hit rate matter so much!

⚙️ Configuring Row Cache

💾

Cache Size

Parameter: row_cache_size_in_mb
Default: 0 (disabled)
Typical: 100MB - 2GB
Per-table setting!

Example:
32GB RAM server:
• JVM heap: 8GB
• Row cache: 2GB
• OS cache: 22GB

Rule: Row cache should be 10-25% of heap

📏

Row Size Limit

Parameter: row_cache_size_in_mb per table
Best for: Small rows (< 1KB)
Avoid for: Large rows (> 10KB)

Example:
• User profiles: 500 bytes → Great! ✓
• Images/blobs: 100KB → Don't cache ❌

Why: Large rows waste cache space. Cache fills with few rows!

🔥

When to Enable

Enable if:
• Read-heavy workload (reads >> writes)
• Small rows (< 1KB typical)
• Hot data (same rows accessed often)
• User profiles, sessions, metadata

Don't enable if:
• Write-heavy (cache invalidated constantly)
• Large rows (wastes space)
• Uniform access (no hot data)
• Time-series (old data never re-read)

💻 Configuration Example (Copy-Paste Ready!)

-- Enable row cache for users table
CREATE TABLE users (
    id text PRIMARY KEY,
    name text,
    email text
) WITH caching = {
    'keys': 'ALL',
    'rows_per_partition': 'ALL'
};

-- In cassandra.yaml
row_cache_size_in_mb: 2048  # 2GB cache
row_cache_save_period: 14400  # Save every 4 hours

🚀 Performance Impact (Real Numbers)

⏱️

Latency Reduction

Without row cache:
• P50: 15ms
• P99: 50ms
• Every query hits disk

With row cache (90% hit rate):
• P50: 2ms (7.5x faster!) ✓
• P99: 8ms (6x faster!) ✓
• 90% skip disk entirely

User experience:
15ms = Noticeable lag
2ms = Feels instant!

📈

Throughput Increase

Without cache:
• 30ms per query average
• 33 queries/sec per thread
• Disk is bottleneck

With cache (90% hit rate):
• 3.5ms per query average
• 285 queries/sec per thread
• 8.6x higher throughput! ✓

Result: Same hardware handles 8x more users!

💰

Cost Savings

Scenario: 10,000 req/sec needed

Without cache:
• Need 30 servers @ $500/mo
• Total: $15,000/month

With cache (90% hit):
• Need 5 servers @ $500/mo
• Total: $2,500/month

Savings: $12,500/month = $150k/year!

RAM investment: $100/server for extra RAM. Pays for itself in days!

🏢 Real Company Row Cache Stories

💾 Instagram: 2GB Row Cache for 1B Users

Challenge: User profiles queried millions of times per second. Each profile: 400 bytes (name, bio, follower count, etc.). Without caching = disk overwhelmed!

Solution: Row Cache Configuration
• Cache size: 2GB per node (50 nodes total)
• Row size: 400 bytes average
• Cached users: 5 million most active per node
• Total cached: 250 million users (25% of user base)

Results:
• Hit rate: 92% (most queries hit cache!)
• P50 latency: 0.8ms (was 18ms without cache)
• P99 latency: 3.2ms (was 45ms)
• Throughput: 10x improvement
• Cost: Avoided buying 200 more servers!

Why it worked:
• Hot users (celebrities, influencers) queried constantly
• Small row size (400 bytes = many rows fit)
• 2GB cache holds 5 million profiles
• Read-heavy workload (profile views >> updates)

Instagram Engineering: "Row cache transformed our infrastructure. 92% hit rate means 92% of queries never touch disk. This is 22.5x faster. $100k investment in RAM saved us $2M in server costs!"

🎮 Discord: Session Cache for 150M Users

Use case: User sessions (online status, current channel, voice state)
Pattern: Same users check online status every 30 seconds!

Configuration:
• Row cache: 4GB per node
• Session data: 800 bytes per user
• Cached sessions: 5 million active users per node
• Covers 95% of online users at any time

Implementation:
• Separate table just for sessions (small rows)
• TTL: 1 hour (auto-expire inactive sessions)
• Row cache enabled: 'rows_per_partition': 'ALL'
• Cache invalidation on logout

Results:
• Hit rate: 97% (!)
• Latency: 0.5ms average
• Queries: 500k/sec per node (was 50k without cache)
• User experience: Instant online status updates

Key insight: Sessions are perfect for row cache - small, frequently accessed, read-heavy. 97% hit rate means disk is barely touched!

📱 Spotify: Tiered Caching Strategy

Challenge: 400M users, varying access patterns

Smart Strategy: Different Cache Sizes per Table

1. User Profiles Table:
• Row cache: 2GB
• Hit rate: 90%
• Why: Frequently accessed, small (600 bytes)

2. Listening History Table:
• Row cache: 500MB
• Hit rate: 60%
• Why: Recent history accessed, but large rows (5KB)

3. Playlist Table:
• Row cache: DISABLED
• Why: Large rows (50KB+), not worth caching
• Use key cache instead

Results:
• Overall P99 latency: 12ms
• Profile queries: 1.5ms (cached)
• Playlist queries: 25ms (not cached, but OK)
• Optimal RAM usage: Each table configured for its pattern

Spotify Engineering: "One size doesn't fit all. Analyze each table's access pattern. Cache hot, small data. Skip cold, large data. This strategy saves RAM while maximizing hit rate where it matters!"

✅ Best Practices: Row Cache Optimization

1️⃣

1. Enable for Right Tables

Perfect candidates:
• User profiles, sessions
• Metadata tables
• Reference data
• Small rows (< 1KB)
• Read-heavy (reads:writes > 10:1)

Avoid for:
• Time-series (old data never re-read)
• Large blobs (> 10KB)
• Write-heavy tables

2️⃣

2. Monitor Hit Rate Constantly

Commands:
• nodetool info | grep -A 3 "Row Cache"
• Check hit rate, size, entries

Target: 80%+ hit rate

If hit rate drops:
• Cache too small → increase size
• Access pattern changed → re-evaluate
• Too much churn → check invalidations

3️⃣

3. Size Appropriately

Start conservative:
• Begin with 100-500MB
• Monitor hit rate
• Increase if hit rate < 80%

Rule of thumb:
• Row cache should be 10-25% of JVM heap
• 8GB heap → 1-2GB row cache max

Don't exceed: Larger cache = longer GC pauses!

4️⃣

4. Understand GC Impact

Row cache lives in JVM heap
• Larger cache = more GC pressure
• GC pauses impact all queries!

Symptoms of too large:
• GC pauses > 500ms
• Latency spikes every few minutes

Solution:
• Reduce cache size
• Or use off-heap key cache instead

5️⃣

5. Test Before Production

Don't guess - measure!
• Enable cache in staging first
• Run production-like load tests
• Measure hit rate, latency, GC

Key metrics:
• Hit rate (aim for 80%+)
• P99 latency improvement
• GC pause frequency/duration

If not helping: Don't use it!

6️⃣

6. Warm Up Cache

Problem: After restart, cache empty!
• First 1000 queries slow (all misses)
• Takes time to warm up

Solutions:
• row_cache_save_period: Save cache to disk
• Pre-warm: Query hot keys at startup
• Rolling restart: One node at a time

Result: Minimize cold start impact

🚨 Common Mistakes

  • ❌ Enabling for ALL tables: Wastes RAM on cold data
  • ❌ Too large cache: GC pauses hurt more than cache helps
  • ❌ Caching large rows: Few rows fit, low hit rate
  • ❌ Not monitoring hit rate: Don't know if it's helping!
  • ❌ Ignoring GC metrics: Cache might be causing pauses
  • ❌ Same size for all tables: One size doesn't fit all
  • ❌ Enabling on write-heavy tables: Constant invalidation, low hit rate

💼 Interview Questions & Answers

1
Explain what row cache is and when you should use it in Cassandra

Complete Answer:

Row cache is an in-memory cache in Cassandra that stores complete rows in the JVM heap to dramatically speed up read queries by avoiding disk I/O.

What It Is:

  • Location: JVM heap memory (same memory space as Cassandra process)
  • Stores: Complete deserialized rows (not just keys)
  • Speed: 0.5ms access (vs 30ms disk read = 60x faster)
  • Size: Configurable per-table (typically 100MB-2GB)
  • Eviction: LRU (Least Recently Used) - oldest unused rows removed when full
  • Persistence: Can be saved to disk between restarts

How It Works:

  1. Query arrives: SELECT * FROM users WHERE id='user123'
  2. Check cache: Is this row already cached? (0.5ms lookup)
  3. Cache HIT: Return immediately, skip all disk I/O
  4. Cache MISS: Read from memtable/SSTables (30ms), then store in cache for next time

When to Use Row Cache (Perfect Scenarios):

1. Read-Heavy Tables:

  • Read:write ratio > 10:1
  • Example: User profiles, product catalogs
  • Why: Writes invalidate cache entries, so write-heavy defeats purpose

2. Small Rows:

  • Row size < 1KB ideal
  • Example: User sessions (800 bytes), metadata
  • Why: 2GB cache holds 2.5 million 800-byte rows vs only 20,000 100KB rows
  • More rows cached = higher hit rate

3. Hot Data Access Pattern:

  • Same rows accessed repeatedly
  • Example: Celebrity profiles on Instagram, popular products
  • Why: High hit rate (80%+) is essential for cache to be worthwhile

4. Latency-Sensitive Applications:

  • Need sub-5ms P99 latency
  • Example: Real-time dashboards, gaming leaderboards
  • Why: Cache makes 95%+ of queries sub-millisecond

When NOT to Use Row Cache (Avoid These):

1. Large Rows:

  • Row size > 10KB
  • Example: Images, documents, large JSON blobs
  • Problem: Cache fills with few rows, low hit rate

2. Write-Heavy Tables:

  • Write:read ratio > 1:5
  • Example: Time-series data, logs
  • Problem: Every write invalidates cache entry, constant churn

3. Uniform Access Pattern:

  • Every row accessed equally (no "hot" data)
  • Example: Analytical scans, full table queries
  • Problem: Low hit rate, cache constantly thrashing

4. Time-Series Data:

  • Data accessed once then never again
  • Example: Logs, metrics with TTL
  • Problem: Cache misses on every new data point

Real-World Example (Instagram):

  • Scenario: User profiles table, 1 billion users
  • Row size: 400 bytes per profile
  • Access pattern: Celebrity profiles accessed millions of times/day
  • Configuration: 2GB row cache per node
  • Capacity: 5 million profiles cached per node
  • Hit rate: 92% (hot users always cached)
  • Impact: P99 latency 3ms (was 45ms without cache)
  • Benefit: 15x faster, avoided 200 additional servers

Trade-offs to Consider:

  • Memory cost: Uses JVM heap (increases GC pressure)
  • GC impact: Larger cache = longer GC pauses
  • Sweet spot: 10-25% of heap size
  • Monitor: Hit rate must be 80%+ to justify memory cost

Key Insight: Row cache is NOT a universal optimization. It's a specialized tool for specific workloads: read-heavy, small rows, hot data. When used correctly (90%+ hit rate), it's transformative - 60x faster reads. When misused (low hit rate, large rows), it wastes RAM and increases GC pauses. Always test and measure before enabling in production!

2
Explain cache hit vs cache miss and why hit rate matters so much

Complete Answer:

Cache hit and miss represent the two possible outcomes when checking if data exists in cache. Hit rate is the most critical cache metric because it directly determines your read performance and cost efficiency.

Cache HIT - The Good Path:

  • Definition: Requested data found in cache
  • Path: Query → Check cache → Found! → Return immediately
  • Latency: ~0.5ms (memory access)
  • Disk I/O: Zero (huge win!)
  • Example: Query user123, it's cached from previous query

Cache MISS - The Slow Path:

  • Definition: Requested data NOT in cache
  • Path: Query → Check cache → Not found → Check memtable → Check SSTables on disk → Load & decompress → Store in cache → Return
  • Latency: ~30ms (disk access = 60x slower)
  • Disk I/O: Full read path executed
  • Example: First time querying user456

Hit Rate Formula:

Hit Rate = (Cache Hits / Total Queries) × 100
Example: 900 hits, 100 misses = 900/(900+100) × 100 = 90%

Why Hit Rate Matters - The Math:

Scenario: 1000 queries/second

With 90% hit rate (GOOD):

  • 900 queries hit cache @ 0.5ms = 450ms total
  • 100 queries miss cache @ 30ms = 3,000ms total
  • Total time: 3,450ms
  • Average latency: 3.45ms per query ✓

With 50% hit rate (BAD - cache too small):

  • 500 queries hit cache @ 0.5ms = 250ms total
  • 500 queries miss cache @ 30ms = 15,000ms total
  • Total time: 15,250ms
  • Average latency: 15.25ms per query ❌

Impact: 90% vs 50% hit rate = 4.4x performance difference!

Real-World Impact Example (Instagram):

Before row cache (0% hit rate):

  • Every query hits disk
  • P50: 18ms, P99: 45ms
  • Disk IOPS maxed out at 50k queries/sec
  • Needed 50 database servers

After row cache (92% hit rate):

  • 92% skip disk entirely
  • P50: 0.8ms, P99: 3.2ms
  • Throughput: 500k queries/sec (10x!)
  • Needed only 5 servers

Why Each Percentage Point Matters:

  • 80% → 85%: 5% more queries skip 30ms disk read = ~1.5ms average latency improvement
  • 85% → 90%: Another ~1.5ms improvement
  • 90% → 95%: Another ~1.5ms improvement
  • Pattern: Linear improvement, every % counts!

Hit Rate Thresholds:

  • 95%+: Excellent - cache sized perfectly
  • 85-95%: Very good - operating efficiently
  • 75-85%: Good - room for improvement
  • 60-75%: Mediocre - consider increasing cache size
  • <60%: Poor - cache too small OR wrong data cached

Factors Affecting Hit Rate:

1. Cache Size:

  • Larger cache = more rows fit = higher hit rate
  • But: Larger cache = more GC pressure
  • Sweet spot: 10-25% of heap

2. Working Set Size:

  • Working set = frequently accessed rows
  • If working set fits in cache → high hit rate
  • If working set larger than cache → low hit rate

3. Access Pattern:

  • Skewed (hot data): High hit rate possible
  • Uniform (all data equal): Low hit rate inevitable
  • Example: 80/20 rule - 80% of queries hit 20% of data

4. Write Frequency:

  • Every write invalidates cached row
  • High write rate = constant cache invalidation = low hit rate

Monitoring Hit Rate:

nodetool info | grep -A 3 "Row Cache"

# Output example:
Row Cache              : entries 1250000, size 1.8 GiB, capacity 2 GiB
                         45000000 hits, 5000000 requests, 0.900 recent hit rate
# 90% hit rate = excellent!

Improving Low Hit Rate:

  • If <80%: Increase cache size
  • Still low: Wrong table for caching (uniform access or large rows)
  • Frequent invalidation: Table too write-heavy
  • Decision: If can't achieve 80%+, disable cache (wasting RAM)

Cost Implications:

Example: 10,000 queries/sec needed

50% hit rate (inefficient):

  • Average 15ms/query
  • Need 30 servers @ $500/mo = $15,000/month

90% hit rate (efficient):

  • Average 3.5ms/query
  • Need 5 servers @ $500/mo = $2,500/month
  • Savings: $12,500/month = $150k/year!

Key Insight: Hit rate is THE most important cache metric. The difference between 50% and 90% is the difference between acceptable and terrible performance - literally 4-5x. Monitor it obsessively. If you can't achieve 80%+ hit rate, the cache is wasting resources and should be disabled. When hit rate is high (90%+), cache provides 10-50x performance improvement and massive cost savings!

Advertisement

Responsive Ad