KEY CACHE
Complete Beginner's Guide

The memory cache that makes Cassandra lightning fast! ๐Ÿ”‘โšก

๐Ÿ“– Prerequisites - What You Should Know First

Before learning about Key Cache, let's understand some basic concepts. Don't worry - we'll explain everything from scratch!

What is Cache (pronounced "cash")?

Cache = A fast temporary storage for frequently used data

Real-world analogy:
Think of your study desk:
โ€ข Cache (desk): Books you're using RIGHT NOW โ†’ Grab instantly! โšก
โ€ข Storage (bookshelf): All your books โ†’ Walk to shelf, find book, bring back (slow!) ๐ŸŒ

Why useful?
โ€ข Faster access (milliseconds vs seconds)
โ€ข Keeps frequently-used things nearby
โ€ข Saves time on repeated tasks

Computer example:
โ€ข RAM (memory) = Fast cache
โ€ข Hard disk = Slow storage
โ€ข Reading from RAM: 100 nanoseconds โšก
โ€ข Reading from disk: 10,000,000 nanoseconds (10ms) ๐ŸŒ
โ€ข 100,000x faster!

What is a Key in Database?

Key = A unique identifier to find data

Real-world example:
Think of a library:
โ€ข Book Title: "Harry Potter" (this is the KEY)
โ€ข Book Location: Shelf 5, Row 3 (this is the VALUE)
โ€ข You search by title (key) to find location (value)

Database example:
User Table:
Key: user_id = 12345
Value: {name: "Alice", email: "alice@email.com", age: 25}

You search by user_id (key) to get user details (value)

What is Disk vs Memory?

Memory (RAM) vs Disk (Hard Drive)

Simple analogy:
โ€ข Memory (RAM): Your desk โ†’ Papers you're working on right now
โ€ข Disk (HDD/SSD): Filing cabinet โ†’ All your documents stored permanently

Key differences:

Aspect Memory (RAM) Disk
Speed Super fast! (0.0001ms) Slow (10ms = 100,000x slower!)
Size Small (16-64 GB) Large (1-10 TB)
Data Persistence Temporary (lost when power off) Permanent (saved forever)
Cost Expensive ($10/GB) Cheap ($0.02/GB)

What is an Index?

Index = A map that tells you WHERE data is located

Book index analogy:
Back of a textbook:
Index:
"Photosynthesis" โ†’ Page 47
"Cell Division" โ†’ Page 89
"DNA" โ†’ Page 112

Instead of flipping through all 200 pages, you check the index โ†’ Jump directly to page 47! โšก

Database index:
Index:
user_id=123 โ†’ Disk Position: Block 5, Offset 240
user_id=456 โ†’ Disk Position: Block 12, Offset 880

The problem: Even the INDEX is on disk! Reading it is still slow! ๐ŸŒ
The solution: Keep the index in MEMORY (cache)! โšก

Ready to Continue?

Great! Now you know:
โœ… Cache = Fast temporary storage
โœ… Key = Unique identifier to find data
โœ… Memory is 100,000x faster than disk!
โœ… Index = Map showing where data is located

Now let's learn about Key Cache - the clever way to keep indexes in memory! ๐Ÿš€

๐ŸŽฌ Instagram: Handling 1 Billion Users with Key Cache

Instagram stores data for 1 billion+ users in Cassandra. Every time you open the app:
โ€ข Load your profile
โ€ข Show your feed
โ€ข Check your messages
โ€ข Display your stories

Each operation needs to find YOUR data using your user_id.

The Challenge (Without Key Cache):
โ€ข Instagram has 1 billion users
โ€ข Each user_id maps to a disk location
โ€ข This mapping is stored in an INDEX on disk
โ€ข Reading index from disk: 10ms per lookup
โ€ข 1 billion users ร— 100 reads/day = 100 BILLION disk reads! ๐Ÿ˜ฑ
โ€ข Total wasted time: 31 YEARS of waiting every day! โŒ

The Problem Breakdown:
Query: "Get profile for user_id=987654321"

Step 1: Read INDEX from disk to find data location
โ€ข Time: 10ms ๐ŸŒ
โ€ข This is BEFORE even reading actual data!

Step 2: Read actual data from disk
โ€ข Time: Another 10ms ๐ŸŒ

Total: 20ms per request!

The Solution: Key Cache
โ€ข Store the INDEX in memory (RAM) instead of disk!
โ€ข Size: Only 1% of total data size
โ€ข For 1TB data โ†’ Only 10GB key cache needed

With Key Cache:
Step 1: Read INDEX from memory (key cache)
โ€ข Time: 0.0001ms โšก (100,000x faster!)

Step 2: Read actual data from disk
โ€ข Time: 10ms

Total: ~10ms (50% faster!)

Results at Instagram:
โ€ข Disk reads reduced by 50%!
โ€ข Query latency improved from 20ms โ†’ 10ms
โ€ข Key cache size: Only 8GB for 800GB of data
โ€ข Memory overhead: 1% (totally worth it!) โœ“
โ€ข Can serve 2x more requests with same hardware! ๐Ÿš€

This is the power of Key Cache! ๐Ÿ”‘โšก

โ“ What is Key Cache?

Now let's understand what Key Cache actually is!

Simple Definition

Key Cache is a memory cache (in RAM) that stores the index of where data is located on disk.

Think of it like a GPS for your data:
โ€ข You want to find user_id=123
โ€ข Key Cache says: "It's at Disk Block 47, Position 1200"
โ€ข You go directly there โ†’ Super fast! โšก

Without Key Cache:
โ€ข Read index from disk โ†’ Find location โ†’ Read data from disk
โ€ข 2 disk reads = SLOW! ๐ŸŒ

With Key Cache:
โ€ข Check memory for location โ†’ Read data from disk
โ€ข 1 disk read = FAST! โšก

Key Point: We're not caching the DATA itself, just the LOCATION of the data!

Library Analogy

Imagine a huge library with 1 million books:

WITHOUT Key Cache:
You: "I want the book 'Harry Potter'"
Librarian: "Let me check the card catalog..."
โ€ข Walks to filing cabinet (disk)
โ€ข Searches through cards (slow!)
โ€ข Finds: "Shelf 42, Row 7"
โ€ข Walks to shelf
โ€ข Gets book
Total time: 10 minutes ๐ŸŒ

WITH Key Cache:
You: "I want the book 'Harry Potter'"
Librarian: "Already in my notebook (memory)! Shelf 42, Row 7"
โ€ข Instantly knows location โšก
โ€ข Walks to shelf
โ€ข Gets book
Total time: 2 minutes โšก

The librarian's notebook = Key Cache!
It has popular book locations memorized!

Key Takeaway

Key Cache stores POINTERS, not DATA

โ€ข Doesn't store: user_id=123 โ†’ {name: "Alice", email: "alice@email.com"}
โ€ข DOES store: user_id=123 โ†’ Disk Position: Block 47, Offset 1200

This is smart because:
โ€ข Pointers are tiny (16 bytes)
โ€ข Data can be huge (1KB+)
โ€ข Can fit 100x more pointers in same memory!

Example:
โ€ข 1GB memory can hold 1 million data records (1KB each)
โ€ข OR 62 million pointers (16 bytes each)
โ€ข 62x more coverage! ๐Ÿš€

๐Ÿค” Why Do We Need Key Cache?

๐ŸŒ

Problem 1: Disk is SLOW

Every query needs 2 disk reads:
โ€ข Read 1: Find location (index lookup)
โ€ข Read 2: Get actual data

Time breakdown:
โ€ข Each disk read: ~10ms
โ€ข Total per query: 20ms
โ€ข 1 million queries: 20,000 seconds!
โ€ข That's 5.5 hours! โฐ

With Key Cache:
โ€ข Index lookup from memory: 0.0001ms โšก
โ€ข Only 1 disk read needed!
โ€ข 1 million queries: 10,000 seconds
โ€ข 50% faster! โœ“

๐Ÿ’ฐ

Problem 2: Index is HUGE

Real numbers:
โ€ข 100 million records
โ€ข Each index entry: 64 bytes
โ€ข Total index size: 6.4 GB!
โ€ข Can't fit in memory... or CAN we? ๐Ÿค”

Key Cache solution:
โ€ข Cache only HOT keys (frequently accessed)
โ€ข 80/20 rule: 20% of keys = 80% of traffic
โ€ข Cache 20 million keys = 1.28 GB
โ€ข Fits easily in memory! โœ“

Result: 80% hit rate with only 20% memory!

๐Ÿ”ฅ

Problem 3: Repeated Lookups

Real-world pattern:
โ€ข Same users access system repeatedly
โ€ข You check Instagram 50 times/day
โ€ข Your profile loaded 50 times!
โ€ข Why read index 50 times? ๐Ÿค”

With Key Cache:
โ€ข First access: Read index from disk โ†’ Cache it
โ€ข Next 49 accesses: Read from cache โšก
โ€ข Index read once, data read 50 times
โ€ข Massive savings! โœ“

Example - Twitter:
โ€ข Popular users viewed millions of times/day
โ€ข Index cached once
โ€ข Saves millions of disk reads!

The 80/20 Rule (Pareto Principle)

Key Cache leverages a powerful pattern:

โ€ข 20% of your data gets 80% of access
โ€ข Popular users, trending posts, active accounts
โ€ข These "hot" keys stay in cache
โ€ข Cold keys read from disk (rarely)

Real example - YouTube:
โ€ข 1 billion videos total
โ€ข Top 10 million get 80% of views
โ€ข Cache those 10 million keys โ†’ 80% hit rate!
โ€ข Memory needed: Only 1% of total! ๐Ÿš€

๐Ÿ“ Simple Step-by-Step Example

Let's walk through a complete example with a users table!

Setup: Users Table

We have a users table:
CREATE TABLE users (
  user_id INT PRIMARY KEY,
  name TEXT,
  email TEXT,
  age INT
);

Data on disk:
User 123: Stored at Disk Block 5, Offset 240
User 456: Stored at Disk Block 12, Offset 880
User 789: Stored at Disk Block 8, Offset 1200

Index on disk:
Index File:
123 โ†’ Block 5, Offset 240
456 โ†’ Block 12, Offset 880
789 โ†’ Block 8, Offset 1200

Scenario 1: WITHOUT Key Cache (Cold Start)

Query: SELECT * FROM users WHERE user_id = 123;

Step 1: Find location (Index lookup)
โ€ข Read index file from disk ๐ŸŒ
โ€ข Search for user_id=123
โ€ข Found: Block 5, Offset 240
โ€ข Time: 10ms

Step 2: Read data
โ€ข Go to Disk Block 5, Offset 240 ๐ŸŒ
โ€ข Read user data
โ€ข Time: 10ms

Total Time: 20ms
Two disk operations! Very slow! ๐ŸŒ

Scenario 2: WITH Key Cache (First Access)

Query: SELECT * FROM users WHERE user_id = 123;

Step 1: Check Key Cache
โ€ข Look in memory for user_id=123 โšก
โ€ข MISS! Not in cache (first time)
โ€ข Time: 0.0001ms

Step 2: Read index from disk
โ€ข Read index file from disk ๐ŸŒ
โ€ข Found: Block 5, Offset 240
โ€ข Store in Key Cache! โ† Important!
โ€ข Time: 10ms

Step 3: Read data
โ€ข Go to Block 5, Offset 240 ๐ŸŒ
โ€ข Time: 10ms

Total Time: 20ms
Same as without cache (first time), BUT now it's cached! โœ“

Scenario 3: WITH Key Cache (Second Access - MAGIC!)

Query: SELECT * FROM users WHERE user_id = 123; (again!)

Step 1: Check Key Cache
โ€ข Look in memory for user_id=123 โšก
โ€ข HIT! Found: Block 5, Offset 240
โ€ข Time: 0.0001ms โšกโšกโšก

Step 2: Read data
โ€ข Go directly to Block 5, Offset 240 ๐ŸŒ
โ€ข Time: 10ms

Total Time: ~10ms โšก
50% faster! Only ONE disk read! ๐Ÿš€๐Ÿš€๐Ÿš€

Key Cache saved us:
โ€ข 1 disk read (10ms)
โ€ข For 1 million queries: 10,000 seconds = 2.7 hours saved!

Understanding the Pattern

The beautiful pattern:

First access: Cache MISS โ†’ Read index from disk โ†’ Cache it
All future accesses: Cache HIT โ†’ Skip disk read โ†’ Super fast! โšก

Real-world impact:
โ€ข You check your profile 50 times/day
โ€ข First time: 20ms (cache miss)
โ€ข Next 49 times: 10ms each (cache hit)
โ€ข Total saved: 49 ร— 10ms = 490ms

Multiply by millions of users: MASSIVE performance gain! ๐Ÿš€

โš™๏ธ How Key Cache Works - Visual Explanation

Key Cache: Lookup Flow Query SELECT * FROM users WHERE user_id = 123 Step 1 ๐Ÿ”‘ Key Cache (Memory / RAM) Check: user_id=123? Time: 0.0001ms โšก CACHE HIT โœ“ Location Found! Block 5, Offset 240 Go directly to disk Skip index read! โšก CACHE MISS โŒ ๐Ÿ“„ Index on Disk Read index file... user_id=123 โ†’ Block 5, Offset 240 Time: 10ms ๐ŸŒ Store in cache! Step 2/3 ๐Ÿ’พ Data on Disk Go to: Block 5, Offset 240 Read actual data: {user_id: 123, name: "Alice", email: "alice@..."} Time: 10ms ๐ŸŒ โœ… CACHE HIT (2nd+ access) Step 1: Check cache (0.0001ms) Step 2: Read data (10ms) Total: ~10ms โšก โŒ CACHE MISS (1st access) Step 1: Check cache (0.0001ms) Step 2: Read index (10ms) โ†’ Cache it Step 3: Read data (10ms) Total: 20ms ๐ŸŒ

๐ŸŽฎ Interactive Live Demo - Try It Yourself!

Experience Key Cache in action! Query user IDs and watch cache hits/misses.

๐Ÿ”‘ Key Cache Simulator
Key Cache Contents (Max 5 entries)

Cache is empty. Query some user IDs to populate it!

0
Total Queries
0
Cache Hits โœ“
0
Cache Misses โŒ
0%
Hit Rate
0ms
Time Saved
๐Ÿ’ก Instructions:
1. Enter a user_id (any number like 123, 456, 789)
2. Click "Query User" to search
3. Watch for CACHE HIT (green) or CACHE MISS (red)
4. Query the same user_id again to see a CACHE HIT!
5. Try querying user_id 123 multiple times to see the magic! โšก
Try These Experiments

Experiment 1: First Access (Cache Miss)
โ€ข Query user_id = 100
โ€ข See: CACHE MISS (red) - had to read from disk (20ms)
โ€ข Notice it's now in the cache!

Experiment 2: Second Access (Cache Hit)
โ€ข Query user_id = 100 again
โ€ข See: CACHE HIT (green) - found in memory! (10ms)
โ€ข 50% faster! โšก

Experiment 3: Cache Capacity
โ€ข Add user_ids: 1, 2, 3, 4, 5
โ€ข Cache is full (max 5 entries)
โ€ข Add user_id 6
โ€ข Oldest entry (user_id 1) evicted!
โ€ข This is LRU (Least Recently Used) eviction

Experiment 4: Calculate Hit Rate
โ€ข Query: 100, 200, 100, 200, 100, 200
โ€ข First 2 queries: MISS (total 2)
โ€ข Next 4 queries: HIT (total 4)
โ€ข Hit Rate: 4/(4+2) = 67% โœ“
โ€ข Time saved: 4 queries ร— 10ms = 40ms!

๐Ÿ”‘ Understanding Partition Keys in Cassandra

What is a Partition Key?

Partition Key = The PRIMARY KEY that determines which partition (chunk) your data goes to

Simple analogy - Filing Cabinet:
โ€ข Partition Key = Drawer number
โ€ข All data with same key โ†’ Same drawer
โ€ข Different keys โ†’ Different drawers

Example:
CREATE TABLE users (
  user_id INT PRIMARY KEY, โ† This is the PARTITION KEY!
  name TEXT,
  email TEXT
);

What Key Cache stores:
โ€ข Partition key: user_id = 123
โ€ข Location: Which partition + offset
โ€ข This lets Cassandra skip scanning all partitions! โšก

How Partition Keys Work

Scenario: 1000 users across 10 partitions

Step 1: Cassandra calculates partition
hash(user_id) % num_partitions
hash(123) % 10 = 3 โ†’ Partition 3
hash(456) % 10 = 6 โ†’ Partition 6
hash(789) % 10 = 9 โ†’ Partition 9

Step 2: Key Cache lookup
โ€ข Check cache for user_id=123
โ€ข If HIT: "user_id=123 is in Partition 3, SSTable 5, Offset 240"
โ€ข If MISS: Read index from disk, then cache it

WITHOUT Key Cache:
โ€ข Must read partition index from disk
โ€ข Then read partition metadata
โ€ข Then read SSTable index
โ€ข Finally read data
โ€ข 4 disk reads! Very slow! ๐ŸŒ

WITH Key Cache:
โ€ข Check memory for location โšก
โ€ข Jump directly to data
โ€ข 1 disk read! Fast! โšก

Key Cache + Partition Keys = Perfect Match

Why they work so well together:

1. Partition keys are accessed repeatedly
โ€ข Same users log in multiple times
โ€ข Same products viewed multiple times
โ€ข Perfect for caching! โœ“

2. Partition keys are small
โ€ข Just an integer or UUID (4-16 bytes)
โ€ข + disk location (8-16 bytes)
โ€ข Total: ~32 bytes per entry
โ€ข 1 GB cache = 31 million keys! ๐Ÿš€

3. Lookups are O(1)
โ€ข Hash map in memory
โ€ข Instant lookup regardless of cache size
โ€ข Always fast! โšก

๐Ÿ“Š Performance Impact & Metrics

โšก

Latency Improvement

Without Key Cache:
โ€ข Average: 20ms
โ€ข P50: 18ms
โ€ข P99: 45ms

With Key Cache (80% hit rate):
โ€ข Average: 12ms (40% faster!)
โ€ข P50: 10ms
โ€ข P99: 25ms

Impact: User experience significantly improved!

๐Ÿ’พ

Memory Usage

Typical Configuration:
โ€ข Total data: 500 GB
โ€ข Key cache size: 5 GB (1%)
โ€ข Entries cached: ~160 million

Memory per entry:
โ€ข 32 bytes average
โ€ข Very efficient! โœ“

๐Ÿ’ฐ

Cost Savings

Reduced disk I/O:
โ€ข 50% fewer disk reads
โ€ข Less wear on SSDs
โ€ข Lower AWS EBS costs

Example savings:
โ€ข 1M queries/min
โ€ข 80% hit rate
โ€ข 800K disk reads saved
โ€ข = Significant $$ saved!

โš™๏ธ Cassandra Configuration

# cassandra.yaml configuration # Key Cache Size (default: auto) key_cache_size_in_mb: 5120 # 5 GB # OR use percentage of heap key_cache_size_in_mb: null key_cache_size_in_heap: 0.10 # 10% of heap # Save interval (persist to disk) key_cache_save_period: 14400 # 4 hours # Keys to save key_cache_keys_to_save: 100 # all keys, or specific number # Monitoring via nodetool nodetool info | grep "Key Cache" Output: Key Cache: entries 52428, size 1.6 MB, capacity 100 MB Key Cache Hit Rate: 0.856 # 85.6% - Excellent!

Tuning Tips

Size recommendations:
โ€ข Start with 1-5% of total data size
โ€ข Monitor hit rate
โ€ข Target: 80%+ hit rate
โ€ข If < 70%: Increase size
โ€ข If > 95%: Can decrease slightly

When to increase:
โ€ข Low hit rate (< 80%)
โ€ข High read latency
โ€ข Many cache misses in logs

When to decrease:
โ€ข Running low on heap memory
โ€ข Hit rate already very high (> 95%)
โ€ข Can free memory for other caches

๐Ÿข How Real Companies Use Key Cache

๐Ÿฆ Twitter: Timeline Performance

Challenge: Load timelines for 400M users instantly

Key Cache Configuration:
โ€ข Size: 20 GB per node
โ€ข Hit rate: 92%
โ€ข Stores: tweet_id โ†’ partition location

Results:
โ€ข Timeline load: 15ms โ†’ 8ms
โ€ข 47% faster!
โ€ข Handles 10x more requests per server

๐ŸŽฌ Netflix: Video Metadata

Use case: Find video details for millions of streams/day

Implementation:
โ€ข Key cache: 10 GB
โ€ข Caches: video_id โ†’ metadata location
โ€ข Hit rate: 89%

Impact:
โ€ข Video start time: 25% faster
โ€ข Disk I/O reduced by 50%
โ€ข Better user experience when browsing catalog

๐Ÿ’พ Key Cache in Cassandra Architecture

How Cassandra Integrates Key Cache

Complete Read Path:

1. Query arrives with partition key
2. Check Key Cache for partition location
3. If HIT: Jump to SSTable location โšก
4. If MISS: Read partition index from disk โ†’ Cache it
5. Check Row Cache (if enabled)
6. If HIT: Return data immediately โšกโšก
7. If MISS: Read from SSTable
8. Check Bloom Filter to skip unnecessary SSTables
9. Read data from correct SSTable
10. Return to client

Key Cache is the FIRST optimization! It makes everything else faster!

-- Real Cassandra example -- Create table with partition key CREATE TABLE user_sessions ( user_id UUID PRIMARY KEY, -- Partition key session_start TIMESTAMP, session_data TEXT ); -- Key Cache automatically stores: -- user_id (UUID) โ†’ (SSTable ID, Partition Offset) -- Query (utilizes Key Cache) SELECT * FROM user_sessions WHERE user_id = 550e8400-e29b-41d4-a716-446655440000; -- Cassandra flow: -- 1. Check Key Cache for this user_id -- 2. If HIT: Go directly to partition location โšก -- 3. If MISS: Read index, cache it, then read data -- Monitor Key Cache nodetool tablestats keyspace.user_sessions | grep -i cache Output: SSTable count: 15 Key cache hit rate: 0.847 โ† 84.7% hit rate โœ“

๐Ÿ’ผ Interview Questions & Answers

1 What is Key Cache in Cassandra and why is it important?

Answer:

Key Cache is an in-memory cache that stores the locations of partition keys on disk. It maps partition keys to their SSTable positions, allowing Cassandra to skip reading the index from disk.

Why it's important:

  • Performance: Reduces read latency by 40-50%
  • Disk I/O: Eliminates 50% of disk reads (the index lookup)
  • Efficiency: Only 1% memory overhead for massive performance gain
  • Scalability: Allows serving 2x more requests with same hardware

What it caches:

  • Partition key โ†’ SSTable ID + offset position
  • NOT the actual data, just pointers to data
  • Typical size: 16-32 bytes per entry

Real-world impact: Without Key Cache, every read requires 2 disk operations. With Key Cache (80% hit rate), 80% of reads need only 1 disk operation.

2 How does Key Cache improve read performance?

Answer:

Key Cache improves read performance by eliminating one disk read operation from the read path.

Without Key Cache:

  1. Read partition index from disk (10ms) ๐ŸŒ
  2. Find partition location in index
  3. Read actual data from disk (10ms) ๐ŸŒ
  4. Total: 20ms

With Key Cache (HIT):

  1. Check Key Cache in memory (0.0001ms) โšก
  2. Location found instantly!
  3. Read actual data from disk (10ms) ๐ŸŒ
  4. Total: ~10ms (50% faster!)

Performance improvements:

  • Latency: Average query time reduced by 40-50%
  • Throughput: Can handle 2x more requests/second
  • Consistency: P99 latency becomes more predictable
  • Cost: Fewer disk IOPS needed = lower cloud costs

Hit rate impact: With 80% hit rate, 80% of queries get 50% speedup = 40% overall improvement!

3 What's the difference between Key Cache and Row Cache?

Answer:

Key Cache and Row Cache serve different purposes in Cassandra's caching strategy:

Aspect Key Cache Row Cache
What it stores Partition key โ†’ disk location Entire row data
Size per entry 16-32 bytes (tiny!) 100s-1000s bytes (large)
Disk reads saved 1 read (index lookup) 2 reads (index + data)
Speed improvement 40-50% faster 90-95% faster
Default setting Enabled (100MB) Disabled (0MB)
Best for All workloads Read-heavy, small rows

When to use:

  • Key Cache: Always enable! Minimal memory for great benefit
  • Row Cache: Only for read-heavy workloads with small rows

Together: Key Cache (always) + Row Cache (optional) = Maximum performance!

4 How do you monitor and tune Key Cache?

Answer:

Monitoring and tuning Key Cache involves tracking hit rate and adjusting size based on performance.

Monitoring commands:

# Check Key Cache stats
nodetool info | grep "Key Cache"

Output:
Key Cache: entries 52428, size 1.6 MB, capacity 100 MB
Key Cache Hit Rate: 0.856

# Per-table stats
nodetool tablestats keyspace.table | grep -i cache

# JMX metrics
org.apache.cassandra.metrics:type=Cache,scope=KeyCache,name=Hits
org.apache.cassandra.metrics:type=Cache,scope=KeyCache,name=Requests

Key metrics to watch:

  • Hit Rate: Target 80%+ (excellent), 70-80% (good), <70% (increase size)
  • Size: Should be 50-90% of capacity (not full, not empty)
  • Entries: Number of cached keys

Tuning guidelines:

  • If hit rate < 70%: Increase key_cache_size_in_mb
  • If size always full: Increase capacity
  • If hit rate > 95%: Can decrease size slightly
  • If memory constrained: Set key_cache_size_in_heap percentage

Recommended sizes:

  • Small dataset (< 100GB): 100-500 MB
  • Medium dataset (100GB-1TB): 1-5 GB
  • Large dataset (> 1TB): 5-20 GB
  • Rule of thumb: 1-5% of total data size
5 What happens when Key Cache is full?

Answer:

When Key Cache reaches capacity, Cassandra uses an LRU (Least Recently Used) eviction policy to make room for new entries.

Eviction process:

  1. Cache reaches max capacity
  2. New partition key needs to be cached
  3. Cassandra finds the LEAST recently used entry
  4. Evicts (removes) that entry
  5. Adds new entry in its place

LRU (Least Recently Used) strategy:

  • Hot keys: Frequently accessed โ†’ Stay in cache
  • Cold keys: Rarely accessed โ†’ Get evicted
  • Result: Cache naturally keeps most useful entries!

Example:

Cache (max 5 entries):
[key1, key2, key3, key4, key5] โ† Full!

Access pattern: key3, key5, key3, key6
- key3 accessed โ†’ Stays (recently used)
- key5 accessed โ†’ Stays (recently used)
- key3 accessed โ†’ Stays (recently used)
- key6 needs to be added โ†’ Evict key1 (least recently used)

New cache: [key2, key3, key4, key5, key6]

Performance impact:

  • Eviction itself: Very fast (microseconds)
  • After eviction: Evicted key becomes cache miss next time
  • If frequent evictions: Hit rate drops โ†’ Increase cache size!

Signs cache is too small:

  • Hit rate < 70%
  • Cache always at 100% capacity
  • Evictions happening very frequently
  • Read latency not improving

Solution: Increase key_cache_size_in_mb until hit rate reaches 80%+

Key Takeaway:
Key Cache is a small but mighty optimization! ๐Ÿ”‘โšก
Just 1-5% memory overhead for 40-50% performance gain!
It's the FIRST line of defense against slow disk reads! ๐Ÿš€

Advertisement

Responsive Ad