KEY CACHE
Complete Beginner's Guide
The memory cache that makes Cassandra lightning fast! ๐โก
๐ Prerequisites - What You Should Know First
Before learning about Key Cache, let's understand some basic concepts. Don't worry - we'll explain everything from scratch!
What is Cache (pronounced "cash")?
Cache = A fast temporary storage for frequently used data
Real-world analogy:
Think of your study desk:
โข Cache (desk): Books you're using RIGHT NOW โ Grab instantly! โก
โข Storage (bookshelf): All your books โ Walk to shelf, find book, bring back (slow!) ๐
Why useful?
โข Faster access (milliseconds vs seconds)
โข Keeps frequently-used things nearby
โข Saves time on repeated tasks
Computer example:
โข RAM (memory) = Fast cache
โข Hard disk = Slow storage
โข Reading from RAM: 100 nanoseconds โก
โข Reading from disk: 10,000,000 nanoseconds (10ms) ๐
โข 100,000x faster!
What is a Key in Database?
Key = A unique identifier to find data
Real-world example:
Think of a library:
โข Book Title: "Harry Potter" (this is the KEY)
โข Book Location: Shelf 5, Row 3 (this is the VALUE)
โข You search by title (key) to find location (value)
Database example:
User Table:
Key: user_id = 12345
Value: {name: "Alice", email: "alice@email.com", age: 25}
You search by user_id (key) to get user details (value)
What is Disk vs Memory?
Memory (RAM) vs Disk (Hard Drive)
Simple analogy:
โข Memory (RAM): Your desk โ Papers you're working on right now
โข Disk (HDD/SSD): Filing cabinet โ All your documents stored permanently
Key differences:
| Aspect | Memory (RAM) | Disk |
|---|---|---|
| Speed | Super fast! (0.0001ms) | Slow (10ms = 100,000x slower!) |
| Size | Small (16-64 GB) | Large (1-10 TB) |
| Data Persistence | Temporary (lost when power off) | Permanent (saved forever) |
| Cost | Expensive ($10/GB) | Cheap ($0.02/GB) |
What is an Index?
Index = A map that tells you WHERE data is located
Book index analogy:
Back of a textbook:
Index:
"Photosynthesis" โ Page 47
"Cell Division" โ Page 89
"DNA" โ Page 112
Instead of flipping through all 200 pages, you check the index โ Jump directly to page 47! โก
Database index:
Index:
user_id=123 โ Disk Position: Block 5, Offset 240
user_id=456 โ Disk Position: Block 12, Offset 880
The problem: Even the INDEX is on disk! Reading it is still slow! ๐
The solution: Keep the index in MEMORY (cache)! โก
Ready to Continue?
Great! Now you know:
โ
Cache = Fast temporary storage
โ
Key = Unique identifier to find data
โ
Memory is 100,000x faster than disk!
โ
Index = Map showing where data is located
Now let's learn about Key Cache - the clever way to keep indexes in memory! ๐
๐ฌ Instagram: Handling 1 Billion Users with Key Cache
Instagram stores data for 1 billion+ users in Cassandra. Every time you open the app:
โข Load your profile
โข Show your feed
โข Check your messages
โข Display your stories
Each operation needs to find YOUR data using your user_id.
The Challenge (Without Key Cache):
โข Instagram has 1 billion users
โข Each user_id maps to a disk location
โข This mapping is stored in an INDEX on disk
โข Reading index from disk: 10ms per lookup
โข 1 billion users ร 100 reads/day = 100 BILLION disk reads! ๐ฑ
โข Total wasted time: 31 YEARS of waiting every day! โ
The Problem Breakdown:
Query: "Get profile for user_id=987654321"
Step 1: Read INDEX from disk to find data location
โข Time: 10ms ๐
โข This is BEFORE even reading actual data!
Step 2: Read actual data from disk
โข Time: Another 10ms ๐
Total: 20ms per request!
The Solution: Key Cache
โข Store the INDEX in memory (RAM) instead of disk!
โข Size: Only 1% of total data size
โข For 1TB data โ Only 10GB key cache needed
With Key Cache:
Step 1: Read INDEX from memory (key cache)
โข Time: 0.0001ms โก (100,000x faster!)
Step 2: Read actual data from disk
โข Time: 10ms
Total: ~10ms (50% faster!)
Results at Instagram:
โข Disk reads reduced by 50%!
โข Query latency improved from 20ms โ 10ms
โข Key cache size: Only 8GB for 800GB of data
โข Memory overhead: 1% (totally worth it!) โ
โข Can serve 2x more requests with same hardware! ๐
This is the power of Key Cache! ๐โก
โ What is Key Cache?
Now let's understand what Key Cache actually is!
Simple Definition
Key Cache is a memory cache (in RAM) that stores the index of where data is located on disk.
Think of it like a GPS for your data:
โข You want to find user_id=123
โข Key Cache says: "It's at Disk Block 47, Position 1200"
โข You go directly there โ Super fast! โก
Without Key Cache:
โข Read index from disk โ Find location โ Read data from disk
โข 2 disk reads = SLOW! ๐
With Key Cache:
โข Check memory for location โ Read data from disk
โข 1 disk read = FAST! โก
Key Point: We're not caching the DATA itself, just the LOCATION of the data!
Library Analogy
Imagine a huge library with 1 million books:
WITHOUT Key Cache:
You: "I want the book 'Harry Potter'"
Librarian: "Let me check the card catalog..."
โข Walks to filing cabinet (disk)
โข Searches through cards (slow!)
โข Finds: "Shelf 42, Row 7"
โข Walks to shelf
โข Gets book
Total time: 10 minutes ๐
WITH Key Cache:
You: "I want the book 'Harry Potter'"
Librarian: "Already in my notebook (memory)! Shelf 42, Row 7"
โข Instantly knows location โก
โข Walks to shelf
โข Gets book
Total time: 2 minutes โก
The librarian's notebook = Key Cache!
It has popular book locations memorized!
Key Takeaway
Key Cache stores POINTERS, not DATA
โข Doesn't store: user_id=123 โ {name: "Alice", email: "alice@email.com"}
โข DOES store: user_id=123 โ Disk Position: Block 47, Offset 1200
This is smart because:
โข Pointers are tiny (16 bytes)
โข Data can be huge (1KB+)
โข Can fit 100x more pointers in same memory!
Example:
โข 1GB memory can hold 1 million data records (1KB each)
โข OR 62 million pointers (16 bytes each)
โข 62x more coverage! ๐
๐ค Why Do We Need Key Cache?
Problem 1: Disk is SLOW
Every query needs 2 disk reads:
โข Read 1: Find location (index lookup)
โข Read 2: Get actual data
Time breakdown:
โข Each disk read: ~10ms
โข Total per query: 20ms
โข 1 million queries: 20,000 seconds!
โข That's 5.5 hours! โฐ
With Key Cache:
โข Index lookup from memory: 0.0001ms โก
โข Only 1 disk read needed!
โข 1 million queries: 10,000 seconds
โข 50% faster! โ
Problem 2: Index is HUGE
Real numbers:
โข 100 million records
โข Each index entry: 64 bytes
โข Total index size: 6.4 GB!
โข Can't fit in memory... or CAN we? ๐ค
Key Cache solution:
โข Cache only HOT keys (frequently accessed)
โข 80/20 rule: 20% of keys = 80% of traffic
โข Cache 20 million keys = 1.28 GB
โข Fits easily in memory! โ
Result: 80% hit rate with only 20% memory!
Problem 3: Repeated Lookups
Real-world pattern:
โข Same users access system repeatedly
โข You check Instagram 50 times/day
โข Your profile loaded 50 times!
โข Why read index 50 times? ๐ค
With Key Cache:
โข First access: Read index from disk โ Cache it
โข Next 49 accesses: Read from cache โก
โข Index read once, data read 50 times
โข Massive savings! โ
Example - Twitter:
โข Popular users viewed millions of times/day
โข Index cached once
โข Saves millions of disk reads!
The 80/20 Rule (Pareto Principle)
Key Cache leverages a powerful pattern:
โข 20% of your data gets 80% of access
โข Popular users, trending posts, active accounts
โข These "hot" keys stay in cache
โข Cold keys read from disk (rarely)
Real example - YouTube:
โข 1 billion videos total
โข Top 10 million get 80% of views
โข Cache those 10 million keys โ 80% hit rate!
โข Memory needed: Only 1% of total! ๐
๐ Simple Step-by-Step Example
Let's walk through a complete example with a users table!
Setup: Users Table
We have a users table:
CREATE TABLE users (
user_id INT PRIMARY KEY,
name TEXT,
email TEXT,
age INT
);
Data on disk:
User 123: Stored at Disk Block 5, Offset 240
User 456: Stored at Disk Block 12, Offset 880
User 789: Stored at Disk Block 8, Offset 1200
Index on disk:
Index File:
123 โ Block 5, Offset 240
456 โ Block 12, Offset 880
789 โ Block 8, Offset 1200
Scenario 1: WITHOUT Key Cache (Cold Start)
Query: SELECT * FROM users WHERE user_id = 123;
Step 1: Find location (Index lookup)
โข Read index file from disk ๐
โข Search for user_id=123
โข Found: Block 5, Offset 240
โข Time: 10ms
Step 2: Read data
โข Go to Disk Block 5, Offset 240 ๐
โข Read user data
โข Time: 10ms
Total Time: 20ms
Two disk operations! Very slow! ๐
Scenario 2: WITH Key Cache (First Access)
Query: SELECT * FROM users WHERE user_id = 123;
Step 1: Check Key Cache
โข Look in memory for user_id=123 โก
โข MISS! Not in cache (first time)
โข Time: 0.0001ms
Step 2: Read index from disk
โข Read index file from disk ๐
โข Found: Block 5, Offset 240
โข Store in Key Cache! โ Important!
โข Time: 10ms
Step 3: Read data
โข Go to Block 5, Offset 240 ๐
โข Time: 10ms
Total Time: 20ms
Same as without cache (first time), BUT now it's cached! โ
Scenario 3: WITH Key Cache (Second Access - MAGIC!)
Query: SELECT * FROM users WHERE user_id = 123; (again!)
Step 1: Check Key Cache
โข Look in memory for user_id=123 โก
โข HIT! Found: Block 5, Offset 240
โข Time: 0.0001ms โกโกโก
Step 2: Read data
โข Go directly to Block 5, Offset 240 ๐
โข Time: 10ms
Total Time: ~10ms โก
50% faster! Only ONE disk read! ๐๐๐
Key Cache saved us:
โข 1 disk read (10ms)
โข For 1 million queries: 10,000 seconds = 2.7 hours saved!
Understanding the Pattern
The beautiful pattern:
First access: Cache MISS โ Read index from disk โ Cache it
All future accesses: Cache HIT โ Skip disk read โ Super fast! โก
Real-world impact:
โข You check your profile 50 times/day
โข First time: 20ms (cache miss)
โข Next 49 times: 10ms each (cache hit)
โข Total saved: 49 ร 10ms = 490ms
Multiply by millions of users: MASSIVE performance gain! ๐
โ๏ธ How Key Cache Works - Visual Explanation
๐ฎ Interactive Live Demo - Try It Yourself!
Experience Key Cache in action! Query user IDs and watch cache hits/misses.
Key Cache Contents (Max 5 entries)
Cache is empty. Query some user IDs to populate it!
1. Enter a user_id (any number like 123, 456, 789)
2. Click "Query User" to search
3. Watch for CACHE HIT (green) or CACHE MISS (red)
4. Query the same user_id again to see a CACHE HIT!
5. Try querying user_id 123 multiple times to see the magic! โก
Try These Experiments
Experiment 1: First Access (Cache Miss)
โข Query user_id = 100
โข See: CACHE MISS (red) - had to read from disk (20ms)
โข Notice it's now in the cache!
Experiment 2: Second Access (Cache Hit)
โข Query user_id = 100 again
โข See: CACHE HIT (green) - found in memory! (10ms)
โข 50% faster! โก
Experiment 3: Cache Capacity
โข Add user_ids: 1, 2, 3, 4, 5
โข Cache is full (max 5 entries)
โข Add user_id 6
โข Oldest entry (user_id 1) evicted!
โข This is LRU (Least Recently Used) eviction
Experiment 4: Calculate Hit Rate
โข Query: 100, 200, 100, 200, 100, 200
โข First 2 queries: MISS (total 2)
โข Next 4 queries: HIT (total 4)
โข Hit Rate: 4/(4+2) = 67% โ
โข Time saved: 4 queries ร 10ms = 40ms!
๐ Understanding Partition Keys in Cassandra
What is a Partition Key?
Partition Key = The PRIMARY KEY that determines which partition (chunk) your data goes to
Simple analogy - Filing Cabinet:
โข Partition Key = Drawer number
โข All data with same key โ Same drawer
โข Different keys โ Different drawers
Example:
CREATE TABLE users (
user_id INT PRIMARY KEY, โ This is the PARTITION KEY!
name TEXT,
email TEXT
);
What Key Cache stores:
โข Partition key: user_id = 123
โข Location: Which partition + offset
โข This lets Cassandra skip scanning all partitions! โก
How Partition Keys Work
Scenario: 1000 users across 10 partitions
Step 1: Cassandra calculates partition
hash(user_id) % num_partitions
hash(123) % 10 = 3 โ Partition 3
hash(456) % 10 = 6 โ Partition 6
hash(789) % 10 = 9 โ Partition 9
Step 2: Key Cache lookup
โข Check cache for user_id=123
โข If HIT: "user_id=123 is in Partition 3, SSTable 5, Offset 240"
โข If MISS: Read index from disk, then cache it
WITHOUT Key Cache:
โข Must read partition index from disk
โข Then read partition metadata
โข Then read SSTable index
โข Finally read data
โข 4 disk reads! Very slow! ๐
WITH Key Cache:
โข Check memory for location โก
โข Jump directly to data
โข 1 disk read! Fast! โก
Key Cache + Partition Keys = Perfect Match
Why they work so well together:
1. Partition keys are accessed repeatedly
โข Same users log in multiple times
โข Same products viewed multiple times
โข Perfect for caching! โ
2. Partition keys are small
โข Just an integer or UUID (4-16 bytes)
โข + disk location (8-16 bytes)
โข Total: ~32 bytes per entry
โข 1 GB cache = 31 million keys! ๐
3. Lookups are O(1)
โข Hash map in memory
โข Instant lookup regardless of cache size
โข Always fast! โก
๐ Performance Impact & Metrics
Latency Improvement
Without Key Cache:
โข Average: 20ms
โข P50: 18ms
โข P99: 45ms
With Key Cache (80% hit rate):
โข Average: 12ms (40% faster!)
โข P50: 10ms
โข P99: 25ms
Impact: User experience significantly improved!
Memory Usage
Typical Configuration:
โข Total data: 500 GB
โข Key cache size: 5 GB (1%)
โข Entries cached: ~160 million
Memory per entry:
โข 32 bytes average
โข Very efficient! โ
Cost Savings
Reduced disk I/O:
โข 50% fewer disk reads
โข Less wear on SSDs
โข Lower AWS EBS costs
Example savings:
โข 1M queries/min
โข 80% hit rate
โข 800K disk reads saved
โข = Significant $$ saved!
โ๏ธ Cassandra Configuration
# cassandra.yaml configuration
# Key Cache Size (default: auto)
key_cache_size_in_mb: 5120 # 5 GB
# OR use percentage of heap
key_cache_size_in_mb: null
key_cache_size_in_heap: 0.10 # 10% of heap
# Save interval (persist to disk)
key_cache_save_period: 14400 # 4 hours
# Keys to save
key_cache_keys_to_save: 100 # all keys, or specific number
# Monitoring via nodetool
nodetool info | grep "Key Cache"
Output:
Key Cache: entries 52428, size 1.6 MB, capacity 100 MB
Key Cache Hit Rate: 0.856 # 85.6% - Excellent!
Tuning Tips
Size recommendations:
โข Start with 1-5% of total data size
โข Monitor hit rate
โข Target: 80%+ hit rate
โข If < 70%: Increase size
โข If > 95%: Can decrease slightly
When to increase:
โข Low hit rate (< 80%)
โข High read latency
โข Many cache misses in logs
When to decrease:
โข Running low on heap memory
โข Hit rate already very high (> 95%)
โข Can free memory for other caches
๐ข How Real Companies Use Key Cache
๐ฆ Twitter: Timeline Performance
Challenge: Load timelines for 400M users instantly
Key Cache Configuration:
โข Size: 20 GB per node
โข Hit rate: 92%
โข Stores: tweet_id โ partition location
Results:
โข Timeline load: 15ms โ 8ms
โข 47% faster!
โข Handles 10x more requests per server
๐ฌ Netflix: Video Metadata
Use case: Find video details for millions of streams/day
Implementation:
โข Key cache: 10 GB
โข Caches: video_id โ metadata location
โข Hit rate: 89%
Impact:
โข Video start time: 25% faster
โข Disk I/O reduced by 50%
โข Better user experience when browsing catalog
๐พ Key Cache in Cassandra Architecture
How Cassandra Integrates Key Cache
Complete Read Path:
1. Query arrives with partition key
2. Check Key Cache for partition location
3. If HIT: Jump to SSTable location โก
4. If MISS: Read partition index from disk โ Cache it
5. Check Row Cache (if enabled)
6. If HIT: Return data immediately โกโก
7. If MISS: Read from SSTable
8. Check Bloom Filter to skip unnecessary SSTables
9. Read data from correct SSTable
10. Return to client
Key Cache is the FIRST optimization! It makes everything else faster!
-- Real Cassandra example
-- Create table with partition key
CREATE TABLE user_sessions (
user_id UUID PRIMARY KEY, -- Partition key
session_start TIMESTAMP,
session_data TEXT
);
-- Key Cache automatically stores:
-- user_id (UUID) โ (SSTable ID, Partition Offset)
-- Query (utilizes Key Cache)
SELECT * FROM user_sessions
WHERE user_id = 550e8400-e29b-41d4-a716-446655440000;
-- Cassandra flow:
-- 1. Check Key Cache for this user_id
-- 2. If HIT: Go directly to partition location โก
-- 3. If MISS: Read index, cache it, then read data
-- Monitor Key Cache
nodetool tablestats keyspace.user_sessions | grep -i cache
Output:
SSTable count: 15
Key cache hit rate: 0.847 โ 84.7% hit rate โ
๐ผ Interview Questions & Answers
1 What is Key Cache in Cassandra and why is it important?
Answer:
Key Cache is an in-memory cache that stores the locations of partition keys on disk. It maps partition keys to their SSTable positions, allowing Cassandra to skip reading the index from disk.
Why it's important:
- Performance: Reduces read latency by 40-50%
- Disk I/O: Eliminates 50% of disk reads (the index lookup)
- Efficiency: Only 1% memory overhead for massive performance gain
- Scalability: Allows serving 2x more requests with same hardware
What it caches:
- Partition key โ SSTable ID + offset position
- NOT the actual data, just pointers to data
- Typical size: 16-32 bytes per entry
Real-world impact: Without Key Cache, every read requires 2 disk operations. With Key Cache (80% hit rate), 80% of reads need only 1 disk operation.
2 How does Key Cache improve read performance?
Answer:
Key Cache improves read performance by eliminating one disk read operation from the read path.
Without Key Cache:
- Read partition index from disk (10ms) ๐
- Find partition location in index
- Read actual data from disk (10ms) ๐
- Total: 20ms
With Key Cache (HIT):
- Check Key Cache in memory (0.0001ms) โก
- Location found instantly!
- Read actual data from disk (10ms) ๐
- Total: ~10ms (50% faster!)
Performance improvements:
- Latency: Average query time reduced by 40-50%
- Throughput: Can handle 2x more requests/second
- Consistency: P99 latency becomes more predictable
- Cost: Fewer disk IOPS needed = lower cloud costs
Hit rate impact: With 80% hit rate, 80% of queries get 50% speedup = 40% overall improvement!
3 What's the difference between Key Cache and Row Cache?
Answer:
Key Cache and Row Cache serve different purposes in Cassandra's caching strategy:
| Aspect | Key Cache | Row Cache |
|---|---|---|
| What it stores | Partition key โ disk location | Entire row data |
| Size per entry | 16-32 bytes (tiny!) | 100s-1000s bytes (large) |
| Disk reads saved | 1 read (index lookup) | 2 reads (index + data) |
| Speed improvement | 40-50% faster | 90-95% faster |
| Default setting | Enabled (100MB) | Disabled (0MB) |
| Best for | All workloads | Read-heavy, small rows |
When to use:
- Key Cache: Always enable! Minimal memory for great benefit
- Row Cache: Only for read-heavy workloads with small rows
Together: Key Cache (always) + Row Cache (optional) = Maximum performance!
4 How do you monitor and tune Key Cache?
Answer:
Monitoring and tuning Key Cache involves tracking hit rate and adjusting size based on performance.
Monitoring commands:
# Check Key Cache stats nodetool info | grep "Key Cache" Output: Key Cache: entries 52428, size 1.6 MB, capacity 100 MB Key Cache Hit Rate: 0.856 # Per-table stats nodetool tablestats keyspace.table | grep -i cache # JMX metrics org.apache.cassandra.metrics:type=Cache,scope=KeyCache,name=Hits org.apache.cassandra.metrics:type=Cache,scope=KeyCache,name=Requests
Key metrics to watch:
- Hit Rate: Target 80%+ (excellent), 70-80% (good), <70% (increase size)
- Size: Should be 50-90% of capacity (not full, not empty)
- Entries: Number of cached keys
Tuning guidelines:
- If hit rate < 70%: Increase key_cache_size_in_mb
- If size always full: Increase capacity
- If hit rate > 95%: Can decrease size slightly
- If memory constrained: Set key_cache_size_in_heap percentage
Recommended sizes:
- Small dataset (< 100GB): 100-500 MB
- Medium dataset (100GB-1TB): 1-5 GB
- Large dataset (> 1TB): 5-20 GB
- Rule of thumb: 1-5% of total data size
5 What happens when Key Cache is full?
Answer:
When Key Cache reaches capacity, Cassandra uses an LRU (Least Recently Used) eviction policy to make room for new entries.
Eviction process:
- Cache reaches max capacity
- New partition key needs to be cached
- Cassandra finds the LEAST recently used entry
- Evicts (removes) that entry
- Adds new entry in its place
LRU (Least Recently Used) strategy:
- Hot keys: Frequently accessed โ Stay in cache
- Cold keys: Rarely accessed โ Get evicted
- Result: Cache naturally keeps most useful entries!
Example:
Cache (max 5 entries): [key1, key2, key3, key4, key5] โ Full! Access pattern: key3, key5, key3, key6 - key3 accessed โ Stays (recently used) - key5 accessed โ Stays (recently used) - key3 accessed โ Stays (recently used) - key6 needs to be added โ Evict key1 (least recently used) New cache: [key2, key3, key4, key5, key6]
Performance impact:
- Eviction itself: Very fast (microseconds)
- After eviction: Evicted key becomes cache miss next time
- If frequent evictions: Hit rate drops โ Increase cache size!
Signs cache is too small:
- Hit rate < 70%
- Cache always at 100% capacity
- Evictions happening very frequently
- Read latency not improving
Solution: Increase key_cache_size_in_mb until hit rate reaches 80%+
Key Takeaway:
Key Cache is a small but mighty optimization! ๐โก
Just 1-5% memory overhead for 40-50% performance gain!
It's the FIRST line of defense against slow disk reads! ๐
Responsive Ad