Cassandra Skinny Rows
Master the narrow partition pattern! Learn key-value optimization, entity storage, when to use single-row partitions with animations, production examples, and expert-level design patterns.
📖 The Story: The Dictionary You Know
How do you find a word in a dictionary?
📖 The Dictionary Pattern
Step 1: Go to "C" section
Step 2: Navigate to "Ca"
Step 3: Find "Cassandra"
Step 4: Read ONE entry
┌──────────────────────────────────┐
│ Word: Cassandra │
│ Definition: A distributed... │
│ Example: We use Cassandra... │
│ Synonyms: NoSQL, Database │
└──────────────────────────────────┘
ONE word = ONE entry = ONE lookup!
Key Characteristics:
- Each word = Separate entry
- One lookup = One result
- Fast direct access
- No need to scan multiple entries
📄 This is Skinny Rows in Cassandra!
Each user_id = ONE partition = ONE row
user_id='alice' →
┌─────────────────────────────────┐
│ ONE PARTITION with ONE ROW │
│ user_id: alice │
│ name: Alice Smith │
│ email: alice@example.com │
│ created: 2024-01-15 │
└─────────────────────────────────┘
ONE key = ONE partition = ONE row = FAST!
Perfect For:
- User profiles (1 user = 1 row)
- Product catalog (1 product = 1 row)
- Configuration (1 setting = 1 row)
- Session store (1 session = 1 row)
- Cache entries (1 key = 1 value)
📄 What Are Skinny Rows?
Understanding the narrow partition pattern.
Complete Definition
Skinny Row (Narrow Partition): A Cassandra partition containing exactly ONE row or very few rows (typically < 10).
Core Pattern:
- One key → One partition
- One partition → One (or very few) rows
- No clustering columns (or used minimally)
- Think: Key-value store, entity lookup, dictionary
PRIMARY KEY (user_id)
↑
Partition Key ONLY (no clustering!)
Result:
user_id='alice' → ONE partition with ONE row
user_id='bob' → ONE partition with ONE row
user_id='carol' → ONE partition with ONE row
Each lookup returns exactly ONE row!
Skinny vs Wide Rows: Visual Comparison
Skinny Rows
Structure:
- Each entity = 1 partition
- Each partition = 1 row
- Direct key lookup
- Like dictionary entry
Best For:
- User profiles
- Product details
- Configuration
- Session data
Wide Rows
Structure:
- Each entity = 1 partition
- Each partition = many rows
- Range queries
- Like spreadsheet
Best For:
- Time-series data
- Activity logs
- Message history
- Event streams
🔑 The Key-Value Pattern
Using Cassandra as a distributed key-value store.
Perfect Use Case: Session Store
session_id UUID PRIMARY KEY,
user_id UUID,
login_time TIMESTAMP,
last_activity TIMESTAMP,
ip_address TEXT,
user_agent TEXT
) WITH default_time_to_live = 86400; -- 24 hour TTL
Pattern Characteristics:
- Primary Key = Partition Key Only: No clustering columns
- One Session = One Row: Direct lookup by session_id
- TTL Enabled: Sessions automatically expire
- Fast Reads: Single partition lookup (1-2ms)
- Fast Writes: Simple insert/update
Query Examples:
SELECT * FROM user_sessions
WHERE session_id = ?;
-- Update last activity:
UPDATE user_sessions
SET last_activity = ?
WHERE session_id = ?;
-- Delete session (logout):
DELETE FROM user_sessions
WHERE session_id = ?;
Performance Benefits
Why Skinny Rows Are Fast:
- O(1) Hash Lookup: Direct partition location
- No Scanning: Single row, no iteration
- Cache-Friendly: Small, frequently accessed data
- Predictable Performance: Consistent 1-2ms reads
- Write Optimization: No clustering overhead
• Read latency: 1-2ms (single partition)
• Write latency: 1-2ms (single partition)
• Throughput: 100K+ ops/sec per node
• Scalability: Linear with cluster size
🗃️ Entity Storage Pattern
Storing complete entities in single rows.
User Profiles
user_id UUID PRIMARY KEY,
username TEXT,
email TEXT,
full_name TEXT,
created_at TIMESTAMP,
profile_image_url TEXT
);
Pattern:
- 1 user = 1 row
- All user data together
- Fast user lookup
- Simple updates
Product Catalog
product_id UUID PRIMARY KEY,
name TEXT,
price DECIMAL,
description TEXT,
category TEXT,
stock_count INT
);
Pattern:
- 1 product = 1 row
- Product details co-located
- Fast product page load
- Easy inventory updates
Configuration
config_key TEXT PRIMARY KEY,
config_value TEXT,
updated_at TIMESTAMP,
updated_by TEXT
);
Pattern:
- 1 setting = 1 row
- Fast config lookup
- Easy updates
- Version tracking
⚖️ Skinny vs Wide Rows: Decision Guide
When to use each pattern.
Use Skinny Rows When:
- Entity Lookup: Get by ID/key
- Key-Value Store: Session, cache
- Single Object: User, product, config
- No History Needed: Current state only
- Simple Updates: Overwrite entire row
- TTL Usage: Auto-expiring data
- Distributed Writes: Even load
• User profiles
• Product details
• Session data
• Configuration
• Cache entries
Use Wide Rows When:
- Time-Series: Events over time
- History/Logs: Activity tracking
- Related Data: Multiple items per entity
- Range Queries: Get by time range
- Append-Only: Adding new events
- Sequential Access: Read multiple rows
- Co-Location: All data together
• User activity logs
• IoT sensor data
• Message history
• Transaction history
• Photo albums
⚡ Performance Characteristics
Understanding skinny row performance.
Performance Analysis
• Latency: 1-2ms (p99)
• Complexity: O(1) hash lookup
• Network hops: 1 (direct to node)
• Disk seeks: 1 (single partition)
• Data returned: Single row (small)
WRITE Performance:
• Latency: 1-2ms (p99)
• Complexity: O(1) direct insert
• No clustering overhead
• Memtable write: Fast
• Distribution: Excellent (spread across nodes)
THROUGHPUT:
• Single node: 50-100K ops/sec
• 10 node cluster: 500K-1M ops/sec
• Linear scaling with nodes
Why So Fast:
- Hash-Based Lookup: O(1) to find partition
- No Clustering Scan: No iteration needed
- Small Data Size: Fits in cache
- Even Distribution: All nodes utilized
- Simple Structure: No complex sorting
✅ Strengths
- Predictable 1-2ms latency
- Linear scalability
- Even load distribution
- Simple to reason about
- Cache-friendly
- No hotspots
⚠️ Limitations
- Can't query related data together
- No time-range queries
- More network hops for batch reads
- Can't leverage DESC clustering
- Higher coordinator overhead for bulk ops
🎯 Production Use Cases
When major companies use skinny rows.
💳 Stripe: Customer Records
customer_id UUID PRIMARY KEY,
email TEXT,
name TEXT,
payment_method_id TEXT,
created_at TIMESTAMP,
metadata MAP
);
Why Skinny Rows:
- Customer lookup by ID (API calls)
- Simple customer updates
- No history in main table
- Fast API response (< 10ms)
- 100M+ customers, distributed evenly
🎮 Epic Games: Player Profiles
Pattern: 1 player = 1 row with current state
- Fast player login (fetch profile)
- Real-time updates (level, XP)
- Distributed across all nodes
- 300M+ players, 1-2ms lookups
- Separate tables for match history (wide rows)
🏪 Amazon: Product Catalog
Pattern: 1 product = 1 row with current details
- Product page loads (< 5ms)
- Price/stock updates
- 100M+ products distributed
- Separate tables for reviews (wide rows)
- Separate tables for price history (wide rows)
🌍 Complete Production Examples
Session Management
WITH default_time_to_live = 3600
Scale: Billions of sessions, auto-expire after 1 hour
Shopping Carts
items LIST
Scale: 100M+ active carts, 1-2ms updates
API Keys
rate_limit_remaining INT
Scale: Millions of keys, sub-ms lookups for auth
✅ Best Practices for Skinny Rows
DO This
- Use for Entities: User, product, order
- High Cardinality Keys: UUIDs, unique IDs
- Enable TTL: For temporary data
- Keep Rows Small: < 1KB ideal
- Use Collections: For related data
- Leverage Caching: Frequently accessed
- Monitor Distribution: Even spread
DON'T Do This
- Use for Time-Series: Use wide rows instead
- Store Large BLOBs: > 1MB per row
- Scan Tables: No WHERE clause
- Use Secondary Indexes: If avoidable
- Over-Update: Batch when possible
- Forget Replication: Plan for RF
⚠️ Skinny Row Anti-Patterns
❌ Anti-Pattern: Using Skinny Rows for Time-Series
BAD: Storing each event as separate partition
Problem: Can't query "all events for user" efficiently
GOOD: Use wide rows for related events
Benefit: All user events in one partition
💼 Interview Questions & Expert Answers
Answer: Skinny rows are partitions containing exactly ONE row (or very few). Use for: entity lookup (users, products), key-value stores (sessions, cache), configuration, and any data where you fetch by single ID with no history needed. Benefits: O(1) lookups, 1-2ms latency, even distribution across nodes.
Answer: Skinny: PRIMARY KEY (id) → 1 partition = 1 row. Use for entities. Wide: PRIMARY KEY (id, timestamp) → 1 partition = many rows. Use for time-series. Skinny = dictionary lookup, Wide = spreadsheet.
Answer: O(1) hash-based partition lookup, no clustering column scanning, single row read, small data size fits in cache, predictable 1-2ms p99 latency regardless of cluster size.
Answer: Absolutely! PRIMARY KEY (session_id) WITH default_time_to_live = 3600. Each session = 1 row, auto-expires, 1-2ms lookups, scales to billions. Used by major web applications for distributed session management.
Answer: Read/write: 1-2ms p99. Throughput: 50-100K ops/sec per node. Scales linearly. Even distribution prevents hotspots. Cache-friendly for frequently accessed data. Predictable performance.
🎓 Chapter Summary: Skinny Row Mastery
You now master Cassandra's narrow partition pattern!
Key Concepts:
- Definition: 1 partition = 1 row (or very few)
- Pattern: PRIMARY KEY (id) with no clustering
- Use Cases: Entities, sessions, configs, key-value
- Performance: O(1) lookups, 1-2ms latency
- Distribution: Even spread across nodes
vs Wide Rows:
Skinny: Entity lookup, current state, key-value
Wide: Time-series, history, related data
🚀 You can now choose the right pattern for your data!
Responsive Ad