Section 2: Data Modeling Fundamentals

Cassandra Skinny Rows

Master the narrow partition pattern! Learn key-value optimization, entity storage, when to use single-row partitions with animations, production examples, and expert-level design patterns.

📖 The Story: The Dictionary You Know

How do you find a word in a dictionary?

📖 The Dictionary Pattern

Look up "Cassandra":

Step 1: Go to "C" section
Step 2: Navigate to "Ca"
Step 3: Find "Cassandra"
Step 4: Read ONE entry

┌──────────────────────────────────┐
│ Word: Cassandra │
│ Definition: A distributed... │
│ Example: We use Cassandra... │
│ Synonyms: NoSQL, Database │
└──────────────────────────────────┘

ONE word = ONE entry = ONE lookup!

Key Characteristics:

  • Each word = Separate entry
  • One lookup = One result
  • Fast direct access
  • No need to scan multiple entries

📄 This is Skinny Rows in Cassandra!

PRIMARY KEY (user_id)

Each user_id = ONE partition = ONE row

user_id='alice' →
┌─────────────────────────────────┐
│ ONE PARTITION with ONE ROW │
│ user_id: alice │
│ name: Alice Smith │
│ email: alice@example.com │
│ created: 2024-01-15 │
└─────────────────────────────────┘

ONE key = ONE partition = ONE row = FAST!

Perfect For:

  • User profiles (1 user = 1 row)
  • Product catalog (1 product = 1 row)
  • Configuration (1 setting = 1 row)
  • Session store (1 session = 1 row)
  • Cache entries (1 key = 1 value)

📄 What Are Skinny Rows?

Understanding the narrow partition pattern.

Complete Definition

Skinny Row (Narrow Partition): A Cassandra partition containing exactly ONE row or very few rows (typically < 10).

Core Pattern:

  • One key → One partition
  • One partition → One (or very few) rows
  • No clustering columns (or used minimally)
  • Think: Key-value store, entity lookup, dictionary
Example Schema:
PRIMARY KEY (user_id)
↑
Partition Key ONLY (no clustering!)

Result:
user_id='alice' → ONE partition with ONE row
user_id='bob' → ONE partition with ONE row
user_id='carol' → ONE partition with ONE row

Each lookup returns exactly ONE row!

Skinny vs Wide Rows: Visual Comparison

📄

Skinny Rows

PRIMARY KEY (entity_id)

Structure:

  • Each entity = 1 partition
  • Each partition = 1 row
  • Direct key lookup
  • Like dictionary entry

Best For:

  • User profiles
  • Product details
  • Configuration
  • Session data
📊

Wide Rows

PRIMARY KEY (entity_id, timestamp)

Structure:

  • Each entity = 1 partition
  • Each partition = many rows
  • Range queries
  • Like spreadsheet

Best For:

  • Time-series data
  • Activity logs
  • Message history
  • Event streams

🔑 The Key-Value Pattern

Using Cassandra as a distributed key-value store.

Perfect Use Case: Session Store

CREATE TABLE user_sessions (
  session_id UUID PRIMARY KEY,
  user_id UUID,
  login_time TIMESTAMP,
  last_activity TIMESTAMP,
  ip_address TEXT,
  user_agent TEXT
) WITH default_time_to_live = 86400; -- 24 hour TTL

Pattern Characteristics:

  • Primary Key = Partition Key Only: No clustering columns
  • One Session = One Row: Direct lookup by session_id
  • TTL Enabled: Sessions automatically expire
  • Fast Reads: Single partition lookup (1-2ms)
  • Fast Writes: Simple insert/update

Query Examples:

-- Get session (1-2ms):
SELECT * FROM user_sessions
WHERE session_id = ?;

-- Update last activity:
UPDATE user_sessions
SET last_activity = ?
WHERE session_id = ?;

-- Delete session (logout):
DELETE FROM user_sessions
WHERE session_id = ?;

Performance Benefits

Why Skinny Rows Are Fast:

  • O(1) Hash Lookup: Direct partition location
  • No Scanning: Single row, no iteration
  • Cache-Friendly: Small, frequently accessed data
  • Predictable Performance: Consistent 1-2ms reads
  • Write Optimization: No clustering overhead
Performance Profile:
• Read latency: 1-2ms (single partition)
• Write latency: 1-2ms (single partition)
• Throughput: 100K+ ops/sec per node
• Scalability: Linear with cluster size

🗃️ Entity Storage Pattern

Storing complete entities in single rows.

👤

User Profiles

CREATE TABLE users (
  user_id UUID PRIMARY KEY,
  username TEXT,
  email TEXT,
  full_name TEXT,
  created_at TIMESTAMP,
  profile_image_url TEXT
);

Pattern:

  • 1 user = 1 row
  • All user data together
  • Fast user lookup
  • Simple updates
📦

Product Catalog

CREATE TABLE products (
  product_id UUID PRIMARY KEY,
  name TEXT,
  price DECIMAL,
  description TEXT,
  category TEXT,
  stock_count INT
);

Pattern:

  • 1 product = 1 row
  • Product details co-located
  • Fast product page load
  • Easy inventory updates
⚙️

Configuration

CREATE TABLE app_config (
  config_key TEXT PRIMARY KEY,
  config_value TEXT,
  updated_at TIMESTAMP,
  updated_by TEXT
);

Pattern:

  • 1 setting = 1 row
  • Fast config lookup
  • Easy updates
  • Version tracking

⚖️ Skinny vs Wide Rows: Decision Guide

When to use each pattern.

Use Skinny Rows When:

  • Entity Lookup: Get by ID/key
  • Key-Value Store: Session, cache
  • Single Object: User, product, config
  • No History Needed: Current state only
  • Simple Updates: Overwrite entire row
  • TTL Usage: Auto-expiring data
  • Distributed Writes: Even load
Examples:
• User profiles
• Product details
• Session data
• Configuration
• Cache entries

Use Wide Rows When:

  • Time-Series: Events over time
  • History/Logs: Activity tracking
  • Related Data: Multiple items per entity
  • Range Queries: Get by time range
  • Append-Only: Adding new events
  • Sequential Access: Read multiple rows
  • Co-Location: All data together
Examples:
• User activity logs
• IoT sensor data
• Message history
• Transaction history
• Photo albums

⚡ Performance Characteristics

Understanding skinny row performance.

Performance Analysis

READ Performance:
• Latency: 1-2ms (p99)
• Complexity: O(1) hash lookup
• Network hops: 1 (direct to node)
• Disk seeks: 1 (single partition)
• Data returned: Single row (small)

WRITE Performance:
• Latency: 1-2ms (p99)
• Complexity: O(1) direct insert
• No clustering overhead
• Memtable write: Fast
• Distribution: Excellent (spread across nodes)

THROUGHPUT:
• Single node: 50-100K ops/sec
• 10 node cluster: 500K-1M ops/sec
• Linear scaling with nodes

Why So Fast:

  • Hash-Based Lookup: O(1) to find partition
  • No Clustering Scan: No iteration needed
  • Small Data Size: Fits in cache
  • Even Distribution: All nodes utilized
  • Simple Structure: No complex sorting

✅ Strengths

  • Predictable 1-2ms latency
  • Linear scalability
  • Even load distribution
  • Simple to reason about
  • Cache-friendly
  • No hotspots

⚠️ Limitations

  • Can't query related data together
  • No time-range queries
  • More network hops for batch reads
  • Can't leverage DESC clustering
  • Higher coordinator overhead for bulk ops

🎯 Production Use Cases

When major companies use skinny rows.

💳 Stripe: Customer Records

CREATE TABLE customers (
  customer_id UUID PRIMARY KEY,
  email TEXT,
  name TEXT,
  payment_method_id TEXT,
  created_at TIMESTAMP,
  metadata MAP
);

Why Skinny Rows:

  • Customer lookup by ID (API calls)
  • Simple customer updates
  • No history in main table
  • Fast API response (< 10ms)
  • 100M+ customers, distributed evenly

🎮 Epic Games: Player Profiles

Pattern: 1 player = 1 row with current state

  • Fast player login (fetch profile)
  • Real-time updates (level, XP)
  • Distributed across all nodes
  • 300M+ players, 1-2ms lookups
  • Separate tables for match history (wide rows)

🏪 Amazon: Product Catalog

Pattern: 1 product = 1 row with current details

  • Product page loads (< 5ms)
  • Price/stock updates
  • 100M+ products distributed
  • Separate tables for reviews (wide rows)
  • Separate tables for price history (wide rows)

🌍 Complete Production Examples

Session Management

PRIMARY KEY (session_id)
WITH default_time_to_live = 3600

Scale: Billions of sessions, auto-expire after 1 hour

Shopping Carts

PRIMARY KEY (cart_id)
items LIST>

Scale: 100M+ active carts, 1-2ms updates

API Keys

PRIMARY KEY (api_key)
rate_limit_remaining INT

Scale: Millions of keys, sub-ms lookups for auth

✅ Best Practices for Skinny Rows

✅

DO This

  • Use for Entities: User, product, order
  • High Cardinality Keys: UUIDs, unique IDs
  • Enable TTL: For temporary data
  • Keep Rows Small: < 1KB ideal
  • Use Collections: For related data
  • Leverage Caching: Frequently accessed
  • Monitor Distribution: Even spread
❌

DON'T Do This

  • Use for Time-Series: Use wide rows instead
  • Store Large BLOBs: > 1MB per row
  • Scan Tables: No WHERE clause
  • Use Secondary Indexes: If avoidable
  • Over-Update: Batch when possible
  • Forget Replication: Plan for RF

⚠️ Skinny Row Anti-Patterns

❌ Anti-Pattern: Using Skinny Rows for Time-Series

BAD: Storing each event as separate partition

PRIMARY KEY (event_id) -- Each event = new partition!

Problem: Can't query "all events for user" efficiently

GOOD: Use wide rows for related events

PRIMARY KEY (user_id, event_timestamp) -- Wide row!

Benefit: All user events in one partition

💼 Interview Questions & Expert Answers

1 What are skinny rows and when should you use them?
Skinny rows have one partition with one row, perfect for entity lookup and key-value patterns.
▼

Answer: Skinny rows are partitions containing exactly ONE row (or very few). Use for: entity lookup (users, products), key-value stores (sessions, cache), configuration, and any data where you fetch by single ID with no history needed. Benefits: O(1) lookups, 1-2ms latency, even distribution across nodes.

2 What's the difference between skinny and wide rows?
Skinny: 1 partition = 1 row (entity). Wide: 1 partition = many rows (time-series, history).
▼

Answer: Skinny: PRIMARY KEY (id) → 1 partition = 1 row. Use for entities. Wide: PRIMARY KEY (id, timestamp) → 1 partition = many rows. Use for time-series. Skinny = dictionary lookup, Wide = spreadsheet.

3 Why are skinny rows fast?
O(1) hash lookup to partition, single row read, no clustering scan, cache-friendly small data.
▼

Answer: O(1) hash-based partition lookup, no clustering column scanning, single row read, small data size fits in cache, predictable 1-2ms p99 latency regardless of cluster size.

4 Can you use skinny rows for session storage?
Yes! Perfect pattern with TTL for auto-expiring sessions, 1-2ms lookups, billions of sessions supported.
▼

Answer: Absolutely! PRIMARY KEY (session_id) WITH default_time_to_live = 3600. Each session = 1 row, auto-expires, 1-2ms lookups, scales to billions. Used by major web applications for distributed session management.

5 What's the performance profile of skinny rows?
1-2ms reads/writes, 50-100K ops/sec per node, linear scaling, excellent distribution, no hotspots.
▼

Answer: Read/write: 1-2ms p99. Throughput: 50-100K ops/sec per node. Scales linearly. Even distribution prevents hotspots. Cache-friendly for frequently accessed data. Predictable performance.

🎓 Chapter Summary: Skinny Row Mastery

You now master Cassandra's narrow partition pattern!

Key Concepts:

  • Definition: 1 partition = 1 row (or very few)
  • Pattern: PRIMARY KEY (id) with no clustering
  • Use Cases: Entities, sessions, configs, key-value
  • Performance: O(1) lookups, 1-2ms latency
  • Distribution: Even spread across nodes

vs Wide Rows:

Skinny: Entity lookup, current state, key-value
Wide: Time-series, history, related data

🚀 You can now choose the right pattern for your data!

Advertisement

Responsive Ad