Data Modeling Guide
Master Cassandra data modeling - principles, patterns, anti-patterns, and real-world examples!
🎯 Core Principle
Query-first, not entity-first
🔑 Key Design
Partition + Clustering keys
📊 Denormalization
Duplicate data for queries
⚡ Performance
One partition = one query
🎯 Core Modeling Principles
The Golden Rule
Design tables based on QUERIES, not entities!
Cassandra modeling is fundamentally different from relational databases. You model for how you READ, not how data is structured.
Cassandra vs Relational Modeling
❌ Relational (SQL)
- 📐 Normalize data (3NF)
- 🔗 JOIN tables at query time
- 🎯 Design entities first
- 📊 Query flexibility (any JOIN)
- ⚠️ Slow for large datasets
✅ Cassandra (NoSQL)
- 📦 Denormalize data
- 🚫 NO JOINs (pre-joined)
- 🔍 Design queries first
- ⚡ Fast, fixed queries only
- ✅ Scalable to petabytes
The Three Fundamental Rules
1️⃣ Spread Data Evenly Across Cluster
Goal: Avoid "hot partitions" where one node handles all traffic.
How: Choose partition keys with high cardinality (many unique values).
Example: user_id ✅ (millions) vs country ❌ (hundreds)
2️⃣ Minimize Partition Reads
Goal: One query = one partition lookup (or very few).
How: Include partition key in WHERE clause.
Example: WHERE user_id = '123' ✅ vs WHERE email = 'alice@...' ❌ (full scan)
3️⃣ Minimize Partition Size
Goal: Keep partitions < 100MB (ideally < 10MB).
How: Use composite partition keys to split large datasets.
Example: (user_id, date) instead of just user_id for time-series
Data Duplication is OK!
In Cassandra, disk is cheap, but queries are expensive. Duplicate data across multiple tables optimized for different queries. This is normal and expected!
📋 The Modeling Process
Step-by-Step Methodology
Step 1: Define Application Queries
List EVERY query your application needs. Be specific!
Step 2: Identify Access Patterns
For each query, determine:
- What is known? (partition key)
- What order? (clustering key)
- How much data? (partition size)
- How often? (read/write ratio)
Step 3: Design Tables for Queries
Create ONE table per query pattern:
Step 4: Optimize & Validate
- ✅ Each query touches one partition?
- ✅ Partitions < 100MB?
- ✅ Data distributed evenly?
- ✅ No hot partitions?
- ✅ Write patterns sustainable?
🔑 Primary Key Design
Primary Key Anatomy
Types of Primary Keys
Simple Partition Key
Composite Partition Key
Compound Primary Key (Partition + Clustering)
Partition Key Selection Guide
✅ Common Design Patterns
Pattern 1: One-to-Many Relationship
User has many posts
Pattern 2: Time-Series Data
Sensor readings over time
Pattern 3: Multiple Access Patterns (Duplication)
Blog posts by author AND by category
Pattern 4: Bucketing Large Partitions
Prevent unlimited partition growth
Pattern 5: Wide Rows (Skinny vs Fat Partitions)
❌ Anti-Patterns to Avoid
Anti-Pattern 1: Using Secondary Indexes for Everything
Anti-Pattern 2: Unbounded Partition Growth
Anti-Pattern 3: Low-Cardinality Partition Keys
Anti-Pattern 4: Treating Cassandra Like SQL
Anti-Pattern 5: Large Partitions
Rule: Partitions should be < 100MB (ideally < 10MB)
Signs of large partitions:
- Slow queries on specific keys
- Timeouts on certain partitions
- Uneven disk usage across nodes
- GC pressure on specific nodes
Solution: Add bucketing column to partition key!
💼 Real-World Examples
Example 1: E-Commerce Order System
Example 2: Social Media Feed
Example 3: IoT Sensor Data
🏆 Best Practices Summary
✅ DO These
- Model queries first
- Denormalize data
- Duplicate tables per query
- Use high-cardinality partition keys
- Keep partitions < 100MB
- Bucket time-series data
- Include partition key in WHERE
- Use clustering for sort order
- Test with production data volume
- Monitor partition sizes
❌ DON'T Do These
- Try to normalize like SQL
- Use JOINs (they don't exist!)
- Use ALLOW FILTERING in production
- Create unbounded partitions
- Use low-cardinality partition keys
- Query without partition key
- Overuse secondary indexes
- Ignore partition size warnings
- Model entities before queries
- Expect flexible querying
The Modeling Mindset
Think in queries, not entities.
Embrace data duplication. Disk is cheap, queries are expensive.
Design for scalability. How will this work with 100TB of data?
One query = one partition lookup. This is the golden rule.
📚 Recommended Reading Order
- Start: Understand your queries (application requirements)
- Learn: Partition key selection (data distribution)
- Practice: One-to-many patterns (most common)
- Master: Time-series bucketing (prevents growth)
- Advanced: Multi-table duplication (query optimization)
Responsive Ad