Cassandra Anti-Patterns
Learn what NOT to do! Master the 10 most common mistakes that destroy performance, cause outages, and waste resources. Real disasters, real fixes, production examples!
⚠️ Why Anti-Patterns Matter
Imagine building a beautiful house on quicksand. That's what happens when you use Cassandra anti-patterns. Your application works great with 100 users... then collapses with 10,000.
💥 Real Production Disasters
- Company A: One unbounded partition crashed entire cluster (3am outage)
- Company B: Query without partition key = $50K/month AWS bill
- Company C: Hot partition = 1 node handling 90% of traffic
- Company D: Tombstone accumulation = reads taking 30+ seconds
Good News: All these disasters are 100% preventable! Learn these 10 anti-patterns and you'll avoid 90% of Cassandra production issues.
Unbounded Partitions (The Cluster Killer)
💥 The Disaster: E-Commerce Product Reviews
What They Did:
product_id UUID,
review_id TIMEUUID,
rating INT,
comment TEXT,
PRIMARY KEY (product_id, review_id)
);
// Looks innocent... but it's a BOMB! 💣
What Went Wrong:
Popular product got 5 MILLION reviews. All 5M rows in ONE partition!
- Partition size: 500MB (limit is 100MB!)
- Read latency: 30+ seconds
- Compaction: Never finished
- Result: Node crashed at 3am 💥
❌ WRONG: Unbounded
// All reviews for product
// in ONE partition!
// Can grow FOREVER!
✅ RIGHT: Bucketed
// Split by month!
// Each month = new partition
// Bounded growth! ✅
How to Detect & Fix
Detection:
Look for:
- SSTable count: > 100 (lots of compactions)
- Partition size: > 100MB
- Read latency: > 100ms
Fix:
- Add bucketing column (date, hash, etc.)
- Create new table with bucket in partition key
- Migrate data to new table
- Update application queries to include bucket
Querying Without Partition Key
💸 The Disaster: "Find All Premium Users"
What They Did:
WHERE subscription_type = 'premium'
ALLOW FILTERING;
// ALLOW FILTERING = "Please scan EVERYTHING"
What Happened:
- 10 million user records
- Cassandra scanned ALL 10M rows across ALL nodes
- Query took 45 seconds
- AWS bill jumped from $2K → $50K/month
- CPU at 100% constantly
❌ WRONG: No Partition Key
WHERE subscription_type = 'premium'
ALLOW FILTERING;
// Full cluster scan!
// All 10M users checked
✅ RIGHT: Separate Table
subscription_type TEXT,
user_id UUID,
PRIMARY KEY (subscription_type, user_id)
);
SELECT * FROM users_by_subscription
WHERE subscription_type = 'premium';
⚠️ The Golden Rule
NEVER use ALLOW FILTERING in production. If you need it, create a new table with the correct partition key!
Hot Partitions (The Bottleneck)
🔥 The Disaster: Celebrity Twitter Account
Schema:
// Celebrity with 50M followers
// Every tweet = 50M writes to SAME partition!
Result:
- One node handling 90% of writes
- That node: 100% CPU, overheating
- Other nodes: Idle (wasted capacity)
- Timeline queries: Timing out
❌ Wrong Approach
Single partition = bottleneck
- 1 node overloaded
- Can't scale horizontally
- Celebrity effect amplified
✅ Right Approach
Fan-out + bucketing
- Distribute across multiple partitions
- Use hash bucketing
- Application-level aggregation
Secondary Index Abuse
⚠️ The Truth About Secondary Indexes
What developers think: "It's like SQL indexes!"
Reality: Secondary indexes in Cassandra are SLOW and should rarely be used.
Why They're Bad:
- Query ALL nodes (can't hash to one node)
- High cardinality = slow
- Low cardinality = hot partitions in index
- No good use case! Create proper table instead
❌ Secondary Index
SELECT * FROM users WHERE email = ?;
// Queries ALL nodes = SLOW
✅ Dedicated Table
email TEXT PRIMARY KEY,
user_id UUID
);
// Direct partition lookup = FAST
Reading Before Writing
The Anti-Pattern
Reading a value before writing it back (like updating a counter manually).
Why It's Bad:
- 2x network round trips
- Race conditions
- Wasted read capacity
Solution: Use lightweight transactions (CAS) or counter columns.
Logged Batch Misuse
Using LOGGED BATCH for performance (it actually SLOWS writes). Only use batches for atomicity across partitions, not performance.
Counter Column Overuse
Using counters for everything. Counters are eventual consistent and have performance overhead. Use only when you truly need distributed counting.
Large Row Anti-Pattern
Storing large BLOBs (images, videos) directly in Cassandra. Keep rows under 10MB. Store large files in object storage (S3) and reference IDs in Cassandra.
Collection Abuse
Storing thousands of items in a SET/LIST/MAP. Collections are stored in single column, loaded entirely on read. Keep collections under 100 items.
Tombstone Accumulation
The Silent Killer
Deleting data frequently without proper TTL or gc_grace_seconds tuning. Tombstones pile up, reads become SLOW.
Solution:
- Use TTL for time-series data
- Tune gc_grace_seconds appropriately
- Monitor tombstone warnings
- Consider time-bucketing to drop whole partitions
🔍 How to Detect Anti-Patterns
Monitoring Tools
nodetool cfstats- Partition sizesnodetool tablehistograms- Latency distributionnodetool tpstats- Thread pool stats- Cassandra logs - Tombstone warnings
Red Flags
- Read latency > 100ms
- Partition size > 100MB
- Tombstone warnings in logs
- Compaction always running
Application Signs
- Queries using ALLOW FILTERING
- Timeouts in production
- One node CPU at 100%
- Increasing AWS costs
✅ The Ultimate Checklist
Before Going to Production:
- ✅ All partitions bounded (< 100MB)
- ✅ Every query has partition key
- ✅ No ALLOW FILTERING in code
- ✅ No secondary indexes (use separate tables)
- ✅ Hot partitions identified and mitigated
- ✅ TTL set for time-series data
- ✅ Collections kept small (< 100 items)
- ✅ No large BLOBs stored directly
- ✅ Monitoring alerts configured
- ✅ Load tested with production data volumes
Follow these rules and you'll avoid 90% of Cassandra production disasters! 🎉
Google AdSense - Responsive Ad Unit