Section 2: Data Modeling Fundamentals

Cassandra Anti-Patterns

Learn what NOT to do! Master the 10 most common mistakes that destroy performance, cause outages, and waste resources. Real disasters, real fixes, production examples!

⚠️ Why Anti-Patterns Matter

Imagine building a beautiful house on quicksand. That's what happens when you use Cassandra anti-patterns. Your application works great with 100 users... then collapses with 10,000.

💥 Real Production Disasters

  • Company A: One unbounded partition crashed entire cluster (3am outage)
  • Company B: Query without partition key = $50K/month AWS bill
  • Company C: Hot partition = 1 node handling 90% of traffic
  • Company D: Tombstone accumulation = reads taking 30+ seconds

Good News: All these disasters are 100% preventable! Learn these 10 anti-patterns and you'll avoid 90% of Cassandra production issues.

1 🔥 CRITICAL

Unbounded Partitions (The Cluster Killer)

💥 The Disaster: E-Commerce Product Reviews

What They Did:

CREATE TABLE reviews (
  product_id UUID,
  review_id TIMEUUID,
  rating INT,
  comment TEXT,
  PRIMARY KEY (product_id, review_id)
);

// Looks innocent... but it's a BOMB! 💣

What Went Wrong:

Popular product got 5 MILLION reviews. All 5M rows in ONE partition!

  • Partition size: 500MB (limit is 100MB!)
  • Read latency: 30+ seconds
  • Compaction: Never finished
  • Result: Node crashed at 3am 💥
Unbounded Partition Growth Over Time Month 1 Month 3 Month 6 Month 12 10K rows 1MB ✅ 100K rows 10MB ⚠️ 1M rows 100MB 🔥 5M rows 500MB 💥 CRASHED! 100MB LIMIT

❌ WRONG: Unbounded

PRIMARY KEY (product_id, review_id)

// All reviews for product
// in ONE partition!
// Can grow FOREVER!

✅ RIGHT: Bucketed

PRIMARY KEY ((product_id, year_month), review_id)

// Split by month!
// Each month = new partition
// Bounded growth! ✅

How to Detect & Fix

Detection:

nodetool cfstats keyspace.table_name

Look for:
- SSTable count: > 100 (lots of compactions)
- Partition size: > 100MB
- Read latency: > 100ms

Fix:

  1. Add bucketing column (date, hash, etc.)
  2. Create new table with bucket in partition key
  3. Migrate data to new table
  4. Update application queries to include bucket
2 💸 EXPENSIVE

Querying Without Partition Key

💸 The Disaster: "Find All Premium Users"

What They Did:

SELECT * FROM users
WHERE subscription_type = 'premium'
ALLOW FILTERING;

// ALLOW FILTERING = "Please scan EVERYTHING"

What Happened:

  • 10 million user records
  • Cassandra scanned ALL 10M rows across ALL nodes
  • Query took 45 seconds
  • AWS bill jumped from $2K → $50K/month
  • CPU at 100% constantly
Query Without Partition Key = Full Cluster Scan Node 1 Scan 3M users Node 2 Scan 3.5M users Node 3 Scan 3.5M users Result: 45 seconds, 100% CPU, $$$$$ Coordinator merges results from ALL nodes

❌ WRONG: No Partition Key

SELECT * FROM users
WHERE subscription_type = 'premium'
ALLOW FILTERING;

// Full cluster scan!
// All 10M users checked

✅ RIGHT: Separate Table

CREATE TABLE users_by_subscription (
  subscription_type TEXT,
  user_id UUID,
  PRIMARY KEY (subscription_type, user_id)
);

SELECT * FROM users_by_subscription
WHERE subscription_type = 'premium';

⚠️ The Golden Rule

NEVER use ALLOW FILTERING in production. If you need it, create a new table with the correct partition key!

3 🔥 HOT SPOT

Hot Partitions (The Bottleneck)

🔥 The Disaster: Celebrity Twitter Account

Schema:

PRIMARY KEY (user_id, tweet_id)

// Celebrity with 50M followers
// Every tweet = 50M writes to SAME partition!

Result:

  • One node handling 90% of writes
  • That node: 100% CPU, overheating
  • Other nodes: Idle (wasted capacity)
  • Timeline queries: Timing out

❌ Wrong Approach

Single partition = bottleneck

  • 1 node overloaded
  • Can't scale horizontally
  • Celebrity effect amplified

✅ Right Approach

Fan-out + bucketing

  • Distribute across multiple partitions
  • Use hash bucketing
  • Application-level aggregation
4 🐌 SLOW

Secondary Index Abuse

⚠️ The Truth About Secondary Indexes

What developers think: "It's like SQL indexes!"

Reality: Secondary indexes in Cassandra are SLOW and should rarely be used.

Why They're Bad:

  • Query ALL nodes (can't hash to one node)
  • High cardinality = slow
  • Low cardinality = hot partitions in index
  • No good use case! Create proper table instead

❌ Secondary Index

CREATE INDEX ON users (email);

SELECT * FROM users WHERE email = ?;

// Queries ALL nodes = SLOW

✅ Dedicated Table

CREATE TABLE users_by_email (
  email TEXT PRIMARY KEY,
  user_id UUID
);

// Direct partition lookup = FAST
5 ⏱️ INEFFICIENT

Reading Before Writing

The Anti-Pattern

Reading a value before writing it back (like updating a counter manually).

Why It's Bad:

  • 2x network round trips
  • Race conditions
  • Wasted read capacity

Solution: Use lightweight transactions (CAS) or counter columns.

6

Logged Batch Misuse

Using LOGGED BATCH for performance (it actually SLOWS writes). Only use batches for atomicity across partitions, not performance.

7

Counter Column Overuse

Using counters for everything. Counters are eventual consistent and have performance overhead. Use only when you truly need distributed counting.

8

Large Row Anti-Pattern

Storing large BLOBs (images, videos) directly in Cassandra. Keep rows under 10MB. Store large files in object storage (S3) and reference IDs in Cassandra.

9

Collection Abuse

Storing thousands of items in a SET/LIST/MAP. Collections are stored in single column, loaded entirely on read. Keep collections under 100 items.

10

Tombstone Accumulation

The Silent Killer

Deleting data frequently without proper TTL or gc_grace_seconds tuning. Tombstones pile up, reads become SLOW.

Solution:

  • Use TTL for time-series data
  • Tune gc_grace_seconds appropriately
  • Monitor tombstone warnings
  • Consider time-bucketing to drop whole partitions

🔍 How to Detect Anti-Patterns

Monitoring Tools

  • nodetool cfstats - Partition sizes
  • nodetool tablehistograms - Latency distribution
  • nodetool tpstats - Thread pool stats
  • Cassandra logs - Tombstone warnings

Red Flags

  • Read latency > 100ms
  • Partition size > 100MB
  • Tombstone warnings in logs
  • Compaction always running

Application Signs

  • Queries using ALLOW FILTERING
  • Timeouts in production
  • One node CPU at 100%
  • Increasing AWS costs

✅ The Ultimate Checklist

Before Going to Production:

  • ✅ All partitions bounded (< 100MB)
  • ✅ Every query has partition key
  • ✅ No ALLOW FILTERING in code
  • ✅ No secondary indexes (use separate tables)
  • ✅ Hot partitions identified and mitigated
  • ✅ TTL set for time-series data
  • ✅ Collections kept small (< 100 items)
  • ✅ No large BLOBs stored directly
  • ✅ Monitoring alerts configured
  • ✅ Load tested with production data volumes

Follow these rules and you'll avoid 90% of Cassandra production disasters! 🎉

Advertisement

Google AdSense - Responsive Ad Unit