Practice Debugging
Fix real-world Cassandra errors and performance issues!
- Query works but takes 30+ seconds
- CPU usage spikes to 100% during query
- Affects all nodes in cluster
- Query scans entire products table
ALLOW FILTERING scans ALL partitions, causing full table scan. The problem is that 'price' is not part of the PRIMARY KEY, so Cassandra can't efficiently filter. Think about how to redesign the data model to support this query pattern.
🔍 Root Cause
ALLOW FILTERING forces a full table scan across all partitions. With a large products table, this means reading millions of rows and filtering them in memory. This is why the query is slow.
🛠️ Fix: Create Price Range Table
💡 Alternative: Secondary Index (with caution)
Warning: Secondary indexes have limitations. Better to design proper tables.
- Never use ALLOW FILTERING in production queries
- Design tables based on query patterns (query-first design)
- Use bucketing for range queries
- Avoid secondary indexes on high-cardinality columns
- Compaction warnings for large partitions
- Read queries timing out
- Node crashes during compaction
- Excessive disk usage on some nodes
The problem is that sensor_id is the only partition key. All readings for a sensor go to one partition forever. For time-series data, you need to "bucket" data by time to keep partitions bounded. Think: how can you split data by time periods?
🔍 Root Cause
Using only sensor_id as partition key creates unbounded partitions. One sensor collecting data every second creates ~31 million rows per year in a single partition, far exceeding the 100MB recommendation.
🛠️ Fix: Time Bucketing
📝 Insert Example
🔍 Query Pattern
- Always use time bucketing for time-series data
- Keep partitions under 100MB (ideally under 10MB)
- Choose bucket size based on write rate
- Use TimeWindowCompactionStrategy (TWCS) for time-series
- Monitor partition sizes with nodetool tablestats
- Read performance degrades over time
- Tombstone warnings in logs
- High disk I/O during reads
- Query latency increases linearly with time
Every DELETE creates a tombstone marker. Cassandra must read and skip all tombstones during queries. High delete rates + long gc_grace_seconds = tombstone buildup. Consider: do you really need to DELETE, or can you use TTL or a different pattern?
🔍 Root Cause
Frequent DELETEs create tombstones. During reads, Cassandra must scan through all tombstones (120,000!) to find live data (15,000 rows). Tombstones aren't removed until gc_grace_seconds expires AND compaction runs.
🛠️ Fix 1: Use TTL Instead of DELETE
Why it works: TTL still creates tombstones, but they're more predictable and manageable.
🛠️ Fix 2: Reduce gc_grace_seconds
Trade-off: Must run repairs more frequently to prevent zombie data.
🛠️ Fix 3: Redesign Without Deletes
🔧 Immediate Fix: Manual Compaction
Forces compaction to remove eligible tombstones now.
- Minimize DELETEs; use TTL when possible
- Lower gc_grace_seconds for high-churn tables
- Monitor tombstone warnings
- Consider redesigning to avoid deletes
- Run regular compactions on high-delete tables
- Write timeouts during high traffic
- Works fine during low traffic
- Some nodes have high write latency
- Client sees "2 nodes required but only 1 responded"
QUORUM requires 2 out of 3 nodes to respond within timeout. If nodes are overloaded or have high GC pauses, they can't respond in time. Check node performance metrics and consider either fixing slow nodes or adjusting consistency level.
🔍 Root Cause
QUORUM (2/3 nodes) must respond within 2 seconds. During high load, one or more nodes can't respond fast enough. Common causes: high GC pauses, overloaded nodes, network issues, or hardware problems.
🔍 Step 1: Diagnose Node Performance
🛠️ Fix 1: Increase Timeout
Note: This is a temporary fix. Find root cause!
🛠️ Fix 2: Lower Consistency Level
🛠️ Fix 3: Tune JVM Heap
🛠️ Fix 4: Add More Nodes
If nodes are genuinely overloaded, scale horizontally by adding more nodes to distribute load.
- Monitor node performance metrics continuously
- Set appropriate timeouts based on P99 latency
- Use appropriate consistency levels per use case
- Tune JVM for consistent GC pauses
- Scale cluster before hitting resource limits
- One node has 5x more data than others
- That node is slower for reads/writes
- Uneven "Owns" percentages
- Cluster is not utilizing resources evenly
Unbalanced clusters usually indicate either: (1) poor partition key choice creating "hot" partitions, or (2) incorrect token assignment. Check if certain partition keys are storing way more data than others, or if initial_token was manually misconfigured.
🔍 Root Cause Analysis
Two main causes: (1) Hot partitions - one or more partition keys have massively more data, or (2) Bad token distribution - tokens weren't properly assigned when nodes joined.
🔍 Diagnose: Check for Hot Partitions
🛠️ Fix 1: If Hot Partitions (Data Model Issue)
🛠️ Fix 2: If Bad Token Distribution
🛠️ Fix 3: Run Cleanup After Decommission
- Use vnodes (num_tokens: 256) for automatic balancing
- Choose partition keys with high cardinality
- Monitor "nodetool status" regularly
- Avoid manual token assignment unless necessary
- Test data distribution in staging first
Responsive Ad