Optimization

Performance Tuning

Optimize Cassandra for maximum speed and efficiency!

Scenario #1
High Impact
🐌 Slow Read Queries (P95 = 200ms)
Your user dashboard queries are slow. Users complain about lag when viewing their order history. P95 latency is 200ms, but should be under 50ms.
Current Metrics
200ms
P95 Read Latency
15ms
Target Latency
5000/s
Read Requests
82%
Cache Hit Rate
Current Query
SELECT * FROM orders_by_user WHERE user_id = ? LIMIT 100; -- Table has 20+ columns including large text fields -- Fetching unnecessary data!
🔍 Performance Analysis
  • Query selects ALL columns (SELECT *)
  • Large TEXT columns (order_details, shipping_address)
  • Low cache hit rate (82%)
  • No prepared statements used
⚡ Performance Optimizations

🎯 Optimization 1: Select Only Needed Columns

-- Before: SELECT * fetches all 20+ columns SELECT * FROM orders_by_user ... -- After: Select only what you need SELECT order_id, order_date, total_amount, status FROM orders_by_user WHERE user_id = ? LIMIT 100;

Impact: Reduces data transfer by 60-70%, improves cache efficiency.

🎯 Optimization 2: Use Prepared Statements

// Python - prepare once, execute many prepared = session.prepare(""" SELECT order_id, order_date, total_amount, status FROM orders_by_user WHERE user_id = ? """) # Reuse prepared statement (faster!) result = session.execute(prepared, [user_id])

Impact: 10-30% latency improvement from query plan caching.

🎯 Optimization 3: Increase Row Cache

-- Enable row cache for hot tables ALTER TABLE orders_by_user WITH caching = { 'keys': 'ALL', 'rows_per_partition': '100' };

Impact: Cache hit rate improves from 82% → 95%.

🎯 Optimization 4: Add Indexes to Application

// Create index on application side const orderCache = new Map(); function getOrders(userId) { if (orderCache.has(userId)) { return orderCache.get(userId); } const orders = cassandra.execute(query, [userId]); orderCache.set(userId, orders); return orders; }

Impact: Eliminates DB calls for frequently accessed data.

MetricBeforeAfterImprovement
P95 Latency200ms12ms94% faster
Data Transfer50KB/query8KB/query84% less
Cache Hit Rate82%95%+13%
Throughput5000 req/s15000 req/s3x higher
✅ Key Takeaways
  • SELECT only columns you need (avoid SELECT *)
  • Always use prepared statements for repeated queries
  • Enable row cache for frequently read partitions
  • Add application-level caching for hot data
  • Monitor P95/P99 latencies, not just averages
Scenario #2
High Impact
📝 Low Write Throughput (500 writes/sec)
Your IoT application needs to ingest 10,000 sensor readings per second, but you're only achieving 500 writes/sec. The application is making individual INSERT statements for each reading.
Current Metrics
500/s
Current Writes/Sec
10000/s
Target Writes/Sec
50ms
Write Latency
98%
CPU Usage
Current Code (Slow)
# Processing one at a time - BAD! for reading in sensor_readings: session.execute(""" INSERT INTO sensor_data (sensor_id, timestamp, value) VALUES (?, ?, ?) """, (reading.id, reading.time, reading.value)) # Network roundtrip for EACH insert = slow!
🔍 Performance Analysis
  • Individual writes - no batching
  • Network latency (5ms) × 10,000 = 50 seconds!
  • No concurrent writes
  • Synchronous execution blocks on each write
⚡ Performance Optimizations

🎯 Optimization 1: Batch Statements

from cassandra.query import BatchStatement batch = BatchStatement() prepared = session.prepare(""" INSERT INTO sensor_data (sensor_id, timestamp, value) VALUES (?, ?, ?) """) # Batch 100 inserts together for reading in readings[:100]: batch.add(prepared, (reading.id, reading.time, reading.value)) session.execute(batch) # Single network call!

Impact: 100 writes in 1 roundtrip vs 100 roundtrips. 20x faster!

🎯 Optimization 2: Async + Concurrent Execution

from cassandra.concurrent import execute_concurrent_with_args prepared = session.prepare(""" INSERT INTO sensor_data (sensor_id, timestamp, value) VALUES (?, ?, ?) """) # Execute 100 queries concurrently parameters = [(r.id, r.time, r.value) for r in readings] execute_concurrent_with_args( session, prepared, parameters, concurrency=100 )

Impact: 100 concurrent requests = ~100x faster than serial.

🎯 Optimization 3: Tune Write Settings

# cassandra.yaml tuning concurrent_writes: 128 # Default: 32, increase for write-heavy memtable_flush_writers: 4 # Parallel flushes commitlog_sync: periodic # vs batch (faster) commitlog_sync_period_in_ms: 10 # Consistency level tuning consistency_level = ONE # Faster writes (vs QUORUM)

Impact: Increases write capacity by 2-3x.

🎯 Optimization 4: Unlogged Batches

-- For same partition, use UNLOGGED BEGIN UNLOGGED BATCH INSERT INTO sensor_data (...) VALUES (...); INSERT INTO sensor_data (...) VALUES (...); INSERT INTO sensor_data (...) VALUES (...); APPLY BATCH; -- Skip distributed log = faster -- Only use if all inserts go to same partition!

Impact: 30-50% faster than logged batches.

MetricBeforeAfterImprovement
Write Throughput500/s12,000/s24x higher
Write Latency50ms5ms10x faster
CPU Usage98%45%53% less
Network Calls10,00010099% fewer
✅ Key Takeaways
  • Batch 50-100 writes together (not too large!)
  • Use async + concurrent execution for parallel writes
  • Tune concurrent_writes based on workload
  • Use consistency ONE for writes when acceptable
  • UNLOGGED batches for same-partition writes only
Scenario #3
High Impact
💾 High GC Pauses (5+ second pauses)
Your nodes experience 5-10 second GC pauses during peak hours, causing timeouts and impacting availability. Heap is 16GB with frequent full GCs.
Current Metrics
5-10s
GC Pause Time
16GB
Heap Size
92%
Heap Usage
50+
Full GCs/hour
🔍 Performance Analysis
  • Heap constantly at 90%+ usage
  • Frequent full GCs (stop-the-world)
  • Memtables not flushing fast enough
  • Using CMS GC (old generation collector)
⚡ Performance Optimizations

🎯 Optimization 1: Switch to G1GC

# jvm.options - Before (CMS) -XX:+UseConcMarkSweepGC -XX:+CMSParallelRemarkEnabled # After (G1GC - better for large heaps) -XX:+UseG1GC -XX:G1RSetUpdatingPauseTimePercent=5 -XX:MaxGCPauseMillis=500 -XX:InitiatingHeapOccupancyPercent=70

Impact: GC pauses drop from 5-10s → 200-500ms.

🎯 Optimization 2: Right-size Heap

# cassandra-env.sh # Before: Too large (16GB) MAX_HEAP_SIZE="16G" # After: Optimal for G1GC (8-12GB) MAX_HEAP_SIZE="10G" HEAP_NEWSIZE="2G" # Rule: 1/4 to 1/2 of system RAM, max 12GB

Why: Larger heaps = longer GC pauses. Let OS cache do the work!

🎯 Optimization 3: Tune Memtable Settings

# cassandra.yaml memtable_heap_space_in_mb: 2048 memtable_offheap_space_in_mb: 2048 memtable_cleanup_threshold: 0.6 # Flush more aggressively memtable_flush_writers: 4 concurrent_compactors: 4

Impact: Reduces heap pressure by flushing earlier.

🎯 Optimization 4: Reduce Cache Sizes

# cassandra.yaml key_cache_size_in_mb: 100 # Was: 500 (too large) row_cache_size_in_mb: 0 # Disable if causing pressure # Let OS page cache handle caching

Why: OS cache is more efficient than JVM heap cache.

MetricBeforeAfterImprovement
GC Pause Time5-10s200-500ms95% faster
Heap Usage92%65%Healthier
Full GCs/hour50+2-394% fewer
Availability95%99.9%More reliable
✅ Key Takeaways
  • Use G1GC for heaps > 6GB (better pause times)
  • Keep heap between 8-12GB maximum
  • Tune memtable settings to reduce heap pressure
  • Monitor GC logs with nodetool gcstats
  • Let OS page cache handle caching when possible
Scenario #4
Medium Impact
🌐 High Network Transfer (500MB/sec)
Your cross-datacenter replication is saturating network links. Transferring 500MB/sec between datacenters, causing high bandwidth costs and replication lag.
Current Metrics
500MB/s
Network Transfer
$2500/mo
Bandwidth Cost
30s
Replication Lag
None
Compression
🔍 Performance Analysis
  • No compression enabled on internode traffic
  • Using QUORUM across datacenters (slow!)
  • Full hints transfer on node recovery
  • Inefficient compaction strategy
⚡ Performance Optimizations

🎯 Optimization 1: Enable Compression

# cassandra.yaml # Internode compression (DC to DC) internode_compression: dc # Client-server compression client_encryption_options: enabled: true compression: lz4 # SSTable compression ALTER TABLE users WITH compression = { 'sstable_compression': 'LZ4Compressor', 'chunk_length_kb': 64 };

Impact: 60-80% reduction in network transfer.

🎯 Optimization 2: Use LOCAL_QUORUM

// Before: QUORUM (waits for all DCs) consistency_level = QUORUM; // After: LOCAL_QUORUM (local DC only) consistency_level = LOCAL_QUORUM; // Async replication to other DCs // Much faster for writes!

Impact: 5-10x faster writes, lower latency.

🎯 Optimization 3: Tune Streaming Throttle

# cassandra.yaml # Limit streaming bandwidth stream_throughput_outbound_megabits_per_sec: 200 inter_dc_stream_throughput_outbound_megabits_per_sec: 50 # Prevents saturating network links

Why: Prevents repair/bootstrap from overwhelming network.

🎯 Optimization 4: Optimize Table Schema

-- Use appropriate data types (smaller = less transfer) CREATE TABLE events ( event_id uuid, timestamp bigint, -- vs timestamp (smaller) event_type tinyint, -- vs text (much smaller!) data blob -- compressed JSON );

Impact: 40-50% smaller data size.

MetricBeforeAfterImprovement
Network Transfer500MB/s120MB/s76% less
Bandwidth Cost$2500/mo$600/mo$1900 saved
Replication Lag30s5s83% faster
Write Latency150ms20ms87% faster
✅ Key Takeaways
  • Enable internode compression for DC-to-DC traffic
  • Use LOCAL_QUORUM instead of QUORUM for multi-DC
  • Throttle streaming to prevent network saturation
  • Choose efficient data types (tinyint vs text)
  • Enable SSTable compression (LZ4 is fast)
Scenario #5
Medium Impact
🔧 Compaction Causing Performance Spikes
Every few hours, your cluster experiences performance degradation. Read latency spikes to 500ms+ during these periods. Investigation shows major compactions running.
Current Metrics
500ms
Latency Spikes
STCS
Compaction Strategy
80%
Disk I/O During Spike
4 hours
Compaction Duration
🔍 Performance Analysis
  • Using SizeTieredCompactionStrategy (STCS)
  • Large compactions (50GB+) run periodically
  • High disk I/O during compaction
  • Read amplification = 20+ SSTables per read
⚡ Performance Optimizations

🎯 Optimization 1: Switch to Leveled Compaction

-- For read-heavy workloads ALTER TABLE users WITH compaction = { 'class': 'LeveledCompactionStrategy', 'sstable_size_in_mb': 160 }; -- Benefits: -- - Smaller, more frequent compactions -- - Lower read amplification (fewer SSTables) -- - More predictable performance

Impact: Read amplification drops from 20 → 5 SSTables.

🎯 Optimization 2: Time-Window for Time-Series

-- For time-series data (IoT, logs, metrics) ALTER TABLE sensor_readings WITH compaction = { 'class': 'TimeWindowCompactionStrategy', 'compaction_window_size': '24', 'compaction_window_unit': 'HOURS' }; -- Only compacts data within time windows -- Perfect for time-series with TTL

Impact: Eliminates massive compactions for time-series.

🎯 Optimization 3: Tune Compaction Throughput

# cassandra.yaml compaction_throughput_mb_per_sec: 64 # Default: 16 concurrent_compactors: 4 # Default: 1 # Or disable throttling entirely on fast SSDs: compaction_throughput_mb_per_sec: 0 # No limit

Why: Finish compactions faster, reduce duration.

🎯 Optimization 4: Adjust Compaction Strategy

-- Fine-tune STCS if keeping it ALTER TABLE events WITH compaction = { 'class': 'SizeTieredCompactionStrategy', 'min_threshold': 4, -- Compact sooner 'max_threshold': 32, -- More SSTables per compaction 'bucket_high': 1.5, -- Tighter buckets 'bucket_low': 0.5 };

Impact: More frequent, smaller compactions.

MetricBefore (STCS)After (LCS)Improvement
Read Latency Spike500ms50ms90% better
Read Amplification20 SSTables5 SSTables75% fewer
Compaction Duration4 hours15 min94% faster
Performance SpikesEvery 3 hoursNoneEliminated
✅ Key Takeaways
  • LCS for read-heavy workloads (lower read amplification)
  • TWCS for time-series data with TTL
  • Increase concurrent_compactors on multi-core systems
  • Monitor compaction with nodetool compactionstats
  • Choose strategy based on workload characteristics
Advertisement

Responsive Ad