Performance Tuning
Optimize Cassandra for maximum speed and efficiency!
- Query selects ALL columns (SELECT *)
- Large TEXT columns (order_details, shipping_address)
- Low cache hit rate (82%)
- No prepared statements used
🎯 Optimization 1: Select Only Needed Columns
Impact: Reduces data transfer by 60-70%, improves cache efficiency.
🎯 Optimization 2: Use Prepared Statements
Impact: 10-30% latency improvement from query plan caching.
🎯 Optimization 3: Increase Row Cache
Impact: Cache hit rate improves from 82% → 95%.
🎯 Optimization 4: Add Indexes to Application
Impact: Eliminates DB calls for frequently accessed data.
| Metric | Before | After | Improvement |
|---|---|---|---|
| P95 Latency | 200ms | 12ms | 94% faster |
| Data Transfer | 50KB/query | 8KB/query | 84% less |
| Cache Hit Rate | 82% | 95% | +13% |
| Throughput | 5000 req/s | 15000 req/s | 3x higher |
- SELECT only columns you need (avoid SELECT *)
- Always use prepared statements for repeated queries
- Enable row cache for frequently read partitions
- Add application-level caching for hot data
- Monitor P95/P99 latencies, not just averages
- Individual writes - no batching
- Network latency (5ms) × 10,000 = 50 seconds!
- No concurrent writes
- Synchronous execution blocks on each write
🎯 Optimization 1: Batch Statements
Impact: 100 writes in 1 roundtrip vs 100 roundtrips. 20x faster!
🎯 Optimization 2: Async + Concurrent Execution
Impact: 100 concurrent requests = ~100x faster than serial.
🎯 Optimization 3: Tune Write Settings
Impact: Increases write capacity by 2-3x.
🎯 Optimization 4: Unlogged Batches
Impact: 30-50% faster than logged batches.
| Metric | Before | After | Improvement |
|---|---|---|---|
| Write Throughput | 500/s | 12,000/s | 24x higher |
| Write Latency | 50ms | 5ms | 10x faster |
| CPU Usage | 98% | 45% | 53% less |
| Network Calls | 10,000 | 100 | 99% fewer |
- Batch 50-100 writes together (not too large!)
- Use async + concurrent execution for parallel writes
- Tune concurrent_writes based on workload
- Use consistency ONE for writes when acceptable
- UNLOGGED batches for same-partition writes only
- Heap constantly at 90%+ usage
- Frequent full GCs (stop-the-world)
- Memtables not flushing fast enough
- Using CMS GC (old generation collector)
🎯 Optimization 1: Switch to G1GC
Impact: GC pauses drop from 5-10s → 200-500ms.
🎯 Optimization 2: Right-size Heap
Why: Larger heaps = longer GC pauses. Let OS cache do the work!
🎯 Optimization 3: Tune Memtable Settings
Impact: Reduces heap pressure by flushing earlier.
🎯 Optimization 4: Reduce Cache Sizes
Why: OS cache is more efficient than JVM heap cache.
| Metric | Before | After | Improvement |
|---|---|---|---|
| GC Pause Time | 5-10s | 200-500ms | 95% faster |
| Heap Usage | 92% | 65% | Healthier |
| Full GCs/hour | 50+ | 2-3 | 94% fewer |
| Availability | 95% | 99.9% | More reliable |
- Use G1GC for heaps > 6GB (better pause times)
- Keep heap between 8-12GB maximum
- Tune memtable settings to reduce heap pressure
- Monitor GC logs with nodetool gcstats
- Let OS page cache handle caching when possible
- No compression enabled on internode traffic
- Using QUORUM across datacenters (slow!)
- Full hints transfer on node recovery
- Inefficient compaction strategy
🎯 Optimization 1: Enable Compression
Impact: 60-80% reduction in network transfer.
🎯 Optimization 2: Use LOCAL_QUORUM
Impact: 5-10x faster writes, lower latency.
🎯 Optimization 3: Tune Streaming Throttle
Why: Prevents repair/bootstrap from overwhelming network.
🎯 Optimization 4: Optimize Table Schema
Impact: 40-50% smaller data size.
| Metric | Before | After | Improvement |
|---|---|---|---|
| Network Transfer | 500MB/s | 120MB/s | 76% less |
| Bandwidth Cost | $2500/mo | $600/mo | $1900 saved |
| Replication Lag | 30s | 5s | 83% faster |
| Write Latency | 150ms | 20ms | 87% faster |
- Enable internode compression for DC-to-DC traffic
- Use LOCAL_QUORUM instead of QUORUM for multi-DC
- Throttle streaming to prevent network saturation
- Choose efficient data types (tinyint vs text)
- Enable SSTable compression (LZ4 is fast)
- Using SizeTieredCompactionStrategy (STCS)
- Large compactions (50GB+) run periodically
- High disk I/O during compaction
- Read amplification = 20+ SSTables per read
🎯 Optimization 1: Switch to Leveled Compaction
Impact: Read amplification drops from 20 → 5 SSTables.
🎯 Optimization 2: Time-Window for Time-Series
Impact: Eliminates massive compactions for time-series.
🎯 Optimization 3: Tune Compaction Throughput
Why: Finish compactions faster, reduce duration.
🎯 Optimization 4: Adjust Compaction Strategy
Impact: More frequent, smaller compactions.
| Metric | Before (STCS) | After (LCS) | Improvement |
|---|---|---|---|
| Read Latency Spike | 500ms | 50ms | 90% better |
| Read Amplification | 20 SSTables | 5 SSTables | 75% fewer |
| Compaction Duration | 4 hours | 15 min | 94% faster |
| Performance Spikes | Every 3 hours | None | Eliminated |
- LCS for read-heavy workloads (lower read amplification)
- TWCS for time-series data with TTL
- Increase concurrent_compactors on multi-core systems
- Monitor compaction with nodetool compactionstats
- Choose strategy based on workload characteristics
Responsive Ad