Performance

Performance Overview in Cassandra

Turbo-charge your database! Performance overview tools show how efficiently Cassandra handles reads, writes, and scaling — helping you fine-tune your system like a pro mechanic for big-data engines.

📖 The Story: Netflix's 99.99% Uptime Challenge

Netflix serves 200+ million users globally with billions of requests daily. Their viewing history, recommendations, and user preferences run on Cassandra. Here's how they achieved legendary performance...

❌ The Problem: Performance Bottlenecks

Initial Challenges:

  • 🔥 Write Latency: 500ms+ during peak hours (Friday night traffic)
  • 📊 Read Latency: 200ms for user profiles (too slow!)
  • 💾 Disk I/O: 90%+ utilization causing slowdowns
  • 🌡️ GC Pauses: 5-10 second pauses causing timeouts
  • 🔥 Hot Partitions: Popular shows overloading specific nodes
  • 💥 Result: Buffering wheels, angry users, revenue loss!

✅ The Solution: Performance Optimization

What Netflix Did:

1. Data Modeling Optimization

  • ✅ Partitions kept < 100MB (bucketed by date)
  • ✅ Denormalized aggressively for read patterns
  • ✅ Used time-series bucketing for viewing history

2. Tuning Consistency Levels

  • ✅ Writes: LOCAL_ONE (fastest possible)
  • ✅ Reads: LOCAL_QUORUM (balance of speed + consistency)
  • ✅ Critical data: QUORUM only where needed

3. Compaction Strategy

  • ✅ TimeWindowCompactionStrategy for viewing history
  • ✅ LeveledCompactionStrategy for user profiles
  • ✅ Automated compaction during off-peak hours

4. JVM & Memory Tuning

  • ✅ Heap size: 8-16GB (not too large!)
  • ✅ G1GC collector for low pause times
  • ✅ Off-heap memory for caching

5. Hardware Optimization

  • ✅ SSDs for commit log (10x faster writes)
  • ✅ Separate SSDs for data (parallel I/O)
  • ✅ 32+ cores per node for parallelism

The Results:

  • ⚡ Write Latency: 500ms → 5ms (100x faster!)
  • ⚡ Read Latency: 200ms → 10ms (20x faster!)
  • 📊 Throughput: 10x increase in requests/sec
  • 💚 99.99% Uptime: Even during peak hours
  • 😊 Happy Users: No buffering, instant playback!
  • 💰 Cost Savings: 40% reduction in infrastructure

Netflix now handles 1 trillion requests per day with millisecond latency! 🎉

⚡ Cassandra Performance Fundamentals

Understand the core principles that make Cassandra fast (or slow)!

Why Cassandra is FAST

⚡

Write Performance

Blazing Fast Writes!

  • Sequential Writes: Append-only, no random seeks
  • Commit Log First: Write to log, then memtable
  • No Read-Before-Write: Just write!
  • Parallel Distribution: Writes go to multiple nodes
  • No Locks: Lock-free architecture

Result: 10,000-100,000 writes/sec per node! 🚀

📖

Read Performance

Fast Reads with Caching!

  • Bloom Filters: Skip SSTables without data
  • Key Cache: Skip partition index lookups
  • Row Cache: Cache hot rows in memory
  • Compression: Less disk I/O
  • SSD Optimization: Fast random reads

Result: Sub-millisecond reads for cached data! ⚡

📈

Scalability

Linear Scalability!

  • Masterless: No single bottleneck
  • Add Nodes: Throughput increases linearly
  • Distributed Hash: Even data distribution
  • Local Coordination: Minimize network hops
  • Peer-to-Peer: All nodes equal

Result: Scale to 1000s of nodes! 📈

Performance Metrics That Matter

1. Latency (Most Important!)

How fast does a single operation complete?

Operation Target Typical
Single-partition read < 5ms 1-3ms
Single write < 2ms 0.5-1ms
Multi-partition query < 20ms 10-15ms
P99 latency < 50ms 10-30ms

2. Throughput

How many operations per second?

  • Writes: 10,000-100,000 writes/sec per node
  • Reads: 50,000-200,000 reads/sec per node (with caching)
  • Cluster: Scales linearly with nodes
  • Example: 100-node cluster = 10M+ writes/sec!

3. Resource Utilization

Keep these in healthy ranges:

  • CPU: 50-70% average (spikes to 90% OK)
  • Memory: 60-80% heap usage
  • Disk I/O: < 80% utilization
  • Network: < 50% bandwidth
  • GC Pause: < 200ms per pause

🎨 Data Modeling for Performance

80% of performance comes from good data modeling!

🎯 The Golden Rule of Performance

"One query = One partition read"

The fastest Cassandra queries read from a SINGLE partition on a SINGLE node. Every additional partition or node adds latency.

Fast vs Slow Data Models

🐌

SLOW Models

  • Unbounded Partitions: > 100MB per partition
  • Multi-Partition Queries: Scatter-gather across nodes
  • ALLOW FILTERING: Full table scans
  • Large Collections: > 1000 elements
  • Low Cardinality Keys: Hot partitions
  • No Bucketing: Time-series without date buckets

Result: 100ms-10s latency 💥

⚡

FAST Models

  • Small Partitions: < 100MB each
  • Single-Partition Reads: One node, one partition
  • Partition Key in WHERE: Always!
  • Modest Collections: < 100 elements
  • High Cardinality Keys: Even distribution
  • Time Bucketing: Partition by date/hour

Result: 1-5ms latency ⚡

Performance Optimization Examples

-- ❌ SLOW: Unbounded partition (grows forever) CREATE TABLE user_events ( user_id UUID PRIMARY KEY, event_time TIMESTAMP, event_data TEXT ); -- After 1 year: 10M events in ONE partition! 💥 -- Query takes 2-5 seconds! -- ✅ FAST: Bounded partitions with bucketing CREATE TABLE user_events ( user_id UUID, event_date DATE, -- Bucket by day! event_time TIMESTAMP, event_data TEXT, PRIMARY KEY ((user_id, event_date), event_time) ) WITH CLUSTERING ORDER BY (event_time DESC); -- Now: Max ~100K events per partition (manageable!) -- Query takes 5-10ms! ⚡ ----------------------------------- -- ❌ SLOW: Query without partition key SELECT * FROM orders WHERE status = 'shipped' ALLOW FILTERING; -- Scans ALL nodes! Takes 30+ seconds! -- ✅ FAST: Denormalized table per access pattern CREATE TABLE orders_by_status ( status TEXT, order_date DATE, order_id UUID, customer_id UUID, total DECIMAL, PRIMARY KEY ((status, order_date), order_id) ); SELECT * FROM orders_by_status WHERE status = 'shipped' AND order_date = '2025-01-03' LIMIT 100; -- Reads ONE partition! Takes 5ms! ⚡

Partition Size Best Practices

  • Target: 10-100MB per partition (sweet spot)
  • Soft Limit: 100MB (start seeing slowdowns)
  • Hard Limit: 2GB (queries become very slow)
  • Row Count: 100,000 rows max per partition
  • Solution: Use bucketing (by date, by range, by hash)

📖 Read Path Optimization

Optimize every layer of the read path for blazing-fast queries!

Cassandra Read Path Client Query 1. Row Cache Check ⚡ Fastest (cached rows) 2. Bloom Filter Skip SSTables without data 3. Key Cache Skip partition index lookup 4. SSTable Read 🐌 Slowest (disk I/O) Performance Impact Row Cache: 0.1ms Bloom Filter: 0.5ms Key Cache: 1ms SSTable (SSD): 5-10ms SSTable (HDD): 50-100ms

Read Optimization Strategies

1. Enable Row Cache for Hot Data

Cache entire rows in memory for ultra-fast reads!

-- Enable row cache (for frequently accessed tables) ALTER TABLE user_profiles WITH caching = { 'keys': 'ALL', 'rows_per_partition': '100' -- Cache up to 100 rows }; -- Monitor cache hit rate nodetool info -- Look for: Row Cache Hit Rate: 95%

When to Use:

  • ✅ Frequently read data (hot rows)
  • ✅ Small rows (< 1KB)
  • ✅ Read-heavy workload (80%+ reads)
  • ❌ Don't use for: Large rows, write-heavy tables

2. Tune Key Cache Size

Cache partition key locations to skip index lookups!

-- In cassandra.yaml: /* key_cache_size_in_mb: 1024 # 1GB (default: auto) key_cache_save_period: 14400 # Save every 4 hours */ -- Monitor key cache nodetool info -- Key Cache Hit Rate: 99% (excellent!)

Sizing Formula:

key_cache_size = num_partitions × 200 bytes

3. Optimize Bloom Filter False Positive Rate

Lower false positives = fewer unnecessary SSTable reads!

-- Default: 1% false positive rate ALTER TABLE high_traffic_table WITH bloom_filter_fp_chance = 0.01; -- For critical tables: Lower to 0.1% (uses more memory) ALTER TABLE critical_data WITH bloom_filter_fp_chance = 0.001; -- Trade-off: -- 0.1% = 10x less false positives, 1.5x more memory

4. Use Appropriate Consistency Level

Lower consistency = faster reads!

-- Application code (Python example): # Fastest: LOCAL_ONE (one replica in local DC) session.execute(query, consistency_level=ConsistencyLevel.LOCAL_ONE) # Use for: Caching, session data, non-critical reads # Balanced: LOCAL_QUORUM (majority in local DC) session.execute(query, consistency_level=ConsistencyLevel.LOCAL_QUORUM) # Use for: Most production reads (good balance) # Strong: QUORUM (majority across all DCs) session.execute(query, consistency_level=ConsistencyLevel.QUORUM) # Use for: Critical data needing strong consistency

5. Optimize Compression

Better compression = less disk I/O!

-- LZ4: Fastest (default, recommended for most) ALTER TABLE high_throughput_table WITH compression = { 'class': 'LZ4Compressor', 'chunk_length_in_kb': 64 }; -- Snappy: Good balance ALTER TABLE general_purpose WITH compression = { 'class': 'SnappyCompressor' }; -- Deflate: Best compression (slower, for cold data) ALTER TABLE archive_data WITH compression = { 'class': 'DeflateCompressor' };

✍️ Write Path Optimization

Cassandra is write-optimized by design - make it even faster!

Why Cassandra Writes are Fast

⚡ The Write Path (Sub-Millisecond!)

  1. Commit Log (Sequential): Append to commit log on disk (~0.1ms)
  2. Memtable (Memory): Write to in-memory memtable (~0.01ms)
  3. Acknowledge: Return success to client
  4. Background Flush: Memtable → SSTable (async, doesn't block)

Total: 0.5-2ms per write! ⚡

Write Optimization Strategies

1. Separate Commit Log on Fast Disk

Commit log is the bottleneck - put it on the fastest disk!

-- In cassandra.yaml: /* # Put commit log on separate SSD commitlog_directory: /mnt/fast-ssd/commitlog # Increase commitlog size commitlog_total_space_in_mb: 8192 # 8GB # Sync mode commitlog_sync: periodic commitlog_sync_period_in_ms: 10000 # 10 seconds */

Hardware Recommendation:

  • ✅ Dedicated NVMe SSD for commit log
  • ✅ Separate from data SSD (parallel I/O)
  • ✅ Result: 5-10x faster writes!

2. Use Appropriate Consistency Level

Lower consistency = faster writes!

-- Fastest: LOCAL_ONE (one replica in local DC) session.execute(insert, consistency_level=ConsistencyLevel.LOCAL_ONE) # Latency: ~1ms # Use for: High-volume writes, logs, analytics -- Balanced: LOCAL_QUORUM (majority in local DC) session.execute(insert, consistency_level=ConsistencyLevel.LOCAL_QUORUM) # Latency: ~5ms # Use for: Most production writes -- Strong: QUORUM (majority across all DCs) session.execute(insert, consistency_level=ConsistencyLevel.QUORUM) # Latency: ~50ms (cross-DC network) # Use for: Critical data only

3. Use Prepared Statements

10x faster than unprepared!

-- ❌ SLOW: Unprepared statement (every time) session.execute( "INSERT INTO users (id, name) VALUES (1, 'John')" ) # Cassandra must parse query every time! -- ✅ FAST: Prepared statement (parse once) prepared = session.prepare( "INSERT INTO users (id, name) VALUES (?, ?)" ) session.execute(prepared, (1, 'John')) # 10x faster! Cassandra reuses parsed query

4. Batch Writes to Same Partition

Atomic updates to denormalized data!

-- ✅ GOOD: BATCH for same partition (denormalization) BEGIN BATCH INSERT INTO posts_by_user (user_id, post_id, ...) VALUES (...); INSERT INTO posts_by_tag (tag, post_id, ...) VALUES (...); APPLY BATCH; -- ❌ BAD: BATCH for bulk inserts (slower!) BEGIN BATCH INSERT INTO users (...) VALUES (1, ...); INSERT INTO users (...) VALUES (2, ...); INSERT INTO users (...) VALUES (3, ...); APPLY BATCH; -- Use async parallel writes instead!

5. Async Parallel Writes

Maximum throughput for bulk operations!

-- Python example: Async concurrent writes /* from cassandra.concurrent import execute_concurrent_with_args prepared = session.prepare("INSERT INTO users (...) VALUES (?, ?, ?)") # Execute 10,000 inserts in parallel! parameters = [(i, f'user_{i}', f'email_{i}@example.com') for i in range(10000)] results = execute_concurrent_with_args( session, prepared, parameters, concurrency=100 # 100 parallel requests ) # Result: 10,000 inserts in < 1 second! */

Write Performance Killers

  • ❌ Lightweight Transactions (LWT): 4x slower than normal writes (Paxos)
  • ❌ Cross-Partition BATCH: Slower than individual writes
  • ❌ Large Collections: > 1000 elements slow down writes
  • ❌ Too Many Materialized Views: > 3 MVs = -50% write throughput
  • ❌ Secondary Indexes: Each index adds ~10% write overhead

☕ JVM & Memory Tuning

Cassandra runs on the JVM - proper tuning prevents GC pauses!

Heap Size Configuration

⚠️ Common Mistake: Heap Too Large!

Many developers think "more heap = better performance". This is WRONG for Cassandra!

Problem with Large Heap:

  • 🔥 Long GC Pauses: 32GB heap = 5-10 second pauses!
  • 💥 Stop-the-World: All operations freeze during GC
  • ⏱️ Timeouts: Clients timeout waiting for response

✅ Optimal Heap Size

Total RAM Heap Size Off-Heap
16GB 8GB 8GB (OS + cache)
32GB 8-12GB 20-24GB
64GB 12-16GB 48-52GB
128GB+ 16-24GB Rest for cache

Golden Rule:

Heap = 8-16GB (never more than 32GB!)

JVM Configuration

# In jvm.options or jvm11-server.options: # Heap size (8GB recommended) -Xms8G -Xmx8G # Use G1GC (Garbage First Garbage Collector) -XX:+UseG1GC -XX:G1RSetUpdatingPauseTimePercent=5 -XX:MaxGCPauseMillis=200 # GC logging (for monitoring) -Xlog:gc*,gc+age=trace,safepoint:file=/var/log/cassandra/gc.log:time,uptime,level,tags:filecount=10,filesize=10M # Heap dump on OOM (for debugging) -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/log/cassandra/heap_dump.hprof # String deduplication (saves memory) -XX:+UseStringDeduplication

Memory Management Best Practices

✅

DO's

  • 8-16GB Heap: Sweet spot for most workloads
  • Use G1GC: Best for Cassandra (low pause times)
  • Monitor GC Pauses: Should be < 200ms
  • Off-Heap Caching: Use remaining RAM for OS cache
  • Matching Xms/Xmx: Prevents heap resizing
❌

DON'Ts

  • Don't Use > 32GB Heap: GC pauses become huge!
  • Don't Use CMS: Deprecated, use G1GC instead
  • Don't Ignore GC Logs: Monitor for long pauses
  • Don't Swap: Disable swap! (swappiness=0)
  • Don't Overcommit: Leave RAM for OS

GC Monitoring

# Check current GC stats nodetool gcstats # Output shows: /* Interval (ms): 1000 Max GC Elapsed (ms): 45 Total GC Elapsed (ms): 120 Stdev GC Elapsed (ms): 12 */ # Good: Max < 200ms, Total < 500ms per interval # Bad: Max > 1000ms = investigate immediately!

💻 Hardware Optimization

Right hardware = 10x performance improvement without code changes!

Recommended Hardware Configuration

🏆 Production-Grade Node Specs

Component Minimum Recommended High-Performance
CPU Cores 8 cores 16-32 cores 32-64 cores
RAM 16GB 32-64GB 128GB+
Commit Log Disk SSD (SATA) NVMe SSD NVMe (separate)
Data Disk SSD 500GB SSD 1-2TB NVMe 2-4TB
Network 1 Gbps 10 Gbps 25-100 Gbps

Critical Hardware Decisions

1. SSDs are MANDATORY (Never Use HDDs!)

SSDs provide 100x faster random I/O!

🐌

HDD Performance

  • Random Read: 50-100ms
  • Random Write: 50-100ms
  • IOPS: 100-200
  • Throughput: 100-200 MB/s

Result: Slow queries, timeouts! ❌

⚡

SSD Performance

  • Random Read: 0.1-1ms
  • Random Write: 0.1-1ms
  • IOPS: 50,000-500,000
  • Throughput: 500-7000 MB/s

Result: Blazing fast! ⚡

2. Separate Commit Log and Data Disks

Parallel I/O = 2x faster writes!

# cassandra.yaml configuration: # Commit log on separate fast SSD commitlog_directory: /mnt/nvme0/commitlog # Data on different SSD data_file_directories: - /mnt/nvme1/data # Saved caches on data disk saved_caches_directory: /mnt/nvme1/saved_caches

Why This Matters:

  • ✅ Eliminates I/O contention
  • ✅ Commit log gets dedicated IOPS
  • ✅ 2-5x faster write performance

3. CPU Cores for Parallelism

More cores = higher throughput!

  • Compaction: Uses multiple threads (more cores = faster)
  • Concurrent Requests: Each request gets a thread
  • Read/Write Parallelism: Scales with cores
  • Recommendation: 16-32 cores minimum for production

4. Network Bandwidth

10 Gbps minimum for production clusters!

Network Bottlenecks:

  • ⚠️ 1 Gbps: OK for dev/testing only
  • ✅ 10 Gbps: Minimum for production
  • 🚀 25-100 Gbps: High-performance clusters

Why: Replication, streaming, repairs all use network heavily!

5. Disable Swap!

Swapping kills Cassandra performance!

# Disable swap immediately sudo swapoff -a # Disable permanently (edit /etc/fstab, comment out swap) # Or set swappiness to 0: sudo sysctl -w vm.swappiness=0 echo "vm.swappiness=0" | sudo tee -a /etc/sysctl.conf # Verify cat /proc/sys/vm/swappiness # Should output: 0

📊 Monitoring & Troubleshooting

You can't optimize what you can't measure!

Key Metrics to Monitor

⏱️

Latency Metrics

  • Read Latency: P99 < 50ms
  • Write Latency: P99 < 10ms
  • Range Latency: P99 < 100ms
  • CAS Latency: P99 < 50ms
nodetool tablestats keyspace.table
📈

Throughput Metrics

  • Reads/sec: Per table
  • Writes/sec: Per table
  • Pending Tasks: Should be 0
  • Dropped Messages: Should be 0
nodetool tpstats
💾

Resource Metrics

  • CPU: 50-70% avg
  • Heap Usage: 60-80%
  • Disk I/O: < 80%
  • GC Pause: < 200ms
nodetool info nodetool gcstats

Essential Nodetool Commands

-- Check cluster status nodetool status -- Table statistics (latency, reads, writes) nodetool tablestats keyspace_name.table_name -- Thread pool stats (pending, active, blocked) nodetool tpstats -- GC statistics nodetool gcstats -- Compaction stats nodetool compactionstats -- Table histograms (detailed latency breakdown) nodetool tablehistograms keyspace_name table_name -- Proxy histograms (coordinator latency) nodetool proxyhistograms -- Check for slow queries (enable slow query logging) -- In cassandra.yaml: slow_query_log_timeout_in_ms: 500 grep "operations were slow" /var/log/cassandra/system.log

Common Performance Problems & Solutions

Problem #1: High Read Latency

Symptoms: Queries taking 100ms-1s+

Possible Causes & Fixes:

  • ✅ Large partitions: Add bucketing to data model
  • ✅ Low cache hit rate: Enable row cache for hot data
  • ✅ Too many SSTables: Run compaction
  • ✅ GC pauses: Reduce heap size, tune G1GC
  • ✅ Slow disks: Upgrade to NVMe SSDs

Problem #2: High Write Latency

Symptoms: Writes taking 50ms-500ms+

Possible Causes & Fixes:

  • ✅ Slow commit log: Move to separate NVMe SSD
  • ✅ Too many MVs: Reduce to 3-5 max per table
  • ✅ Consistency level too high: Use LOCAL_ONE/LOCAL_QUORUM
  • ✅ Pending compactions: Tune compaction strategy
  • ✅ GC pauses: Reduce heap, monitor gcstats

Problem #3: GC Pauses

Symptoms: Pauses > 1 second, timeouts

Possible Causes & Fixes:

  • ✅ Heap too large: Reduce to 8-16GB
  • ✅ Wrong GC: Use G1GC, not CMS
  • ✅ Memory leaks: Check for large partitions
  • ✅ Insufficient RAM: Increase off-heap memory
# Monitor GC in real-time watch -n 1 nodetool gcstats # Check GC logs tail -f /var/log/cassandra/gc.log

Problem #4: High CPU Usage

Symptoms: CPU constantly > 90%

Possible Causes & Fixes:

  • ✅ Heavy compaction: Schedule during off-peak hours
  • ✅ Too many concurrent requests: Add more nodes
  • ✅ Inefficient queries: Review query patterns
  • ✅ Repair running: Limit repair concurrency

Problem #5: Dropped Messages

Symptoms: nodetool tpstats shows dropped messages

Possible Causes & Fixes:

  • ✅ Overloaded cluster: Add more nodes
  • ✅ GC pauses: Tune JVM settings
  • ✅ Network issues: Check network bandwidth
  • ✅ Timeouts too low: Increase timeout settings

✅ Production Performance Checklist

Complete this checklist before going to production!

🎯 Pre-Production Checklist

Data Modeling

  • ☑️ All partitions < 100MB
  • ☑️ Time-series data uses bucketing
  • ☑️ One table per query pattern
  • ☑️ High cardinality partition keys
  • ☑️ No ALLOW FILTERING in production queries
  • ☑️ All queries use partition key
  • ☑️ Collections < 100 elements

Hardware

  • ☑️ SSDs for all disks (no HDDs!)
  • ☑️ Separate commit log and data disks
  • ☑️ 16-32+ CPU cores per node
  • ☑️ 32-64GB+ RAM per node
  • ☑️ 10 Gbps+ network
  • ☑️ Swap disabled (swappiness=0)

JVM Configuration

  • ☑️ Heap size: 8-16GB (never > 32GB)
  • ☑️ G1GC enabled
  • ☑️ GC logging enabled
  • ☑️ Xms = Xmx (matching heap sizes)
  • ☑️ GC pauses < 200ms

Table Configuration

  • ☑️ Appropriate compaction strategy
  • ☑️ Compression enabled (LZ4/Snappy)
  • ☑️ Bloom filter FP rate tuned
  • ☑️ Caching configured for hot tables
  • ☑️ TTL set for temporary data
  • ☑️ < 5 materialized views per table

Query Optimization

  • ☑️ All queries use prepared statements
  • ☑️ Consistency level: LOCAL_QUORUM or LOCAL_ONE
  • ☑️ LIMIT on all queries
  • ☑️ No SELECT * in production
  • ☑️ Batch only for denormalization

Monitoring

  • ☑️ Metrics collection enabled (Prometheus/Grafana)
  • ☑️ Slow query logging enabled (500ms threshold)
  • ☑️ GC monitoring active
  • ☑️ Alerts configured (latency, CPU, disk)
  • ☑️ nodetool commands scheduled (tablestats, tpstats)

Cluster Configuration

  • ☑️ Replication factor ≥ 3
  • ☑️ NetworkTopologyStrategy for production
  • ☑️ Multiple datacenters for HA
  • ☑️ Regular repair scheduled (weekly)
  • ☑️ Backup strategy in place

🖥️ Interactive Performance Analyzer

Analyze your Cassandra configuration for performance issues!

Performance Configuration Analyzer
🚀 Performance Analyzer Ready!
Enter your configuration and I'll analyze it for performance issues...

Supported Parameters:
• Heap Size, Disk Type, Partition Size
• Consistency Level, CPU Cores, Network
• Compaction Strategy, Row Cache, etc.

💼 Interview Questions & Expert Answers

Master Cassandra performance for your next interview!

1 Explain why Cassandra's write performance is so fast. What makes it different from traditional databases? ▼

Answer: Cassandra achieves fast writes through sequential I/O, no read-before-write, and append-only architecture.

The Write Path (Why It's Fast):

  1. Commit Log (Sequential Write): Data is appended to commit log on disk (~0.1ms). Sequential writes are 100x faster than random writes!
  2. Memtable (In-Memory): Data written to in-memory structure (~0.01ms)
  3. Acknowledge Client: Success returned immediately (total ~0.5-2ms)
  4. Background Flush: Memtable → SSTable happens asynchronously (doesn't block)

Key Differences from Traditional Databases:

Aspect Traditional DB Cassandra
Write Pattern Random I/O (in-place updates) Sequential (append-only)
Read Before Write Yes (for updates) No (upsert model)
Locking Row/table locks Lock-free
Durability WAL + data update Commit log only (initially)
Write Latency 10-50ms 0.5-2ms

Additional Optimizations:

  • No Redo/Undo Logs: Simpler durability mechanism
  • No B-Tree Updates: No complex tree rebalancing
  • Batch Commits: Commit log syncs every 10 seconds by default
  • Parallel Distribution: Writes distributed across multiple nodes

Result: Cassandra can handle 10,000-100,000 writes/sec per node!

2 What are the most common causes of poor read performance in Cassandra and how do you fix them? ▼

Answer: Poor read performance typically stems from large partitions, cache misses, too many SSTables, or inefficient queries.

Common Causes & Solutions:

1. Large Partitions (> 100MB)

Problem: Reading large partitions causes memory pressure and slow queries

Solution:

  • Add bucketing to data model (partition by date, range, or hash)
  • Target: 10-100MB per partition
  • Example: Change PK from (user_id) to ((user_id, date_bucket), timestamp)

2. Low Cache Hit Rate

Problem: Every read goes to disk (slow!)

Solution:

-- Enable row cache for hot data ALTER TABLE hot_table WITH caching = { 'keys': 'ALL', 'rows_per_partition': '100' }; -- Monitor cache hit rate nodetool info -- Look for: Row Cache Hit Rate > 90%

3. Too Many SSTables

Problem: Must read from many files, slows down queries

Solution:

-- Check SSTable count nodetool tablestats keyspace.table -- Force compaction nodetool compact keyspace table -- Tune compaction strategy ALTER TABLE table_name WITH compaction = { 'class': 'LeveledCompactionStrategy' };

4. Multi-Partition Queries

Problem: Scatter-gather across many nodes

Solution:

  • Redesign data model for single-partition reads
  • Create denormalized tables per query pattern
  • Use materialized views for multiple access patterns

5. GC Pauses

Problem: JVM stops responding during garbage collection

Solution:

  • Reduce heap size to 8-16GB (not 32GB+!)
  • Use G1GC with MaxGCPauseMillis=200
  • Monitor: nodetool gcstats
  • Target: GC pauses < 200ms

6. Slow Disks

Problem: HDDs or slow SSDs

Solution:

  • Upgrade to NVMe SSDs (100x faster random I/O)
  • HDD: 50-100ms latency → SSD: 0.1-1ms latency
3 Why is heap size limited to 8-16GB in Cassandra? What happens if you use 32GB+ heap? ▼

Answer: Large heaps (> 16GB) cause long GC pauses that freeze Cassandra, leading to timeouts and poor performance.

The Problem with Large Heaps:

With 32GB Heap:

  • 🔥 GC Pause Times: 5-10 seconds (stop-the-world!)
  • 💥 All Operations Freeze: No reads, no writes during GC
  • ⏱️ Client Timeouts: Requests timeout waiting for response
  • 📉 Throughput Drops: System appears "frozen" regularly
  • 😡 User Experience: App becomes unusable

Why Does This Happen?

  1. More Heap = More Objects: 32GB heap holds millions more objects than 8GB
  2. GC Must Scan Everything: Garbage collector must examine every object
  3. Stop-the-World: Application freezes while GC runs
  4. Time Proportional to Heap Size: 4x heap = ~4x pause time

With 8-16GB Heap:

  • ✅ GC Pause Times: 50-200ms (acceptable!)
  • ✅ No Noticeable Freezes: Operations continue smoothly
  • ✅ No Timeouts: Clients get responses quickly
  • ✅ High Throughput: System stays responsive

Why Not Just Disable GC?

You can't! Java requires garbage collection. Without it, you get OutOfMemoryError and crash.

The Solution: Off-Heap Memory

Cassandra uses off-heap memory for caching!

Total RAM Heap Off-Heap (OS Cache)
64GB 12-16GB 48-52GB
128GB 16GB 112GB

OS page cache uses off-heap memory to cache SSTables - no GC required!

Key Takeaway:

More RAM is good, but add it as off-heap memory, not heap! Keep heap at 8-16GB max.

4 How does consistency level affect performance? What's the trade-off between consistency and latency? ▼

Answer: Lower consistency levels provide faster performance but weaker consistency guarantees. The trade-off is between speed and data accuracy.

Consistency Level Performance Comparison:

Level Nodes Latency Consistency
LOCAL_ONE 1 (local DC) ⚡ 1-2ms Weakest
LOCAL_QUORUM Majority (local) ⚡ 5-10ms Strong (local)
QUORUM Majority (all DCs) 🐌 50-100ms Strong (global)
ALL All replicas 🐌 100-500ms Strongest

Why The Performance Difference?

LOCAL_ONE (Fastest):

  • Coordinator sends request to 1 closest replica
  • Returns immediately when that 1 node responds
  • No waiting for other nodes
  • ⚡ Latency: ~1-2ms

LOCAL_QUORUM (Balanced):

  • Coordinator sends to multiple replicas in local DC
  • Waits for majority (RF=3 → needs 2 nodes)
  • Returns when 2 nodes respond
  • ⚡ Latency: ~5-10ms

QUORUM (Slow - Cross-DC):

  • Coordinator sends to replicas across ALL datacenters
  • Waits for majority globally
  • Cross-datacenter network latency adds 50-100ms+
  • 🐌 Latency: ~50-100ms+

The Consistency Trade-off:

LOCAL_ONE Risk:

If you read from replica A, then immediately read from replica B, you might see stale data (eventual consistency). But it's FAST!

LOCAL_QUORUM Guarantee:

Reading from majority ensures you see the latest write (if write was also LOCAL_QUORUM). Good balance of speed + consistency!

Production Recommendations:

  • Cache/Sessions: LOCAL_ONE (speed matters, stale OK)
  • General Application: LOCAL_QUORUM (best balance)
  • Financial/Critical: QUORUM (strong consistency needed)
  • Never Use ALL: One node down = entire cluster fails!
5 What would you do if you're seeing high CPU usage and slow queries in production? ▼

Answer: Follow a systematic troubleshooting approach: check metrics → identify bottleneck → apply targeted fix.

Step-by-Step Troubleshooting Process:

Step 1: Gather Metrics

# Check overall node stats nodetool info nodetool status # Check thread pools (look for blocked/pending) nodetool tpstats # Check GC stats (long pauses?) nodetool gcstats # Check table-level performance nodetool tablestats keyspace.table # Check for slow queries grep "slow" /var/log/cassandra/system.log

Step 2: Identify the Bottleneck

Scenario A: GC Pauses > 1 Second

Symptoms: nodetool gcstats shows Max GC > 1000ms

Fix:

  • Reduce heap size from 32GB → 12GB
  • Ensure G1GC is enabled
  • Check for memory leaks (large partitions)

Scenario B: High Pending Compactions

Symptoms: nodetool compactionstats shows 50+ pending

Fix:

# Increase compaction throughput nodetool setcompactionthroughput 64 # Or tune in cassandra.yaml: compaction_throughput_mb_per_sec: 64

Scenario C: Large Partitions

Symptoms: nodetool tablestats shows partitions > 100MB

Fix:

  • Add bucketing to data model (requires migration)
  • Example: (user_id) → ((user_id, date_bucket), timestamp)

Scenario D: Too Many SSTables

Symptoms: Each read touches 20+ SSTables

Fix:

# Force major compaction (off-peak hours!) nodetool compact keyspace table # Or change compaction strategy ALTER TABLE table_name WITH compaction = { 'class': 'LeveledCompactionStrategy' };

Scenario E: Inefficient Queries

Symptoms: Queries without partition key, ALLOW FILTERING

Fix:

  • Review application queries (enable slow query log)
  • Create denormalized tables for common queries
  • Add indexes (carefully!) or materialized views

Step 3: Quick Wins (Immediate Relief)

  • ✅ Reduce consistency level (QUORUM → LOCAL_QUORUM)
  • ✅ Enable row cache for hot tables
  • ✅ Force compaction on problematic tables
  • ✅ Increase connection pool size in application
  • ✅ Add more nodes (horizontal scaling)

Step 4: Long-Term Fixes

  • 📊 Fix data model (add bucketing for large partitions)
  • ⚡ Upgrade hardware (NVMe SSDs, more cores)
  • 🎯 Optimize queries (add prepared statements)
  • 📈 Implement proper monitoring (Prometheus + Grafana)

🎓 Chapter Summary: Performance Mastery

Congratulations! You now know how to optimize Cassandra for production-level performance!

Key Performance Principles:

  1. One Query = One Partition: Fastest queries read single partition
  2. Small Partitions: Keep < 100MB for best performance
  3. SSDs Are Mandatory: 100x faster than HDDs
  4. Heap Size 8-16GB: Never exceed 32GB (GC pauses!)
  5. Consistency Level Matters: LOCAL_QUORUM for balance

Performance Optimization Checklist:

  • ✅ Partitions < 100MB (use bucketing)
  • ✅ SSDs for all disks (NVMe preferred)
  • ✅ Separate commit log and data disks
  • ✅ Heap size: 8-16GB max
  • ✅ G1GC enabled with MaxGCPauseMillis=200
  • ✅ Row cache for hot data
  • ✅ Consistency: LOCAL_QUORUM or LOCAL_ONE
  • ✅ Prepared statements everywhere
  • ✅ < 5 materialized views per table
  • ✅ Regular compaction and repair

Target Latencies:

Operation Target P99
Single-partition read < 5ms
Single write < 2ms
Multi-partition query < 20ms
GC pause < 200ms

Common Performance Killers:

  • ❌ Large partitions (> 100MB)
  • ❌ Heap size > 32GB
  • ❌ Using HDDs instead of SSDs
  • ❌ ALLOW FILTERING in production
  • ❌ Too many MVs (> 5 per table)
  • ❌ No partition key in queries
  • ❌ Consistency level ALL

🚀 You're now ready to build blazing-fast Cassandra systems!

Remember Netflix: Following these principles = 99.99% uptime with millisecond latency! 🎉

Advertisement

Responsive Ad