Benchmarking in Cassandra
Measure what really matters! Benchmarking helps you test Cassandraโs performance under different workloads, so you can compare results, tune the system, and ensure it delivers consistent speed and scalability in real-world conditions.
๐ The Story: Marcus's Misleading Benchmark
Marcus ran a benchmark that showed 500K writes/sec. He deployed to production expecting amazing performance. Instead, he got 5K writes/sec - 100x slower! His boss asked: "Why did the benchmark lie?" Marcus learned the hard way: bad benchmarks are worse than no benchmarks.
๐ฑ The Bad Benchmark
What Marcus Did:
What Was Wrong:
- Single Node: Tested only 1 node, not cluster behavior
- No Replication: RF=1 (production needs RF=3!)
- All in Memory: Working set 200MB, server had 128GB RAM
- No Reads: Only writes (production is 80% reads)
- Sequential Keys: Perfect case, production has random keys
- Unrealistic Data: Fixed-size rows, production varies wildly
- No Compaction: Test ran 5 minutes (compaction kicks in after hours)
Production Reality:
The Consequences:
- ๐ฅ Undersized Cluster: Planned for 2 nodes, needed 20
- ๐ธ Cost Overrun: 10x more infrastructure needed
- โฐ Timeline Blown: 3-month delay to add capacity
- ๐ก Users Unhappy: Slow, unreliable for weeks
- ๐ฏ Boss Unhappy: "Why didn't you test properly?"
โ The Proper Benchmark
What Marcus Should Have Done:
Realistic Results:
- ๐ Throughput: 25K ops/sec total (not 500K!)
- โฑ๏ธ Write Latency: P99 = 15ms (not 2ms!)
- โฑ๏ธ Read Latency: P99 = 8ms
- ๐พ Disk I/O: 60% utilized (not 10%!)
- ๐ Compaction: Running continuously
- ๐ฏ Realistic: Matches production!
Proper Capacity Planning:
- โ Expected Load: 50K ops/sec
- โ Benchmark Shows: 25K ops/sec per 3-node cluster
- โ Calculation: Need 6 nodes (2ร for 2x headroom)
- โ Budget: Planned correctly from start
- โ Result: Production launch smooth! ๐
Marcus learned: Benchmark production workload, not toy scenarios! ๐ฏ
๐ฏ Benchmarking Fundamentals
The principles of meaningful performance testing!
๐ The Golden Rules
- Match Production Workload: Same read/write ratio, data size, access patterns
- Match Production Cluster: Same RF, consistency levels, number of nodes
- Run Long Enough: 24+ hours to include compaction, GC cycles
- Realistic Data: Production-like data sizes and distributions
- Monitor Everything: CPU, disk, network, latency percentiles
- Don't Trust Peaks: Sustained performance matters, not bursts
Common Benchmarking Mistakes
Bad Benchmark
- Single node test
- RF=1 (no replication)
- Write-only workload
- Sequential keys
- 5-minute duration
- All data in RAM
- Report peak throughput
100x off from production!
Good Benchmark
- Full cluster (โฅ3 nodes)
- RF=3 (production config)
- Mixed read/write (80/20)
- Random/Gaussian keys
- 24+ hour duration
- Data > RAM (realistic)
- Report sustained P99 latency
Predicts production!
Metrics That Matter
| Metric | Why It Matters | What to Report |
|---|---|---|
| Throughput | How many ops/sec sustained | Sustained rate, not peak |
| P99 Latency | Worst-case user experience | 99th percentile (not average!) |
| Resource Util | Bottlenecks and headroom | CPU, disk, network % |
| GC Pauses | Impact on tail latency | Max pause time, frequency |
| Compaction | Long-term stability | Pending tasks, SSTable count |
Don't Trust Marketing Benchmarks!
Vendor benchmarks often show:
- Peak throughput (not sustained)
- Best-case scenarios (all in memory)
- Single node (not distributed)
- Synthetic workloads (not production-like)
Always run your own benchmarks with your workload!
๐จ cassandra-stress Tool
The official Cassandra benchmarking tool!
Basic Usage
Simple Write Test
Read Test
Mixed Workload (Most Important!)
Advanced Options
Access Patterns
Sequential (Unrealistic)
Uniform (Somewhat Realistic)
Gaussian (Realistic!)
Reading Results
๐ Custom Workload Profiles
Testing your actual schema and queries!
Creating Custom Schema
Running Custom Workload
Time-Series Workload Example
๐ Monitoring During Benchmarks
What to watch while benchmarking!
Essential Metrics
System Resources
Cassandra Metrics
Application Metrics
Monitoring Script
โญ Benchmarking Best Practices
Lessons learned from production benchmarks!
1. Warm Up First
Don't measure cold start performance!
2. Run Long Enough
Short tests miss compaction impact!
3. Test Data Larger Than RAM
Force disk I/O like production!
4. Test With Production RF
RF=3 means 3x more work!
5. Test Failure Scenarios
How does it perform when things go wrong?
The 80/20 Rule
For capacity planning:
- Benchmark shows: 50K ops/sec sustained
- Plan for: 40K ops/sec (80% of benchmark)
- Reason: Leaves headroom for:
- Traffic spikes
- Node failures
- Repairs/compaction
- GC pauses
๐ผ Interview Questions & Expert Answers
Master benchmarking for your interview!
Answer: Single-node RF=1 benchmark doesn't account for: (1) replication overhead (RF=3 means 3x write load), (2) network latency between nodes, (3) coordinator overhead, (4) consistency level coordination, (5) cluster-wide compaction impact. Can be 10-100x off from production performance.
What Single-Node RF=1 Misses:
Production Reality (RF=3, 3 nodes):
Performance Impact Breakdown:
| Factor | Single Node RF=1 | Cluster RF=3 |
|---|---|---|
| Writes per operation | 1 (local) | 3 (distributed) |
| Network hops | 0 | 2 (+ 2-4ms) |
| Coordination overhead | None | CL=QUORUM logic |
| Disk contention | Single disk | 3 disks (better!) |
| Compaction impact | Local only | Cluster-wide |
Key Takeaway: Always benchmark full cluster with production RF. Single-node tests can be 10-100x faster than production reality!
Answer: Short benchmarks miss critical long-term effects: (1) compaction overhead kicks in after hours, (2) GC pressure builds as heap fills, (3) OS cache effectiveness stabilizes, (4) SSTable count grows, (5) disk I/O patterns change. 5-minute test shows "best case", 24-hour shows "sustained reality".
What Happens Over Time:
0-5 Minutes (Honeymoon Phase):
1-4 Hours (Reality Sets In):
8-12 Hours (Steady State):
24+ Hours (Long-Term Stability):
Performance Over Time Graph:
Key Takeaway: 5-minute tests show best-case burst performance. 24-hour tests show sustainable production reality. Always test long enough to see compaction impact!
Answer: Sequential (keys 1,2,3...) has perfect cache locality - all data accessed in order. Gaussian (hot keys + cold keys) mimics production where some data is accessed frequently, some rarely. Sequential can be 10-50x faster due to caching, making it unrealistic for capacity planning.
Sequential Access Pattern:
Gaussian Access Pattern:
Performance Comparison:
| Pattern | Cache Hit Rate | Read Throughput | P99 Latency |
|---|---|---|---|
| Sequential | 95% | 200K ops/sec | 2ms |
| Uniform | 30% | 40K ops/sec | 12ms |
| Gaussian | 45% | 50K ops/sec | 10ms |
| Production | 40-50% | 45-55K ops/sec | 8-15ms |
Choosing the Right Pattern:
- โ Sequential: Never use (except for debugging)
- โ Gaussian: Best for most workloads (hot data + cold data)
- โ Uniform: Good for evenly distributed workloads (rare)
- โญ Custom: Best if you know your actual distribution
Key Takeaway: Sequential access is 4-10x faster due to perfect caching. Always use Gaussian or production-like distributions for realistic capacity planning!
Answer: 8 nodes. Calculation: Need 200K ร 2 = 400K capacity. Benchmark shows 100K on 3 nodes = 33K per node. 400K / 33K = 12 nodes. But account for RF=3 coordination overhead and node failures, so ~8-10 nodes with proper tuning.
Step-by-Step Calculation:
Step 1: Understand Your Benchmark
Step 2: Account for RF=3
Step 3: Calculate for Target Load
Step 4: Account for Scaling Factors
Step 5: Conservative Recommendation
Quick Formula:
Key Takeaway: Always include 2x headroom, account for scaling efficiency (80-90%), and monitor to add capacity before hitting limits. Start conservative, scale up as needed!
Answer: P99 latency shows worst-case user experience (1 in 100 requests). Average hides tail latency - can have 2ms average with 500ms P99. Also critical: sustained throughput (not peak), resource utilization (headroom), GC pause times, and compaction keeping up. P99 > 50ms means poor user experience.
Why Average Latency is Misleading:
Critical Metrics Ranked:
| Metric | Why It Matters | Target |
|---|---|---|
| P99 Latency โญ | Worst-case user experience | < 50ms |
| Sustained Throughput โญ | Can you handle load long-term? | 2x target load |
| Resource Utilization | Headroom for spikes | CPU < 70%, Disk < 75% |
| P999 Latency | Absolute worst case | < 100ms |
| GC Max Pause | Causes latency spikes | < 200ms |
| Compaction Pending | Long-term stability | < 20 tasks |
| Average Latency | Mostly marketing | Ignore this! |
Real-World Example:
How to Read Benchmark Output:
Key Takeaway: P99 shows real user experience. Average is marketing fluff. Always demand P99 < 50ms and P999 < 100ms for production systems!
๐ Chapter Summary: Benchmarking Mastery
You now understand production-grade benchmarking!
Golden Rules:
- ๐ฏ Full Cluster: 3+ nodes with RF=3 (not single node!)
- โฑ๏ธ Run Long: 24+ hours to see compaction
- ๐ Match Production: 80/20 read/write, Gaussian distribution
- ๐พ Data > RAM: Force disk I/O
- ๐ Watch P99: Not average (P99 < 50ms)
- ๐๏ธ 2x Headroom: For spikes and failures
Essential Command:
Red Flags:
- โ 5-minute test (misses compaction!)
- โ Sequential keys (unrealistic cache)
- โ Write-only (production is mixed!)
- โ Reporting average latency (watch P99!)
Remember Marcus: Bad benchmarks are worse than no benchmarks! ๐ฏ
Responsive Ad