Disk Optimization in Cassandra
Make your storage work harder and smarter! Disk optimization helps Cassandra store and access data efficiently, so your database stays fast, stable, and ready for heavy workloads.
📖 The Story: Jake's 500ms Write Crisis
Jake's Cassandra cluster was supposed to handle 100K writes/sec. But writes were taking 500ms instead of 5ms. Users complained. Timeouts everywhere. He had 10 Gbps network, 128GB RAM, 32-core CPUs. But everything was slow. The bottleneck? His disks. Here's what he discovered...
😱 The Problem: HDDs + Single Disk
Jake's Original Setup:
- 💾 Disks: 4TB SATA HDD (spinning disk!)
- 📁 Layout: Everything on one disk (commit log + data)
- ⚡ Performance: 100 IOPS, 10ms latency
- 📊 Reality: Write: 500ms, Read: 200ms
- 😡 Users: Constant timeouts!
What Was Happening:
- Write Request Arrives: Client wants to write data
- Commit Log Write: Must write to disk (fsync) = 10ms on HDD
- Disk Queue Builds: 100K writes/sec but disk only 100 IOPS
- Queue = 1000 deep: Each write waits for 1000 others
- Latency = 1000 × 5ms = 5 seconds! 💥
- Same Disk for Reads: Compaction also fighting for I/O
Why This Failed:
- 💾 HDD Too Slow: 100 IOPS vs 100K writes/sec needed
- 🔀 Single Disk: Commit log + data + compaction fighting
- ⏱️ Seek Time: HDDs need 10ms to seek, SSDs don't
- 📊 Queue Buildup: Can't keep up with write rate
- 🔥 Read Impact: Reads also slow (same bottleneck)
🚀 The Fix: NVMe SSDs + Separate Disks!
Jake's New Setup:
Disk 1 (NVMe SSD): Commit Log Only
Disks 2-3 (SATA SSD): Data
The Results:
- ⚡ Write Latency: 500ms → 2ms (250x faster!)
- ⚡ Read Latency: 200ms → 5ms (40x faster!)
- 📊 Throughput: 1K writes/sec → 100K writes/sec
- 🎯 Disk Util: 100% → 30% (headroom!)
- 😊 Timeouts: 90% → 0%
- 💰 ROI: $2K in SSDs saved $50K/month in lost revenue!
Jake learned: SSDs are not optional, they're mandatory! 🎉
💿 Disk Fundamentals
Why disk performance is critical for Cassandra!
🎯 Why Disks Matter
Cassandra is disk-intensive by design. Every write hits commit log (fsync), reads scan SSTables, compaction rewrites data. Even with caching, disks are the ultimate bottleneck. Fast disks = fast Cassandra. Slow disks = disaster.
HDD vs SSD Comparison
Disk Performance Metrics
| Metric | HDD | SATA SSD | NVMe SSD |
|---|---|---|---|
| Random Read IOPS | 100-200 | 50,000-100,000 | 500,000-1,000,000 |
| Random Write IOPS | 100-200 | 40,000-90,000 | 400,000-800,000 |
| Latency | 10-20ms | 0.5-1ms | 0.1-0.2ms |
| Sequential Read | 150 MB/s | 550 MB/s | 3,500 MB/s |
| Sequential Write | 150 MB/s | 520 MB/s | 3,000 MB/s |
| Price (1TB) | $40 | $100 | $150 |
| Cassandra Use | ❌ Never | ✅ Data disks | ⭐ Commit log |
NEVER Use HDDs for Cassandra!
HDDs are 100-1000x slower than SSDs. Using HDDs means:
- 💥 Write latency: 500ms instead of 2ms
- 💥 Throughput: 100 writes/sec instead of 100K
- 💥 Constant timeouts and errors
- 💥 Unusable in production
SSDs are mandatory, not optional!
📝 Commit Log Optimization
The most critical disk optimization!
⚡ Why Commit Log is Critical
Every write must be written to commit log with fsync before acknowledging. This is THE bottleneck for write performance. Commit log on slow disk = slow writes. Commit log on fast disk = fast writes. It's that simple.
Write Path and Commit Log
Step 1: Write to Commit Log (Disk)
MUST fsync to disk for durability
Step 2: Write to Memtable (Memory)
Fast in-memory operation
Step 3: Acknowledge to Client
Total time = commit log time + network
Separate Commit Log Disk
Same Disk (Bad)
- Problem: Commit log + data + compaction
- Contention: All fighting for I/O
- Seeks: Disk head jumping around
- Result: 10-20ms writes
- Throughput: 1K-5K writes/sec
Never use same disk!
Separate Disks (Good)
- Disk 1 (NVMe): Commit log only
- Disk 2 (SATA): Data + compaction
- No Contention: Isolated I/O
- Result: 0.5-2ms writes
- Throughput: 50K-100K writes/sec
2-5x faster writes!
Commit Log Configuration
Commit Log Sizing
How Much Space for Commit Log?
💾 Data Disk Optimization
Optimizing SSTable storage and compaction!
RAID Configuration
RAID 5/6 (Bad)
- Problem: Write penalty (4x I/O)
- Parity: Calculates parity on every write
- Performance: 75% slower writes
- Why: Read-modify-write cycle
- Use Case: None for Cassandra
Never use RAID 5/6!
RAID 1/10 (OK)
- Mirroring: Data written to 2+ disks
- Performance: Good (no write penalty)
- Capacity: 50% (2 disks = 1 disk space)
- Why OK: Cassandra has replication
- Use Case: If must have RAID
Acceptable but expensive
RAID 0 or JBOD (Best)
- RAID 0: Stripe across disks
- JBOD: Each disk separate
- Performance: Full disk speed
- Capacity: 100% (4 disks = 4x space)
- Why Best: Cassandra handles failures
Recommended for Cassandra!
Why RAID 0/JBOD for Cassandra?
Cassandra's Built-In Fault Tolerance
Cassandra has its own replication (RF=3 means 3 copies). Disk failure just means one replica is down. Other replicas serve requests. RAID adds overhead for redundancy you already have!
File System Choice
Data Disk Sizing
Calculate Required Space
Multiple Data Disks
⚙️ I/O Scheduler & Kernel Tuning
OS-level optimizations for disk performance!
I/O Scheduler Selection
CFQ (Bad for SSDs)
- Design: For HDDs (seek optimization)
- Problem: Adds latency for SSDs
- Overhead: Unnecessary queueing
- Result: 20-30% slower
deadline (Good)
- Design: Low latency
- Works: Well with SSDs
- Overhead: Minimal
- Result: Good performance
noop/none (Best)
- Design: No scheduling
- Perfect: For NVMe SSDs
- Overhead: Zero
- Result: Maximum performance
Configure I/O Scheduler
Kernel Parameters
Read-Ahead Settings
📊 Disk Monitoring
Essential commands to monitor disk performance!
Key Disk Metrics
Health Thresholds
| Metric | Healthy | Warning | Critical |
|---|---|---|---|
| Disk Utilization | < 60% ✅ | 60-80% ⚠️ | > 80% ❌ |
| Await (latency) | < 5ms ✅ | 5-10ms ⚠️ | > 20ms ❌ |
| Queue Size | < 10 ✅ | 10-50 ⚠️ | > 100 ❌ |
| Space Used | < 70% ✅ | 70-85% ⚠️ | > 85% ❌ |
Disk Space Monitoring
SSD Health Monitoring
💼 Interview Questions & Expert Answers
Master disk optimization for your interview!
Answer: SSDs provide 100-1000x more IOPS (50K vs 100) and 100x lower latency (0.5ms vs 15ms) than HDDs. Using HDDs causes write latency of 500ms+ instead of 2ms, making Cassandra unusable for production workloads.
Performance Comparison:
| Metric | HDD | SSD | Difference |
|---|---|---|---|
| IOPS | 100-200 | 50,000-500,000 | 500-5000x faster |
| Latency | 10-20ms | 0.1-1ms | 10-200x faster |
| Sequential | 150 MB/s | 500-3500 MB/s | 3-23x faster |
Why Cassandra Needs High IOPS:
- Commit Log: Every write requires fsync (disk sync)
- Compaction: Rewrites data constantly (background I/O)
- Reads: May need to read from 5-10 SSTables
- Repairs: Stream data across network and disk
Real-World Impact:
Key Takeaway: SSDs aren't a nice-to-have, they're absolutely mandatory. HDDs make Cassandra 100-1000x slower and unusable for any real workload!
Answer: No! Commit log should be on a separate NVMe SSD. When on the same disk, commit log competes with reads, compaction, and SSTable writes, causing I/O contention and 2-5x slower writes. Separate disk = isolated I/O paths.
Problem with Same Disk:
Solution: Separate Disks
Performance Gain:
| Configuration | Write Latency | Throughput |
|---|---|---|
| Same disk (SATA) | 10-20ms | 5K writes/sec |
| Separate (NVMe + SATA) | 0.5-2ms | 50K+ writes/sec |
Cost-Benefit:
- 💰 Cost: $150 for 500GB NVMe SSD
- ⚡ Benefit: 2-5x faster writes, 10x more throughput
- 📊 ROI: Pays for itself immediately in performance
Key Takeaway: Commit log on separate NVMe SSD is one of the highest ROI optimizations for Cassandra! Always separate commit log and data disks.
Answer: RAID 5/6 has 4x write penalty (read-modify-write for parity), making writes 75% slower. Cassandra already has replication (RF=3), so RAID redundancy is unnecessary overhead. RAID 0/JBOD provides full performance and capacity without duplicate redundancy.
RAID 5/6 Write Penalty:
Cassandra's Built-In Redundancy:
RAID Configuration Comparison:
| RAID Level | Performance | Capacity | Recommendation |
|---|---|---|---|
| RAID 5/6 | 25% speed (4x penalty) | 75% (1 disk parity) | ❌ Never use |
| RAID 10 | 100% speed | 50% (mirrored) | ⚠️ Acceptable but costly |
| RAID 0 | 100%+ speed (striped) | 100% | ✅ Recommended |
| JBOD | 100% speed | 100% | ✅ Also good |
Why RAID 0 or JBOD Works:
- ✅ Full Performance: No parity overhead
- ✅ Full Capacity: No mirroring waste
- ✅ Cassandra Handles Failures: RF=3 provides redundancy
- ✅ Simpler: Less hardware complexity
- ✅ Cheaper: More usable space per dollar
Key Takeaway: Don't add RAID redundancy when Cassandra already provides it! Use RAID 0 or JBOD for maximum performance and capacity.
Answer: Monitor disk utilization (< 80%), await latency (< 5ms), queue size (< 10), IOPS (vs capacity), and disk space (< 70% full). High utilization or latency indicates I/O bottleneck requiring faster disks or reducing load.
Key Metrics to Track:
1. Disk Utilization (%util)
2. Await (Average Wait Time)
3. Queue Size (avgqu-sz)
4. IOPS (Operations/Sec)
5. Disk Space
Monitoring Dashboard:
Alert Thresholds:
- ⚠️ WARNING: %util > 60%, await > 5ms, space > 70%
- ❌ CRITICAL: %util > 80%, await > 10ms, space > 85%
Answer: Per-node calculation: (10TB / nodes / RF) × space_amplification × safety_margin. For 10 nodes: (10TB / 10 / 3) × 1.5 (STCS) × 1.5 (snapshots+margin) = 750GB per node. Provision 1TB per node for headroom.
Step-by-Step Sizing:
Step 1: Calculate Per-Node Raw Data
Step 2: Account for Space Amplification
Step 3: Provision Disks
Step 4: Validation
Quick Formula:
Key Takeaway: Always provision 1.5-2.5x raw data size per node to account for compaction, snapshots, and safety margin. Never run disks > 85% full!
🎓 Chapter Summary: Disk Optimization Mastery
You now understand disk optimization at a production level!
Critical Rules:
- 💿 SSDs Mandatory: 100-1000x faster than HDDs
- 📝 Separate Commit Log: NVMe for commit log, SATA for data
- 🔀 RAID 0/JBOD: No RAID 5/6 (Cassandra has replication!)
- ⚙️ I/O Scheduler: 'none' for NVMe, 'deadline' for SATA
- 📊 Monitor: %util < 80%, await < 5ms, space < 70%
Recommended Setup:
Performance Gain:
HDD: 500ms writes, 1K/sec → SSD: 2ms writes, 100K/sec (250x faster!)
Remember Jake: SSDs are not optional! 🚀
Responsive Ad