Advanced Performance Topics

Disk Optimization in Cassandra

Make your storage work harder and smarter! Disk optimization helps Cassandra store and access data efficiently, so your database stays fast, stable, and ready for heavy workloads.

📖 The Story: Jake's 500ms Write Crisis

Jake's Cassandra cluster was supposed to handle 100K writes/sec. But writes were taking 500ms instead of 5ms. Users complained. Timeouts everywhere. He had 10 Gbps network, 128GB RAM, 32-core CPUs. But everything was slow. The bottleneck? His disks. Here's what he discovered...

😱 The Problem: HDDs + Single Disk

Jake's Original Setup:

  • 💾 Disks: 4TB SATA HDD (spinning disk!)
  • 📁 Layout: Everything on one disk (commit log + data)
  • ⚡ Performance: 100 IOPS, 10ms latency
  • 📊 Reality: Write: 500ms, Read: 200ms
  • 😡 Users: Constant timeouts!

What Was Happening:

  1. Write Request Arrives: Client wants to write data
  2. Commit Log Write: Must write to disk (fsync) = 10ms on HDD
  3. Disk Queue Builds: 100K writes/sec but disk only 100 IOPS
  4. Queue = 1000 deep: Each write waits for 1000 others
  5. Latency = 1000 × 5ms = 5 seconds! 💥
  6. Same Disk for Reads: Compaction also fighting for I/O
-- Disk stats showing disaster: iostat -x 1 /* Device r/s w/s util sda 50 100 100% ← Completely saturated! avgqu-sz: 1247 ← 1247 operations waiting! await: 5241ms ← Average wait: 5+ seconds! */ -- HDD specs (terrible for Cassandra): /* IOPS: 100-200 (vs SSD: 50,000-500,000!) Latency: 10-20ms (vs SSD: 0.1-1ms!) Sequential: 150 MB/s (vs SSD: 500-3500 MB/s!) */

Why This Failed:

  • 💾 HDD Too Slow: 100 IOPS vs 100K writes/sec needed
  • 🔀 Single Disk: Commit log + data + compaction fighting
  • ⏱️ Seek Time: HDDs need 10ms to seek, SSDs don't
  • 📊 Queue Buildup: Can't keep up with write rate
  • 🔥 Read Impact: Reads also slow (same bottleneck)

🚀 The Fix: NVMe SSDs + Separate Disks!

Jake's New Setup:

Disk 1 (NVMe SSD): Commit Log Only

-- 500GB NVMe SSD for commit log Device: /dev/nvme0n1 Mount: /var/lib/cassandra/commitlog IOPS: 500,000 ← 5000x better! Latency: 0.1ms ← 100x better! # In cassandra.yaml: commitlog_directory: /var/lib/cassandra/commitlog

Disks 2-3 (SATA SSD): Data

-- 2× 4TB SATA SSDs in RAID 0 for data Devices: /dev/sda, /dev/sdb Mount: /var/lib/cassandra/data IOPS: 100,000 ← 1000x better! Latency: 1ms ← 10x better! # In cassandra.yaml: data_file_directories: - /var/lib/cassandra/data

The Results:

  • ⚡ Write Latency: 500ms → 2ms (250x faster!)
  • ⚡ Read Latency: 200ms → 5ms (40x faster!)
  • 📊 Throughput: 1K writes/sec → 100K writes/sec
  • 🎯 Disk Util: 100% → 30% (headroom!)
  • 😊 Timeouts: 90% → 0%
  • 💰 ROI: $2K in SSDs saved $50K/month in lost revenue!

Jake learned: SSDs are not optional, they're mandatory! 🎉

💿 Disk Fundamentals

Why disk performance is critical for Cassandra!

🎯 Why Disks Matter

Cassandra is disk-intensive by design. Every write hits commit log (fsync), reads scan SSTables, compaction rewrites data. Even with caching, disks are the ultimate bottleneck. Fast disks = fast Cassandra. Slow disks = disaster.

HDD vs SSD Comparison

HDD vs SSD Performance HDD (Spinning Disk) IOPS: 100-200 Latency: 10-20ms Sequential: 150 MB/s Random: 1 MB/s ❌ NOT for Cassandra! SATA SSD IOPS: 50K-100K Latency: 0.5-1ms Sequential: 550 MB/s Random: 500 MB/s ✅ Good for data NVMe SSD IOPS: 500K+ Latency: 0.1ms Sequential: 3500 MB/s Random: 2500 MB/s ⭐ Best for commit log Write Performance Comparison HDD: 100 SATA SSD: 50,000 (500x faster!) NVMe SSD: 500,000 (5000x faster!) Cassandra Recommendation: SSDs Mandatory! NVMe for commit log, SATA SSD for data

Disk Performance Metrics

Metric HDD SATA SSD NVMe SSD
Random Read IOPS 100-200 50,000-100,000 500,000-1,000,000
Random Write IOPS 100-200 40,000-90,000 400,000-800,000
Latency 10-20ms 0.5-1ms 0.1-0.2ms
Sequential Read 150 MB/s 550 MB/s 3,500 MB/s
Sequential Write 150 MB/s 520 MB/s 3,000 MB/s
Price (1TB) $40 $100 $150
Cassandra Use ❌ Never ✅ Data disks ⭐ Commit log

NEVER Use HDDs for Cassandra!

HDDs are 100-1000x slower than SSDs. Using HDDs means:

  • 💥 Write latency: 500ms instead of 2ms
  • 💥 Throughput: 100 writes/sec instead of 100K
  • 💥 Constant timeouts and errors
  • 💥 Unusable in production

SSDs are mandatory, not optional!

📝 Commit Log Optimization

The most critical disk optimization!

⚡ Why Commit Log is Critical

Every write must be written to commit log with fsync before acknowledging. This is THE bottleneck for write performance. Commit log on slow disk = slow writes. Commit log on fast disk = fast writes. It's that simple.

Write Path and Commit Log

Step 1: Write to Commit Log (Disk)

MUST fsync to disk for durability

-- Commit log write performance: HDD: 10-20ms per write ← DISASTER! SATA SSD: 1-2ms ← Acceptable NVMe SSD: 0.1-0.5ms ← BEST! This is the write bottleneck!

Step 2: Write to Memtable (Memory)

Fast in-memory operation

-- Memtable write: 0.01ms (fast!) // But can't acknowledge until commit log done!

Step 3: Acknowledge to Client

Total time = commit log time + network

-- Total write latency: NVMe: 0.5ms + 1ms network = 1.5ms ✅ SATA: 2ms + 1ms network = 3ms ✅ HDD: 15ms + 1ms network = 16ms ❌

Separate Commit Log Disk

❌

Same Disk (Bad)

  • Problem: Commit log + data + compaction
  • Contention: All fighting for I/O
  • Seeks: Disk head jumping around
  • Result: 10-20ms writes
  • Throughput: 1K-5K writes/sec

Never use same disk!

✅

Separate Disks (Good)

  • Disk 1 (NVMe): Commit log only
  • Disk 2 (SATA): Data + compaction
  • No Contention: Isolated I/O
  • Result: 0.5-2ms writes
  • Throughput: 50K-100K writes/sec

2-5x faster writes!

Commit Log Configuration

-- cassandra.yaml configuration # Separate NVMe disk for commit log commitlog_directory: /mnt/nvme/cassandra/commitlog # Commit log sync (CRITICAL!) commitlog_sync: periodic commitlog_sync_period_in_ms: 10000 # Default: 10 seconds # For low-latency writes: commitlog_sync: batch commitlog_sync_batch_window_in_ms: 2 # 2ms batching # Commit log segment size commitlog_segment_size_in_mb: 32 # Default, works well # Compression (optional, saves space) commitlog_compression: - class_name: LZ4Compressor parameters: -

Commit Log Sizing

How Much Space for Commit Log?

-- Formula: commit_log_size = write_rate × flush_interval -- Example 1: 10K writes/sec, 5KB avg 10,000 writes/sec × 5KB = 50MB/sec Memtable flush every 10 minutes = 600 sec 50MB/sec × 600sec = 30GB needed -- Recommendation: 2-3x for safety Provision: 100GB commit log disk -- Example 2: High write rate 100K writes/sec × 2KB = 200MB/sec 200MB × 600sec = 120GB Provision: 250-500GB NVMe

💾 Data Disk Optimization

Optimizing SSTable storage and compaction!

RAID Configuration

❌

RAID 5/6 (Bad)

  • Problem: Write penalty (4x I/O)
  • Parity: Calculates parity on every write
  • Performance: 75% slower writes
  • Why: Read-modify-write cycle
  • Use Case: None for Cassandra

Never use RAID 5/6!

⚠️

RAID 1/10 (OK)

  • Mirroring: Data written to 2+ disks
  • Performance: Good (no write penalty)
  • Capacity: 50% (2 disks = 1 disk space)
  • Why OK: Cassandra has replication
  • Use Case: If must have RAID

Acceptable but expensive

✅

RAID 0 or JBOD (Best)

  • RAID 0: Stripe across disks
  • JBOD: Each disk separate
  • Performance: Full disk speed
  • Capacity: 100% (4 disks = 4x space)
  • Why Best: Cassandra handles failures

Recommended for Cassandra!

Why RAID 0/JBOD for Cassandra?

Cassandra's Built-In Fault Tolerance

Cassandra has its own replication (RF=3 means 3 copies). Disk failure just means one replica is down. Other replicas serve requests. RAID adds overhead for redundancy you already have!

-- With RF=3: // Data replicated across 3 nodes // Disk failure on one node? Other 2 nodes still serve! // RAID redundancy = waste of performance -- RAID 0 advantages: // 1. Full performance (no parity overhead) // 2. Full capacity (no mirroring cost) // 3. Simpler (less hardware complexity) // 4. Cheaper (more usable space)

File System Choice

-- Recommended: XFS or ext4 # XFS (preferred for large files) mkfs.xfs -f -L cassandra_data /dev/sdb mount -o noatime,nodiratime /dev/sdb /var/lib/cassandra/data # ext4 (also good) mkfs.ext4 -L cassandra_data /dev/sdb mount -o noatime,nodiratime,data=writeback /dev/sdb /var/lib/cassandra/data # Mount options: # - noatime: Don't update access times (faster!) # - nodiratime: Don't update directory access times # - data=writeback (ext4): Better write performance # Add to /etc/fstab: /dev/sdb /var/lib/cassandra/data xfs noatime,nodiratime 0 0

Data Disk Sizing

Calculate Required Space

-- Formula: disk_space = (total_data / RF) × space_amplification -- Example: 10TB dataset, RF=3 Per node: 10TB / 3 = 3.33TB raw data -- Space amplification factors: Compaction overhead: 1.5x (STCS) or 1.1x (LCS/TWCS) Snapshots: 1.2x Safety margin: 1.2x Total: 3.33TB × 1.5 × 1.2 × 1.2 = 7.2TB -- Provision: 8TB per node

Multiple Data Disks

-- cassandra.yaml with multiple disks: data_file_directories: - /mnt/ssd1/cassandra/data - /mnt/ssd2/cassandra/data - /mnt/ssd3/cassandra/data -- Cassandra distributes SSTables across disks // Better I/O parallelization // Example: 3× 4TB SSDs = 12TB total

⚙️ I/O Scheduler & Kernel Tuning

OS-level optimizations for disk performance!

I/O Scheduler Selection

❌

CFQ (Bad for SSDs)

  • Design: For HDDs (seek optimization)
  • Problem: Adds latency for SSDs
  • Overhead: Unnecessary queueing
  • Result: 20-30% slower
✅

deadline (Good)

  • Design: Low latency
  • Works: Well with SSDs
  • Overhead: Minimal
  • Result: Good performance
⭐

noop/none (Best)

  • Design: No scheduling
  • Perfect: For NVMe SSDs
  • Overhead: Zero
  • Result: Maximum performance

Configure I/O Scheduler

-- Check current scheduler cat /sys/block/sda/queue/scheduler // [mq-deadline] kyber bfq none -- Set to 'none' for NVMe SSDs echo none | sudo tee /sys/block/nvme0n1/queue/scheduler -- Set to 'deadline' for SATA SSDs echo mq-deadline | sudo tee /sys/block/sda/queue/scheduler -- Make permanent (add to /etc/rc.local or udev rules): echo 'ACTION=="add|change", KERNEL=="nvme[0-9]n[0-9]", ATTR{queue/scheduler}="none"' | \ sudo tee /etc/udev/rules.d/60-scheduler.rules echo 'ACTION=="add|change", KERNEL=="sd[a-z]", ATTR{queue/scheduler}="mq-deadline"' | \ sudo tee -a /etc/udev/rules.d/60-scheduler.rules

Kernel Parameters

-- /etc/sysctl.conf optimizations # Increase read-ahead for sequential I/O vm.min_free_kbytes = 1000000 vm.vfs_cache_pressure = 50 # Dirty page writeback (for better write performance) vm.dirty_ratio = 40 # % of RAM for dirty pages vm.dirty_background_ratio = 10 # When background writeback starts vm.dirty_expire_centisecs = 3000 # 30 seconds vm.dirty_writeback_centisecs = 500 # 5 seconds # Swap (disable or minimize) vm.swappiness = 1 # Emergency only # Apply changes sudo sysctl -p

Read-Ahead Settings

-- Check current read-ahead sudo blockdev --getra /dev/sda // 256 (default, too small for SSDs) -- Set read-ahead for better sequential reads sudo blockdev --setra 8192 /dev/sda # 8192 = 4MB sudo blockdev --setra 8192 /dev/nvme0n1 -- Make permanent (add to /etc/rc.local): echo '/sbin/blockdev --setra 8192 /dev/sda' | sudo tee -a /etc/rc.local echo '/sbin/blockdev --setra 8192 /dev/nvme0n1' | sudo tee -a /etc/rc.local

📊 Disk Monitoring

Essential commands to monitor disk performance!

Key Disk Metrics

-- iostat: Comprehensive disk stats iostat -x 1 /* Device r/s w/s rMB/s wMB/s %util await sda 120 450 15.2 67.3 45.2 8.4 nvme0n1 10 2500 1.2 125.0 32.1 0.3 Key metrics: - r/s, w/s: Reads/writes per second - %util: Disk utilization (< 80% good) - await: Average wait time in ms (< 10ms good) */ -- iotop: Per-process I/O sudo iotop -o -- Watch disk stats in real-time watch -n 1 'iostat -x'

Health Thresholds

Metric Healthy Warning Critical
Disk Utilization < 60% ✅ 60-80% ⚠️ > 80% ❌
Await (latency) < 5ms ✅ 5-10ms ⚠️ > 20ms ❌
Queue Size < 10 ✅ 10-50 ⚠️ > 100 ❌
Space Used < 70% ✅ 70-85% ⚠️ > 85% ❌

Disk Space Monitoring

-- Check disk space df -h /* Filesystem Size Used Avail Use% Mounted on /dev/sda1 4.0T 2.1T 1.9T 53% /var/lib/cassandra/data /dev/nvme0n1 500G 120G 380G 24% /var/lib/cassandra/commitlog */ -- Find large directories du -sh /var/lib/cassandra/data/* -- Cassandra-specific space usage nodetool status // Shows "Load" per node nodetool tablestats // Shows space per table

SSD Health Monitoring

-- Check SSD SMART stats sudo smartctl -a /dev/sda -- Key metrics to watch: /* Wear_Leveling_Count: 100 → 0 (lifespan indicator) Total_LBAs_Written: Track write volume Reallocated_Sector_Ct: Should be 0 */ -- NVMe health sudo nvme smart-log /dev/nvme0n1 -- Monitor write amplification // data_written / host_writes = amplification // Should be < 5x for good wear

💼 Interview Questions & Expert Answers

Master disk optimization for your interview!

1 Why are SSDs mandatory for Cassandra? What happens if you use HDDs? ▼

Answer: SSDs provide 100-1000x more IOPS (50K vs 100) and 100x lower latency (0.5ms vs 15ms) than HDDs. Using HDDs causes write latency of 500ms+ instead of 2ms, making Cassandra unusable for production workloads.

Performance Comparison:

Metric HDD SSD Difference
IOPS 100-200 50,000-500,000 500-5000x faster
Latency 10-20ms 0.1-1ms 10-200x faster
Sequential 150 MB/s 500-3500 MB/s 3-23x faster

Why Cassandra Needs High IOPS:

  • Commit Log: Every write requires fsync (disk sync)
  • Compaction: Rewrites data constantly (background I/O)
  • Reads: May need to read from 5-10 SSTables
  • Repairs: Stream data across network and disk

Real-World Impact:

-- Write performance comparison: // HDD (100 IOPS): Write rate: 100 writes/sec max Latency: 500ms+ (queue builds up) Timeouts: 90% Result: Unusable! ❌ // SSD (50,000 IOPS): Write rate: 100K writes/sec Latency: 2ms Timeouts: 0% Result: Production-ready! ✅

Key Takeaway: SSDs aren't a nice-to-have, they're absolutely mandatory. HDDs make Cassandra 100-1000x slower and unusable for any real workload!

2 Should you put the commit log on the same disk as data? Why or why not? ▼

Answer: No! Commit log should be on a separate NVMe SSD. When on the same disk, commit log competes with reads, compaction, and SSTable writes, causing I/O contention and 2-5x slower writes. Separate disk = isolated I/O paths.

Problem with Same Disk:

-- Single disk scenario: /dev/sda (SATA SSD - 100K IOPS) // Competing operations: 1. Commit log writes: 50K IOPS 2. SSTable reads: 20K IOPS 3. Compaction: 20K IOPS 4. Memtable flushes: 10K IOPS ───────────────────────── Total demand: 100K IOPS Result: - Disk 100% saturated - Queue builds up - Commit log waits for compaction - Write latency: 10-20ms ❌

Solution: Separate Disks

-- Optimized setup: // Disk 1: NVMe SSD (500K IOPS) /dev/nvme0n1 → Commit log ONLY Usage: 50K IOPS (10% utilization) Latency: 0.5ms ✅ // Disk 2-3: SATA SSDs (100K IOPS each) /dev/sda, /dev/sdb → Data + compaction Usage: 50K IOPS total (25% each) Latency: 2-5ms ✅ Result: - No I/O contention - Predictable performance - Write latency: 0.5-2ms ✅ - 2-5x faster writes!

Performance Gain:

Configuration Write Latency Throughput
Same disk (SATA) 10-20ms 5K writes/sec
Separate (NVMe + SATA) 0.5-2ms 50K+ writes/sec

Cost-Benefit:

  • 💰 Cost: $150 for 500GB NVMe SSD
  • ⚡ Benefit: 2-5x faster writes, 10x more throughput
  • 📊 ROI: Pays for itself immediately in performance

Key Takeaway: Commit log on separate NVMe SSD is one of the highest ROI optimizations for Cassandra! Always separate commit log and data disks.

3 Why is RAID 0 or JBOD recommended for Cassandra instead of RAID 5/6? ▼

Answer: RAID 5/6 has 4x write penalty (read-modify-write for parity), making writes 75% slower. Cassandra already has replication (RF=3), so RAID redundancy is unnecessary overhead. RAID 0/JBOD provides full performance and capacity without duplicate redundancy.

RAID 5/6 Write Penalty:

-- RAID 5 write process (4 I/O operations): 1. Read old data block 2. Read old parity block 3. Calculate new parity 4. Write new data block 5. Write new parity block Result: 1 logical write = 4 physical I/O! -- Performance impact: SATA SSD without RAID: 100K write IOPS SATA SSD with RAID 5: 25K write IOPS (75% slower!)

Cassandra's Built-In Redundancy:

-- With RF=3 (replication factor 3): // Data automatically replicated to 3 nodes Node 1: Copy A Node 2: Copy B Node 3: Copy C // If Node 1 disk fails: - Node 1 goes down - Nodes 2 and 3 still serve data - Replace disk, run repair - Data restored from other nodes Result: RAID redundancy is REDUNDANT!

RAID Configuration Comparison:

RAID Level Performance Capacity Recommendation
RAID 5/6 25% speed (4x penalty) 75% (1 disk parity) ❌ Never use
RAID 10 100% speed 50% (mirrored) ⚠️ Acceptable but costly
RAID 0 100%+ speed (striped) 100% ✅ Recommended
JBOD 100% speed 100% ✅ Also good

Why RAID 0 or JBOD Works:

  • ✅ Full Performance: No parity overhead
  • ✅ Full Capacity: No mirroring waste
  • ✅ Cassandra Handles Failures: RF=3 provides redundancy
  • ✅ Simpler: Less hardware complexity
  • ✅ Cheaper: More usable space per dollar

Key Takeaway: Don't add RAID redundancy when Cassandra already provides it! Use RAID 0 or JBOD for maximum performance and capacity.

4 What disk metrics would you monitor to identify performance problems? ▼

Answer: Monitor disk utilization (< 80%), await latency (< 5ms), queue size (< 10), IOPS (vs capacity), and disk space (< 70% full). High utilization or latency indicates I/O bottleneck requiring faster disks or reducing load.

Key Metrics to Track:

1. Disk Utilization (%util)

iostat -x 1 /* Device %util sda 45.2% ← Good (< 60%) sdb 87.5% ← Warning (60-80%) sdc 98.2% ← Critical (> 80%) */ -- What it means: < 60%: Healthy, plenty of headroom 60-80%: Warning, monitor closely > 80%: Critical, I/O bottleneck!

2. Await (Average Wait Time)

/* Device await nvme0n1 0.4ms ← Excellent (< 1ms) sda 3.2ms ← Good (< 5ms) sdb 15.8ms ← Bad (> 10ms) */ -- Indicates: < 1ms: NVMe SSD, no contention < 5ms: SATA SSD, healthy > 10ms: Slow disk or saturation > 20ms: HDD or severe problem

3. Queue Size (avgqu-sz)

/* Device avgqu-sz sda 2.4 ← Good (< 10) sdb 45.7 ← Warning (10-50) sdc 247.2 ← Critical (> 100) */ -- Interpretation: < 10: Normal queue depth 10-50: Getting backed up > 100: Severe backlog, can't keep up!

4. IOPS (Operations/Sec)

/* Device r/s w/s Total nvme0n1 100 2500 2600 vs 500K capacity (0.5%) sda 1200 4500 5700 vs 100K capacity (5.7%) sdb 25000 75000 100K vs 100K capacity (100%!) */ -- Check if approaching disk capacity

5. Disk Space

df -h /* Filesystem Size Used Avail Use% /dev/sda1 4.0T 2.8T 1.2T 70% ← Warning /dev/sdb1 4.0T 3.4T 600G 85% ← Critical! */ -- Thresholds: < 70%: Healthy 70-85%: Warning, plan for more space > 85%: Critical, compaction may fail!

Monitoring Dashboard:

-- Comprehensive monitoring script: #!/bin/bash while true; do clear echo "=== Disk Performance ===" iostat -x 1 1 | grep -E "Device|sd|nvme" echo "" echo "=== Disk Space ===" df -h | grep -E "Filesystem|cassandra" echo "" echo "=== Top I/O Processes ===" iotop -b -n 1 | head -n 15 sleep 5 done

Alert Thresholds:

  • ⚠️ WARNING: %util > 60%, await > 5ms, space > 70%
  • ❌ CRITICAL: %util > 80%, await > 10ms, space > 85%
5 How would you size disks for a Cassandra cluster storing 10TB of data with RF=3? ▼

Answer: Per-node calculation: (10TB / nodes / RF) × space_amplification × safety_margin. For 10 nodes: (10TB / 10 / 3) × 1.5 (STCS) × 1.5 (snapshots+margin) = 750GB per node. Provision 1TB per node for headroom.

Step-by-Step Sizing:

Step 1: Calculate Per-Node Raw Data

-- Given: Total data: 10TB Replication factor (RF): 3 Number of nodes: 10 -- Calculation: Per-node raw = Total / Nodes Per-node raw = 10TB / 10 = 1TB // Note: RF doesn't affect per-node calculation! // RF=3 means cluster stores 30TB total (3 copies) // But each node still stores 1TB

Step 2: Account for Space Amplification

-- Space amplification factors: # 1. Compaction overhead: STCS: 1.5x (needs temp space during compaction) LCS: 1.1x (better space efficiency) TWCS: 1.1x (no cross-window compaction) # 2. Snapshots (for backups): 1.2x (20% for snapshots) # 3. Safety margin: 1.2x (20% headroom) -- Total amplification (STCS): 1TB × 1.5 × 1.2 × 1.2 = 2.16TB per node

Step 3: Provision Disks

-- Option 1: Conservative (STCS) Per node: 2.16TB needed Provision: 2.5TB or 3TB per node -- Option 2: Efficient (LCS/TWCS) 1TB × 1.1 × 1.2 × 1.2 = 1.58TB per node Provision: 2TB per node -- Recommended setup per node: Commit log: 250GB NVMe SSD Data: 2TB SATA SSD (2× 1TB in RAID 0)

Step 4: Validation

-- Cluster totals: 10 nodes × 2TB = 20TB total provisioned 10TB raw data + 10TB overhead/safety -- Cluster with RF=3: Logical capacity: 10TB (3 copies) Physical storage: 30TB (3× 10TB) Provisioned: 20TB (10 nodes × 2TB) // Wait, 20TB < 30TB? // No! Each node stores 1TB of unique data // 10 nodes × 1TB = 10TB unique // × RF=3 = 30TB total (distributed) // Each node: 1TB × 2x overhead = 2TB ✅

Quick Formula:

per_node_disk = (total_data / num_nodes) × amplification -- Amplification factors: STCS: 2.0-2.5x LCS/TWCS: 1.5-2.0x -- Example: (10TB / 10 nodes) × 2x = 2TB per node

Key Takeaway: Always provision 1.5-2.5x raw data size per node to account for compaction, snapshots, and safety margin. Never run disks > 85% full!

🎓 Chapter Summary: Disk Optimization Mastery

You now understand disk optimization at a production level!

Critical Rules:

  • 💿 SSDs Mandatory: 100-1000x faster than HDDs
  • 📝 Separate Commit Log: NVMe for commit log, SATA for data
  • 🔀 RAID 0/JBOD: No RAID 5/6 (Cassandra has replication!)
  • ⚙️ I/O Scheduler: 'none' for NVMe, 'deadline' for SATA
  • 📊 Monitor: %util < 80%, await < 5ms, space < 70%

Recommended Setup:

Disk 1: 250-500GB NVMe → Commit log Disk 2-3: 2-4TB SATA SSDs → Data (RAID 0)

Performance Gain:

HDD: 500ms writes, 1K/sec → SSD: 2ms writes, 100K/sec (250x faster!)

Remember Jake: SSDs are not optional! 🚀

Advertisement

Responsive Ad