Network Tuning in Cassandra
Keep your cluster talking fast! Network tuning makes sure data moves quickly and smoothly between nodes, helping Cassandra stay responsive even under heavy traffic.
๐ The Story: Sarah's 10Gbps Network Running at 100Mbps
Sarah had a beautiful setup: 10 Gbps network cards, fiber connections, enterprise switches. But streaming repairs were taking 10 hours instead of 10 minutes. Network throughput showed 100 Mbps instead of 10 Gbps. 100x slower than hardware capability! Here's what was wrong...
๐ฑ The Problem: Default TCP Settings
Sarah's Hardware:
- ๐ NICs: 10 Gbps Intel X710 (enterprise grade!)
- ๐ Cables: Fiber optic (perfect!)
- ๐ก Switch: Cisco 10Gbps (no issues!)
- ๐ป Latency: 0.1ms between nodes (excellent!)
The Reality:
The Hidden Culprits:
- Tiny TCP Buffers: Default 128KB (need 32MB!)
- Small Window Size: 65KB max (need 16MB+!)
- No TCP Timestamps: Can't scale beyond 1 Gbps
- Wrong Congestion Algorithm: Reno instead of CUBIC
- Connection Tracking Full: Firewall limiting connections
- IRQ Affinity Wrong: All interrupts on CPU0
Why Hardware Didn't Matter:
- ๐พ Buffer Too Small: TCP can't use 10Gbps pipe
- ๐ช Window Scaling Off: Limited to 64KB windows
- โฑ๏ธ No Timestamps: Can't measure RTT accurately
- ๐ฆ Old Congestion Control: Too conservative
- ๐ฅ Conntrack Table: Dropping new connections
๐ The Fix: TCP Tuning for 10Gbps!
Sarah's Optimizations:
The Results:
- โก Throughput: 100 Mbps โ 9.2 Gbps (92x faster!)
- โฑ๏ธ Repair Time: 10 hours โ 12 minutes (50x faster!)
- ๐ Streaming: 98 MB/s โ 1150 MB/s
- ๐ฏ Hardware Utilization: 1% โ 92% (finally!)
- ๐ Cross-DC Replication: Actually works now!
- ๐ฐ ROI: $0 cost, 100x performance gain!
Sarah learned: Network tuning is FREE 100x performance! ๐
๐ Network Fundamentals
Understanding TCP performance and bandwidth-delay product!
๐ฏ Why Network Tuning Matters
Cassandra is extremely network-intensive: streaming repairs, hinted handoffs, read repairs, gossip, client requests. Default Linux TCP settings are optimized for 1 Gbps networks from 2005. Modern 10/25/100 Gbps networks need proper tuning or you'll get 100x slower performance!
Bandwidth-Delay Product (BDP)
TCP Window Scaling
Without Window Scaling (Legacy)
Limited to 64KB window size
With Window Scaling (Modern)
Scales up to 1GB window size
Network Performance Bottlenecks
| Bottleneck | Symptom | Fix |
|---|---|---|
| Small TCP buffers | Low throughput (< 1 Gbps) | Increase tcp_rmem/wmem to 32-128MB |
| No window scaling | Capped at 64KB windows | Enable tcp_window_scaling=1 |
| Wrong congestion algo | Slow ramp-up after packet loss | Use CUBIC or BBR |
| IRQ imbalance | One CPU at 100%, others idle | Configure IRQ affinity/RSS |
| Connection tracking | Connection timeouts, packet drops | Increase nf_conntrack_max |
๐ง TCP Kernel Tuning
Essential kernel parameters for high-performance networking!
Complete TCP Optimization
Parameter Explanations
TCP Buffer Sizes
Most Critical Setting!
Window Scaling
Allows TCP windows > 64KB
Congestion Control
How TCP reacts to packet loss
Connection Tracking
Firewall connection table size
Quick Wins
These three settings alone give 10-100x performance gain:
- tcp_rmem/wmem: 32-128MB (from 6MB default)
- tcp_window_scaling: Enable (often disabled!)
- tcp_congestion_control: CUBIC or BBR (from Reno)
๐๏ธ NIC & Hardware Tuning
Network interface card optimization!
IRQ Affinity & RSS
โ ๏ธ The IRQ Imbalance Problem
By default, all network interrupts go to CPU0. With 10 Gbps traffic, CPU0 hits 100% just handling interrupts while other 31 cores sit idle. This limits throughput to ~1-2 Gbps. Solution: Spread interrupts across all CPUs!
NIC Ring Buffer Sizes
TCP Offloading Features
MTU Size
Standard MTU (1500)
- Size: 1500 bytes
- Compatible: All networks
- Overhead: High (more packets)
- CPU: More interrupts
- Use: Cross-internet, mixed networks
Jumbo Frames (9000)
- Size: 9000 bytes (6x larger!)
- Compatible: Modern datacenter only
- Overhead: Low (fewer packets)
- CPU: Less interrupts (20-30% reduction)
- Use: Same-datacenter Cassandra
โ๏ธ Cassandra Network Configuration
Cassandra-specific network settings!
cassandra.yaml Network Settings
JVM Network Settings
๐ Network Monitoring
Essential commands to monitor network performance!
Network Statistics
TCP Buffer Usage
Connection Tracking
Bandwidth Testing
Health Thresholds
| Metric | Healthy | Warning | Critical |
|---|---|---|---|
| Packet Loss | 0% โ | < 0.1% โ ๏ธ | > 0.5% โ |
| Retransmits | < 0.1% โ | 0.1-1% โ ๏ธ | > 1% โ |
| Bandwidth Util | < 70% โ | 70-85% โ ๏ธ | > 85% โ |
| Latency (same DC) | < 1ms โ | 1-5ms โ ๏ธ | > 10ms โ |
| Conntrack Usage | < 70% โ | 70-90% โ ๏ธ | > 90% โ |
๐ผ Interview Questions & Expert Answers
Master network tuning for your interview!
Answer: Default TCP buffers (6MB) are too small for high bandwidth-delay product networks. With 10 Gbps and 50ms RTT, you need 62.5MB buffers. Solution: Increase tcp_rmem/wmem to 128MB, enable window scaling, and use CUBIC/BBR congestion control.
Diagnosis Steps:
Step 1: Verify Hardware
Step 2: Check TCP Buffer Settings
Step 3: Check Window Scaling
Step 4: Apply Fix
Step 5: Verify Fix
Key Takeaway: TCP buffers must be sized for bandwidth-delay product. Default 6MB is only enough for 1 Gbps with low latency!
Answer: BDP = Bandwidth ร RTT. It represents the amount of data "in flight" on the network. TCP buffer must be โฅ BDP to fully utilize available bandwidth. If buffer < BDP, TCP can't keep the pipe full, limiting throughput to buffer_size / RTT.
Understanding BDP:
Why BDP Matters:
TCP uses a "window" to control how much data can be in flight (sent but not yet acknowledged). The window must be at least BDP size to keep the network pipe full.
Buffer Sizing Guide:
| Scenario | BDP | Buffer Size |
|---|---|---|
| Same DC, 10 Gbps | 2.5 MB | 32 MB (safety margin) |
| Cross-region, 10 Gbps | 62.5 MB | 128 MB (2x BDP) |
| Cross-continent, 10 Gbps | 250 MB | 512 MB (2x BDP) |
Key Takeaway: Always size TCP buffers to 2x BDP. Higher latency = larger buffers needed. Default 6MB only works for low-latency 1 Gbps networks!
Answer: CUBIC is loss-based (reacts to packet loss), while BBR is model-based (measures actual bandwidth and RTT). CUBIC works well for low-latency networks. BBR excels in high-latency, high-bandwidth networks (cross-DC Cassandra). BBR maintains 2-4x higher throughput in lossy networks.
CUBIC (Default on Most Linux):
- Algorithm: Loss-based congestion control
- How it works: Grows window aggressively until packet loss, then reduces by 30%
- Recovery: Fast recovery using cubic function
- Best for: Low-latency networks (< 10ms RTT)
- Problem: Treats packet loss as congestion (not always true!)
BBR (Google's Algorithm):
- Algorithm: Model-based congestion control
- How it works: Measures actual bandwidth and RTT, targets optimal operating point
- Recovery: Doesn't rely on packet loss
- Best for: High-latency (> 10ms RTT), high-bandwidth, or lossy networks
- Advantage: Maintains high throughput even with packet loss
Performance Comparison:
| Scenario | CUBIC | BBR |
|---|---|---|
| Same DC (0.5ms, 0% loss) | 9.8 Gbps โ | 9.7 Gbps โ |
| Cross-region (50ms, 0% loss) | 8.5 Gbps | 9.3 Gbps โ |
| Cross-DC (50ms, 0.1% loss) | 2.3 Gbps โ | 8.9 Gbps โ |
| Lossy network (100ms, 1% loss) | 500 Mbps โ | 7.2 Gbps โ |
Recommendation for Cassandra:
- โ Same Datacenter: CUBIC or BBR (both work well)
- โญ Cross-Datacenter: BBR (2-4x better performance!)
- โญ Multi-region Replication: BBR (essential!)
- โ ๏ธ Requires: Linux kernel 4.9+ for BBR
Key Takeaway: BBR is superior for cross-DC Cassandra replication. It maintains high throughput even with packet loss, making it ideal for WAN connections!
Answer: Check stream_throughput setting (increase from 200 to 400-800 Mbps), verify TCP buffer sizes, check network bandwidth utilization, verify no packet loss/retransmits, and ensure streaming_connections_per_host is adequate (4-16). Also check for IRQ imbalance and enable jumbo frames if possible.
Troubleshooting Steps:
Step 1: Check Cassandra Stream Settings
Step 2: Monitor Actual Bandwidth
Step 3: Verify TCP Settings
Step 4: Check Network Health
Step 5: Optimize Streaming Connections
Step 6: Additional Optimizations
Expected Results:
Key Takeaway: Stream throttling + poor TCP tuning can make repairs 100x slower. Always tune TCP buffers and increase stream_throughput for modern networks!
Answer: Yes for same-datacenter deployments if entire network path supports it. Jumbo frames reduce CPU overhead by 20-30% and improve throughput by 10-15%. But ALL devices (NICs, switches, routers) must support 9000 MTU. Never use for cross-DC or internet traffic.
Benefits of Jumbo Frames:
- โก Less CPU: 6x fewer packets = 20-30% less CPU for networking
- โก Less Interrupts: Fewer packets = fewer NIC interrupts
- โก Better Throughput: 10-15% higher throughput
- โก Lower Latency: Slightly lower per-byte latency
Requirements (ALL must be met):
- โ All Cassandra node NICs support MTU 9000
- โ All switches support jumbo frames
- โ All routers in path support jumbo frames
- โ Same physical datacenter (no internet/WAN)
- โ If ANY device doesn't support โ fragmentation โ worse performance!
Testing for Jumbo Frame Support:
Enabling Jumbo Frames:
When NOT to Use:
| Scenario | Use Jumbo? | Reason |
|---|---|---|
| Same datacenter, modern switches | โ Yes | Full control, 20-30% CPU savings |
| Cross-datacenter (WAN) | โ No | WAN rarely supports, fragmentation |
| Cloud (AWS/Azure/GCP) | โ No | Usually not supported |
| Mixed old/new switches | โ No | Old switches may not support |
| Client traffic (CQL) | โ No | Clients may not support |
Performance Impact:
Key Takeaway: Jumbo frames are worthwhile for same-datacenter Cassandra if ALL network gear supports it. Test thoroughly before enabling. Never use for cross-DC or cloud deployments!
๐ Chapter Summary: Network Tuning Mastery
You now understand network tuning at a production level!
Critical Optimizations:
- ๐ฆ TCP Buffers: 32-128MB (not 6MB default!)
- ๐ช Window Scaling: Enable (allows > 64KB windows)
- ๐ BBR/CUBIC: Use BBR for cross-DC
- โก IRQ Balance: Spread across all CPUs
- ๐ข Conntrack: Increase to 1M connections
Quick Wins:
Performance Gain:
100 Mbps โ 9.4 Gbps (94x faster!) with $0 cost!
Remember Sarah: Network tuning is FREE 100x performance! ๐
Responsive Ad