Advanced Performance Topics

Network Tuning in Cassandra

Keep your cluster talking fast! Network tuning makes sure data moves quickly and smoothly between nodes, helping Cassandra stay responsive even under heavy traffic.

๐Ÿ“– The Story: Sarah's 10Gbps Network Running at 100Mbps

Sarah had a beautiful setup: 10 Gbps network cards, fiber connections, enterprise switches. But streaming repairs were taking 10 hours instead of 10 minutes. Network throughput showed 100 Mbps instead of 10 Gbps. 100x slower than hardware capability! Here's what was wrong...

๐Ÿ˜ฑ The Problem: Default TCP Settings

Sarah's Hardware:

  • ๐ŸŒ NICs: 10 Gbps Intel X710 (enterprise grade!)
  • ๐Ÿ”Œ Cables: Fiber optic (perfect!)
  • ๐Ÿ“ก Switch: Cisco 10Gbps (no issues!)
  • ๐Ÿ’ป Latency: 0.1ms between nodes (excellent!)

The Reality:

-- Repair streaming performance: nodetool repair -pr keyspace -- Expected: 10 minutes (10 Gbps) -- Actual: 10 HOURS! (100 Mbps) -- Network throughput: iftop -i eth0 /* TX: 98.7 Mbps โ† Should be 10,000 Mbps! RX: 102.3 Mbps โ† 100x slower! */

The Hidden Culprits:

  1. Tiny TCP Buffers: Default 128KB (need 32MB!)
  2. Small Window Size: 65KB max (need 16MB+!)
  3. No TCP Timestamps: Can't scale beyond 1 Gbps
  4. Wrong Congestion Algorithm: Reno instead of CUBIC
  5. Connection Tracking Full: Firewall limiting connections
  6. IRQ Affinity Wrong: All interrupts on CPU0
-- Sarah's default TCP settings (terrible!): sysctl net.ipv4.tcp_rmem /* net.ipv4.tcp_rmem = 4096 87380 6291456 โ†‘ Max 6MB receive buffer For 10Gbps ร— 0.1ms latency: BDP = 10Gbps ร— 0.0001s = 1.25 MB (barely fits!) For 10Gbps ร— 10ms latency (cross-DC): BDP = 10Gbps ร— 0.01s = 125 MB (DOESN'T FIT!) Result: Tiny window, throttled throughput! ๐Ÿ’ฅ */

Why Hardware Didn't Matter:

  • ๐Ÿ’พ Buffer Too Small: TCP can't use 10Gbps pipe
  • ๐ŸชŸ Window Scaling Off: Limited to 64KB windows
  • โฑ๏ธ No Timestamps: Can't measure RTT accurately
  • ๐Ÿšฆ Old Congestion Control: Too conservative
  • ๐Ÿ”ฅ Conntrack Table: Dropping new connections

๐Ÿš€ The Fix: TCP Tuning for 10Gbps!

Sarah's Optimizations:

-- /etc/sysctl.conf optimizations # 1. Massive TCP buffers (32MB) net.ipv4.tcp_rmem = 4096 87380 33554432 net.ipv4.tcp_wmem = 4096 65536 33554432 net.core.rmem_max = 33554432 net.core.wmem_max = 33554432 # 2. Enable window scaling (for large windows) net.ipv4.tcp_window_scaling = 1 # 3. Enable TCP timestamps (for high speed) net.ipv4.tcp_timestamps = 1 # 4. Use CUBIC congestion control net.ipv4.tcp_congestion_control = cubic # 5. Increase connection tracking net.netfilter.nf_conntrack_max = 1048576 # Apply sudo sysctl -p

The Results:

  • โšก Throughput: 100 Mbps โ†’ 9.2 Gbps (92x faster!)
  • โฑ๏ธ Repair Time: 10 hours โ†’ 12 minutes (50x faster!)
  • ๐Ÿ“Š Streaming: 98 MB/s โ†’ 1150 MB/s
  • ๐ŸŽฏ Hardware Utilization: 1% โ†’ 92% (finally!)
  • ๐Ÿ˜Š Cross-DC Replication: Actually works now!
  • ๐Ÿ’ฐ ROI: $0 cost, 100x performance gain!

Sarah learned: Network tuning is FREE 100x performance! ๐ŸŽ‰

๐ŸŒ Network Fundamentals

Understanding TCP performance and bandwidth-delay product!

๐ŸŽฏ Why Network Tuning Matters

Cassandra is extremely network-intensive: streaming repairs, hinted handoffs, read repairs, gossip, client requests. Default Linux TCP settings are optimized for 1 Gbps networks from 2005. Modern 10/25/100 Gbps networks need proper tuning or you'll get 100x slower performance!

Bandwidth-Delay Product (BDP)

Bandwidth-Delay Product (BDP) BDP = Bandwidth ร— Round-Trip Time (RTT) Same Datacenter (Low Latency) Bandwidth: 10 Gbps RTT: 0.2ms (200 ฮผs) BDP = 10 Gbps ร— 0.0002s BDP = 2.5 MB Cross-Datacenter (High Latency) Bandwidth: 10 Gbps RTT: 50ms BDP = 10 Gbps ร— 0.05s BDP = 62.5 MB TCP Buffer Must Be โ‰ฅ BDP โŒ Default (6MB buffer) Same DC: 6MB > 2.5MB โœ… Cross-DC: 6MB < 62.5MB โŒ Result: 10% utilization! โœ… Tuned (128MB buffer) Same DC: 128MB > 2.5MB โœ… Cross-DC: 128MB > 62.5MB โœ… Result: 95%+ utilization!

TCP Window Scaling

Without Window Scaling (Legacy)

Limited to 64KB window size

-- Maximum throughput without scaling: Max_throughput = Window_size / RTT -- Example: 10ms RTT Max = 64KB / 0.01s = 6.4 MB/s = 51 Mbps Even with 10 Gbps NIC, stuck at 51 Mbps! โŒ

With Window Scaling (Modern)

Scales up to 1GB window size

-- With window scaling enabled: net.ipv4.tcp_window_scaling = 1 -- Can use up to 1GB windows! Max = 128MB / 0.01s = 12.8 GB/s = 102 Gbps โœ… Now can fully utilize 10/25/100 Gbps links!

Network Performance Bottlenecks

Bottleneck Symptom Fix
Small TCP buffers Low throughput (< 1 Gbps) Increase tcp_rmem/wmem to 32-128MB
No window scaling Capped at 64KB windows Enable tcp_window_scaling=1
Wrong congestion algo Slow ramp-up after packet loss Use CUBIC or BBR
IRQ imbalance One CPU at 100%, others idle Configure IRQ affinity/RSS
Connection tracking Connection timeouts, packet drops Increase nf_conntrack_max

๐Ÿ”ง TCP Kernel Tuning

Essential kernel parameters for high-performance networking!

Complete TCP Optimization

#################################### # /etc/sysctl.conf - NETWORK TUNING # For 10 Gbps+ Cassandra clusters #################################### # ========== TCP BUFFER SIZES ========== # Format: min default max (in bytes) # TCP read buffers (32MB max) net.ipv4.tcp_rmem = 4096 87380 33554432 # 4KB min, 87KB default, 32MB max # TCP write buffers (32MB max) net.ipv4.tcp_wmem = 4096 65536 33554432 # 4KB min, 64KB default, 32MB max # Maximum socket buffer sizes net.core.rmem_max = 33554432 # 32MB net.core.wmem_max = 33554432 # 32MB # For cross-DC or very high bandwidth, use 128MB: # net.ipv4.tcp_rmem = 4096 87380 134217728 # net.ipv4.tcp_wmem = 4096 65536 134217728 # net.core.rmem_max = 134217728 # net.core.wmem_max = 134217728 # ========== TCP FEATURES ========== # Enable window scaling (CRITICAL!) net.ipv4.tcp_window_scaling = 1 # Enable timestamps (for high bandwidth) net.ipv4.tcp_timestamps = 1 # Enable selective acknowledgments net.ipv4.tcp_sack = 1 # Disable slow start after idle net.ipv4.tcp_slow_start_after_idle = 0 # ========== CONGESTION CONTROL ========== # Use CUBIC (modern, aggressive) net.ipv4.tcp_congestion_control = cubic # Or BBR (Google's algorithm, even better for high BDP) # net.ipv4.tcp_congestion_control = bbr # net.core.default_qdisc = fq # ========== SOCKET QUEUE SIZES ========== # Max pending connections net.core.somaxconn = 4096 # Max backlog queue size net.core.netdev_max_backlog = 16384 # ========== CONNECTION TRACKING ========== # Increase connection tracking table net.netfilter.nf_conntrack_max = 1048576 net.nf_conntrack_max = 1048576 # Timeout for established connections net.netfilter.nf_conntrack_tcp_timeout_established = 86400 # ========== PORT RANGE ========== # Increase available local ports net.ipv4.ip_local_port_range = 10000 65535 # ========== TCP KEEPALIVE ========== # Detect dead connections faster net.ipv4.tcp_keepalive_time = 600 # 10 minutes net.ipv4.tcp_keepalive_intvl = 60 # 60 seconds net.ipv4.tcp_keepalive_probes = 3 # 3 probes # ========== APPLY CHANGES ========== # Run: sudo sysctl -p

Parameter Explanations

TCP Buffer Sizes

Most Critical Setting!

-- How to calculate needed buffer size: buffer_size = bandwidth ร— RTT -- Examples: 10 Gbps ร— 0.2ms = 2.5 MB (same DC) 10 Gbps ร— 10ms = 125 MB (cross-region) 10 Gbps ร— 100ms = 1250 MB (cross-continent) -- Recommendations: Same DC (10 Gbps): 32MB buffers Cross-region: 128MB buffers Cross-continent: 256MB+ buffers

Window Scaling

Allows TCP windows > 64KB

net.ipv4.tcp_window_scaling = 1 -- Without this: // Maximum window: 64KB // Maximum throughput: 64KB / RTT // With 10ms RTT: 51 Mbps max! -- With window scaling: // Maximum window: 1GB // Can fully use 100 Gbps networks!

Congestion Control

How TCP reacts to packet loss

-- Available algorithms: # Reno (old, slow recovery) // Default on old systems // Cuts window by 50% on loss // Very conservative # CUBIC (modern, default on most Linux) net.ipv4.tcp_congestion_control = cubic // Fast recovery after loss // Good for high BDP networks # BBR (Google's algorithm, best for Cassandra) net.ipv4.tcp_congestion_control = bbr net.core.default_qdisc = fq // Measures actual bandwidth // Excellent for cross-DC // Requires Linux 4.9+

Connection Tracking

Firewall connection table size

-- Default is often too small: cat /proc/sys/net/netfilter/nf_conntrack_max // 65536 (only 65K connections!) -- Cassandra needs many connections: // - Client connections: 1000s // - Inter-node: 100s per node // - Streaming: 100s // Total: 10K-100K+ connections -- Increase to 1M: net.netfilter.nf_conntrack_max = 1048576

Quick Wins

These three settings alone give 10-100x performance gain:

  1. tcp_rmem/wmem: 32-128MB (from 6MB default)
  2. tcp_window_scaling: Enable (often disabled!)
  3. tcp_congestion_control: CUBIC or BBR (from Reno)

๐ŸŽ›๏ธ NIC & Hardware Tuning

Network interface card optimization!

IRQ Affinity & RSS

โš ๏ธ The IRQ Imbalance Problem

By default, all network interrupts go to CPU0. With 10 Gbps traffic, CPU0 hits 100% just handling interrupts while other 31 cores sit idle. This limits throughput to ~1-2 Gbps. Solution: Spread interrupts across all CPUs!

-- Check current IRQ distribution: cat /proc/interrupts | grep eth0 /* CPU0 CPU1 CPU2 CPU3 ... 123: 15M 120 95 87 ... eth0-TxRx-0 124: 14M 105 88 92 ... eth0-TxRx-1 โ†‘ โ†‘ Unbalanced! All on CPU0! */ -- Enable RSS (Receive Side Scaling): # Check RSS support ethtool -l eth0 /* Combined: 8 โ† Number of RX/TX queues */ -- Configure irqbalance (automatic distribution) sudo systemctl enable irqbalance sudo systemctl start irqbalance -- OR manual IRQ affinity (advanced): # Spread eth0 interrupts across CPUs 0-7 for irq in $(grep eth0 /proc/interrupts | cut -d: -f1); do echo 0-7 > /proc/irq/$irq/smp_affinity_list done

NIC Ring Buffer Sizes

-- Check current ring buffer sizes: ethtool -g eth0 /* Ring parameters for eth0: Pre-set maximums: RX: 4096 TX: 4096 Current settings: RX: 512 โ† Too small! TX: 512 โ† Too small! */ -- Increase to maximum: ethtool -G eth0 rx 4096 tx 4096 -- Make permanent (add to /etc/rc.local): echo 'ethtool -G eth0 rx 4096 tx 4096' >> /etc/rc.local

TCP Offloading Features

-- Check current offloading features: ethtool -k eth0 -- Enable important features: ethtool -K eth0 tso on # TCP Segmentation Offload ethtool -K eth0 gso on # Generic Segmentation Offload ethtool -K eth0 gro on # Generic Receive Offload ethtool -K eth0 lro off # Large Receive Offload (disable!) -- Why disable LRO? // LRO can cause issues with bridging/routing // Use GRO instead (more compatible)

MTU Size

๐Ÿ“ฆ

Standard MTU (1500)

  • Size: 1500 bytes
  • Compatible: All networks
  • Overhead: High (more packets)
  • CPU: More interrupts
  • Use: Cross-internet, mixed networks
๐Ÿ“ฆ๐Ÿ“ฆ๐Ÿ“ฆ

Jumbo Frames (9000)

  • Size: 9000 bytes (6x larger!)
  • Compatible: Modern datacenter only
  • Overhead: Low (fewer packets)
  • CPU: Less interrupts (20-30% reduction)
  • Use: Same-datacenter Cassandra
-- Enable jumbo frames (datacenter only!): ifconfig eth0 mtu 9000 -- Make permanent (/etc/network/interfaces): auto eth0 iface eth0 inet static address 192.168.1.10 netmask 255.255.255.0 mtu 9000 -- Verify: ip link show eth0 | grep mtu // mtu 9000 -- IMPORTANT: ALL devices in path must support 9000 MTU! // - All Cassandra nodes // - All switches // - All routers // If any device doesn't support it, packets will be fragmented!

โš™๏ธ Cassandra Network Configuration

Cassandra-specific network settings!

cassandra.yaml Network Settings

-- cassandra.yaml network configuration # ========== STREAMING ========== # Stream throughput (MB/sec) stream_throughput_outbound_megabits_per_sec: 400 # Default: 200 Mbps (too low for 10 Gbps!) # Set to 400-800 for 10 Gbps # Set to 2000-4000 for 25 Gbps # Inter-DC stream throughput inter_dc_stream_throughput_outbound_megabits_per_sec: 200 # Lower for cross-DC to avoid saturating WAN # ========== CONCURRENT OPERATIONS ========== # Concurrent streaming connections streaming_connections_per_host: 4 # Increase to 8-16 for faster repairs # Concurrent compactors concurrent_compactors: 4 # Set to num_cores / 4 # ========== INTERNODE COMPRESSION ========== # Compression for inter-node traffic internode_compression: dc # Options: # - all: Compress all inter-node (high CPU, low bandwidth) # - dc: Compress only cross-DC (recommended!) # - none: No compression (if bandwidth plentiful) # ========== TIMEOUTS ========== # Read request timeout read_request_timeout_in_ms: 5000 # Range request timeout (scans) range_request_timeout_in_ms: 10000 # Write request timeout write_request_timeout_in_ms: 2000 # Request timeout (streaming, etc) request_timeout_in_ms: 10000 # ========== RPC ========== # Native transport (CQL) threads native_transport_max_threads: 128 # Increase for high client connection count # Max concurrent connections native_transport_max_concurrent_connections: -1 # -1 = unlimited (monitor with nodetool)

JVM Network Settings

-- jvm.options network tuning # Prefer IPv4 -Djava.net.preferIPv4Stack=true # TCP settings for JVM -Dcom.sun.management.jmxremote.rmi.port=7199 # Increase Netty buffer sizes (for streaming) -Dcassandra.netty.buffer.size=16384 # Network thread pool -Dcassandra.max_queued_native_transport_requests=4096

๐Ÿ“Š Network Monitoring

Essential commands to monitor network performance!

Network Statistics

-- Real-time bandwidth monitoring: iftop -i eth0 /* 12.5Mb 25.0Mb 37.5Mb 50.0Mb node1 => node2 9.2Gb โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ node1 => node3 8.7Gb โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ */ -- Network throughput by interface: nload eth0 -- Detailed network stats: netstat -s | grep -i "segments retransmitted" // High retransmits = packet loss -- TCP connection stats: ss -s /* Total: 892 TCP: 742 (estab 689, closed 12, orphaned 0) */ -- Check for packet drops: ifconfig eth0 | grep -E "RX|TX" /* RX packets: 45678234 bytes: 67890123456 RX errors: 0 dropped: 0 โ† Should be 0! TX packets: 43567123 bytes: 65432198765 TX errors: 0 dropped: 0 โ† Should be 0! */

TCP Buffer Usage

-- Check actual TCP buffer usage: ss -tm /* State Recv-Q Send-Q Local:Port Peer:Port ESTAB 0 0 10.0.1.5:9042 10.0.1.6:45678 skmem:(r0,rb8388608,t0,tb8388608,...) โ†‘ โ†‘ Receive buffer Transmit buffer 8MB allocated 8MB allocated */ -- Monitor for buffer exhaustion: cat /proc/net/sockstat /* TCP: inuse 689 orphan 0 tw 12 alloc 702 mem 42567 โ†‘ Pages used by TCP If mem grows continuously โ†’ buffer exhaustion! */

Connection Tracking

-- Check connection tracking usage: cat /proc/sys/net/netfilter/nf_conntrack_count // 45678 cat /proc/sys/net/netfilter/nf_conntrack_max // 1048576 -- If count approaches max โ†’ increase max! -- Connection tracking stats: conntrack -S /* cpu=0 found=0 invalid=0 ignore=0 insert=0 entries: 45678 โ† Current entries max entries: 1048576 */

Bandwidth Testing

-- Test bandwidth between nodes: # On receiver node: iperf3 -s # On sender node: iperf3 -c -t 30 /* [ ID] Interval Transfer Bandwidth [ 4] 0.00-30.00 sec 32.8 GBytes 9.41 Gbits/sec โ†‘ Should be close to 10 Gbps! */ -- Test with parallel streams (simulates Cassandra): iperf3 -c -P 10 -t 30 # -P 10 = 10 parallel streams

Health Thresholds

Metric Healthy Warning Critical
Packet Loss 0% โœ… < 0.1% โš ๏ธ > 0.5% โŒ
Retransmits < 0.1% โœ… 0.1-1% โš ๏ธ > 1% โŒ
Bandwidth Util < 70% โœ… 70-85% โš ๏ธ > 85% โŒ
Latency (same DC) < 1ms โœ… 1-5ms โš ๏ธ > 10ms โŒ
Conntrack Usage < 70% โœ… 70-90% โš ๏ธ > 90% โŒ

๐Ÿ’ผ Interview Questions & Expert Answers

Master network tuning for your interview!

1 Why would a 10 Gbps network only achieve 100 Mbps throughput? How would you diagnose and fix it? โ–ผ

Answer: Default TCP buffers (6MB) are too small for high bandwidth-delay product networks. With 10 Gbps and 50ms RTT, you need 62.5MB buffers. Solution: Increase tcp_rmem/wmem to 128MB, enable window scaling, and use CUBIC/BBR congestion control.

Diagnosis Steps:

Step 1: Verify Hardware

-- Check link speed: ethtool eth0 | grep Speed // Speed: 10000Mb/s โœ… Hardware OK -- Check for errors: ifconfig eth0 | grep errors // RX errors: 0, TX errors: 0 โœ…

Step 2: Check TCP Buffer Settings

sysctl net.ipv4.tcp_rmem /* net.ipv4.tcp_rmem = 4096 87380 6291456 โ†‘ Max 6MB โŒ TOO SMALL! */ -- Calculate needed buffer: BDP = Bandwidth ร— RTT BDP = 10 Gbps ร— 50ms = 62.5 MB 6MB buffer < 62.5MB needed = BOTTLENECK!

Step 3: Check Window Scaling

sysctl net.ipv4.tcp_window_scaling // 0 โŒ DISABLED! Without window scaling: // Max window: 64KB // Max throughput = 64KB / 50ms = 10.4 Mbps // This explains the 100 Mbps cap!

Step 4: Apply Fix

-- /etc/sysctl.conf: net.ipv4.tcp_rmem = 4096 87380 134217728 # 128MB net.ipv4.tcp_wmem = 4096 65536 134217728 # 128MB net.core.rmem_max = 134217728 net.core.wmem_max = 134217728 net.ipv4.tcp_window_scaling = 1 net.ipv4.tcp_congestion_control = cubic sudo sysctl -p

Step 5: Verify Fix

-- Test with iperf3: iperf3 -c remote_host -t 30 -- Before: 98 Mbps -- After: 9.4 Gbps โœ… 96x faster!

Key Takeaway: TCP buffers must be sized for bandwidth-delay product. Default 6MB is only enough for 1 Gbps with low latency!

2 Explain the bandwidth-delay product and how it affects TCP performance. โ–ผ

Answer: BDP = Bandwidth ร— RTT. It represents the amount of data "in flight" on the network. TCP buffer must be โ‰ฅ BDP to fully utilize available bandwidth. If buffer < BDP, TCP can't keep the pipe full, limiting throughput to buffer_size / RTT.

Understanding BDP:

-- BDP Formula: BDP = Bandwidth ร— Round-Trip Time -- Example calculations: # Same datacenter (low latency): Bandwidth: 10 Gbps = 1.25 GB/s RTT: 0.2ms = 0.0002s BDP = 1.25 GB/s ร— 0.0002s = 250 KB # Cross-region (medium latency): Bandwidth: 10 Gbps = 1.25 GB/s RTT: 50ms = 0.05s BDP = 1.25 GB/s ร— 0.05s = 62.5 MB # Cross-continent (high latency): Bandwidth: 10 Gbps = 1.25 GB/s RTT: 200ms = 0.2s BDP = 1.25 GB/s ร— 0.2s = 250 MB

Why BDP Matters:

TCP uses a "window" to control how much data can be in flight (sent but not yet acknowledged). The window must be at least BDP size to keep the network pipe full.

-- Scenario: Buffer < BDP Buffer: 6 MB BDP: 62.5 MB (10 Gbps ร— 50ms) // TCP can only send 6MB before waiting for ACK // Effective throughput: Throughput = Buffer / RTT Throughput = 6 MB / 0.05s = 120 MB/s = 960 Mbps // 10 Gbps link running at 960 Mbps! (9.6% utilization) -- Scenario: Buffer โ‰ฅ BDP Buffer: 128 MB BDP: 62.5 MB // TCP can send full BDP worth of data Throughput = 10 Gbps โœ… (Full utilization!)

Buffer Sizing Guide:

Scenario BDP Buffer Size
Same DC, 10 Gbps 2.5 MB 32 MB (safety margin)
Cross-region, 10 Gbps 62.5 MB 128 MB (2x BDP)
Cross-continent, 10 Gbps 250 MB 512 MB (2x BDP)

Key Takeaway: Always size TCP buffers to 2x BDP. Higher latency = larger buffers needed. Default 6MB only works for low-latency 1 Gbps networks!

3 What is the difference between CUBIC and BBR congestion control algorithms? When would you use each? โ–ผ

Answer: CUBIC is loss-based (reacts to packet loss), while BBR is model-based (measures actual bandwidth and RTT). CUBIC works well for low-latency networks. BBR excels in high-latency, high-bandwidth networks (cross-DC Cassandra). BBR maintains 2-4x higher throughput in lossy networks.

CUBIC (Default on Most Linux):

  • Algorithm: Loss-based congestion control
  • How it works: Grows window aggressively until packet loss, then reduces by 30%
  • Recovery: Fast recovery using cubic function
  • Best for: Low-latency networks (< 10ms RTT)
  • Problem: Treats packet loss as congestion (not always true!)
-- Enable CUBIC: sysctl -w net.ipv4.tcp_congestion_control=cubic -- CUBIC behavior on packet loss: /* Window: 100 packets [Packet loss detected] Window: 70 packets (reduced 30%) Window: 73, 76, 80... (cubic growth) Window: 100 (back to original) */

BBR (Google's Algorithm):

  • Algorithm: Model-based congestion control
  • How it works: Measures actual bandwidth and RTT, targets optimal operating point
  • Recovery: Doesn't rely on packet loss
  • Best for: High-latency (> 10ms RTT), high-bandwidth, or lossy networks
  • Advantage: Maintains high throughput even with packet loss
-- Enable BBR (requires Linux 4.9+): sysctl -w net.ipv4.tcp_congestion_control=bbr sysctl -w net.core.default_qdisc=fq -- BBR behavior on packet loss: /* BBR: "Is this packet loss or congestion?" - Measures actual bandwidth - Measures actual RTT - Calculates optimal window - Maintains throughput! No drastic window reduction on random loss! */

Performance Comparison:

Scenario CUBIC BBR
Same DC (0.5ms, 0% loss) 9.8 Gbps โœ… 9.7 Gbps โœ…
Cross-region (50ms, 0% loss) 8.5 Gbps 9.3 Gbps โœ…
Cross-DC (50ms, 0.1% loss) 2.3 Gbps โŒ 8.9 Gbps โœ…
Lossy network (100ms, 1% loss) 500 Mbps โŒ 7.2 Gbps โœ…

Recommendation for Cassandra:

  • โœ… Same Datacenter: CUBIC or BBR (both work well)
  • โญ Cross-Datacenter: BBR (2-4x better performance!)
  • โญ Multi-region Replication: BBR (essential!)
  • โš ๏ธ Requires: Linux kernel 4.9+ for BBR

Key Takeaway: BBR is superior for cross-DC Cassandra replication. It maintains high throughput even with packet loss, making it ideal for WAN connections!

4 How would you troubleshoot slow repair streaming (10 hours instead of 10 minutes)? โ–ผ

Answer: Check stream_throughput setting (increase from 200 to 400-800 Mbps), verify TCP buffer sizes, check network bandwidth utilization, verify no packet loss/retransmits, and ensure streaming_connections_per_host is adequate (4-16). Also check for IRQ imbalance and enable jumbo frames if possible.

Troubleshooting Steps:

Step 1: Check Cassandra Stream Settings

-- cassandra.yaml: grep stream_throughput /etc/cassandra/cassandra.yaml /* stream_throughput_outbound_megabits_per_sec: 200 โ†‘ Only 200 Mbps! โŒ */ -- For 10 Gbps network, increase to 400-800: stream_throughput_outbound_megabits_per_sec: 800 -- Restart Cassandra

Step 2: Monitor Actual Bandwidth

-- During repair, check bandwidth: iftop -i eth0 /* If showing 100 Mbps instead of 10 Gbps: โ†’ TCP tuning issue! */ -- Check for bottlenecks: nload eth0 // Should show near wire speed during streaming

Step 3: Verify TCP Settings

-- Check buffer sizes: sysctl net.ipv4.tcp_rmem // Should be 32-128MB for 10 Gbps -- Check window scaling: sysctl net.ipv4.tcp_window_scaling // Must be 1! -- If wrong, fix: sudo sysctl -w net.ipv4.tcp_rmem="4096 87380 33554432" sudo sysctl -w net.ipv4.tcp_wmem="4096 65536 33554432" sudo sysctl -w net.ipv4.tcp_window_scaling=1

Step 4: Check Network Health

-- Check for packet loss: netstat -s | grep -i retransmit /* 45678 segments retransmitted If this grows rapidly โ†’ packet loss! */ -- Check interface errors: ifconfig eth0 | grep errors // Should be 0! -- Test actual bandwidth: iperf3 -c other_node -t 30 // Should get 9+ Gbps

Step 5: Optimize Streaming Connections

-- cassandra.yaml: streaming_connections_per_host: 4 # Default -- Increase for faster repairs: streaming_connections_per_host: 8 -- With 8 connections ร— 800 Mbps = 6.4 Gbps total!

Step 6: Additional Optimizations

-- Enable jumbo frames (if datacenter supports): ifconfig eth0 mtu 9000 -- Check IRQ balance: cat /proc/interrupts | grep eth0 // Should be distributed across CPUs -- Verify congestion control: sysctl net.ipv4.tcp_congestion_control // Should be cubic or bbr

Expected Results:

-- Before optimization: Stream rate: 100 MB/s 100GB repair = 1000 seconds = 16 minutes per node 10 nodes = 160 minutes (2.7 hours) -- After optimization: Stream rate: 1000 MB/s (10x faster!) 100GB repair = 100 seconds per node 10 nodes = 16 minutes total โœ…

Key Takeaway: Stream throttling + poor TCP tuning can make repairs 100x slower. Always tune TCP buffers and increase stream_throughput for modern networks!

5 Should you enable jumbo frames (MTU 9000) for Cassandra? What are the trade-offs? โ–ผ

Answer: Yes for same-datacenter deployments if entire network path supports it. Jumbo frames reduce CPU overhead by 20-30% and improve throughput by 10-15%. But ALL devices (NICs, switches, routers) must support 9000 MTU. Never use for cross-DC or internet traffic.

Benefits of Jumbo Frames:

  • โšก Less CPU: 6x fewer packets = 20-30% less CPU for networking
  • โšก Less Interrupts: Fewer packets = fewer NIC interrupts
  • โšก Better Throughput: 10-15% higher throughput
  • โšก Lower Latency: Slightly lower per-byte latency
-- Packet comparison: // Standard MTU (1500 bytes): To send 54MB SSTable: Packets needed: 54MB / 1500 = 36,000 packets Interrupts: 36,000 CPU overhead: High // Jumbo frames (9000 bytes): To send 54MB SSTable: Packets needed: 54MB / 9000 = 6,000 packets Interrupts: 6,000 CPU overhead: 83% less! โœ…

Requirements (ALL must be met):

  • โœ… All Cassandra node NICs support MTU 9000
  • โœ… All switches support jumbo frames
  • โœ… All routers in path support jumbo frames
  • โœ… Same physical datacenter (no internet/WAN)
  • โŒ If ANY device doesn't support โ†’ fragmentation โ†’ worse performance!

Testing for Jumbo Frame Support:

-- Test if path supports 9000 MTU: ping -M do -s 8972 remote_host # -M do = Don't fragment # -s 8972 = 8972 + 28 header = 9000 bytes -- If successful: /* 64 bytes from remote_host: icmp_seq=1 ttl=64 โœ… Path supports jumbo frames! */ -- If failed: /* ping: local error: Message too long โŒ Path does NOT support jumbo frames! */

Enabling Jumbo Frames:

-- 1. Set MTU on interface: sudo ifconfig eth0 mtu 9000 -- 2. Make permanent (/etc/network/interfaces): auto eth0 iface eth0 inet static address 192.168.1.10 netmask 255.255.255.0 mtu 9000 -- 3. Configure switch (example for Cisco): # interface GigabitEthernet1/0/1 # mtu 9000 -- 4. Verify: ip link show eth0 | grep mtu // mtu 9000 โœ…

When NOT to Use:

Scenario Use Jumbo? Reason
Same datacenter, modern switches โœ… Yes Full control, 20-30% CPU savings
Cross-datacenter (WAN) โŒ No WAN rarely supports, fragmentation
Cloud (AWS/Azure/GCP) โŒ No Usually not supported
Mixed old/new switches โŒ No Old switches may not support
Client traffic (CQL) โŒ No Clients may not support

Performance Impact:

-- Measured improvement (10 Gbps network): Standard MTU (1500): Throughput: 8.7 Gbps CPU usage: 35% Jumbo Frames (9000): Throughput: 9.5 Gbps (9% better) CPU usage: 25% (29% less!)

Key Takeaway: Jumbo frames are worthwhile for same-datacenter Cassandra if ALL network gear supports it. Test thoroughly before enabling. Never use for cross-DC or cloud deployments!

๐ŸŽ“ Chapter Summary: Network Tuning Mastery

You now understand network tuning at a production level!

Critical Optimizations:

  • ๐Ÿ“ฆ TCP Buffers: 32-128MB (not 6MB default!)
  • ๐ŸชŸ Window Scaling: Enable (allows > 64KB windows)
  • ๐Ÿš€ BBR/CUBIC: Use BBR for cross-DC
  • โšก IRQ Balance: Spread across all CPUs
  • ๐Ÿ”ข Conntrack: Increase to 1M connections

Quick Wins:

net.ipv4.tcp_rmem = 4096 87380 33554432 net.ipv4.tcp_wmem = 4096 65536 33554432 net.ipv4.tcp_window_scaling = 1 net.ipv4.tcp_congestion_control = bbr net.netfilter.nf_conntrack_max = 1048576

Performance Gain:

100 Mbps โ†’ 9.4 Gbps (94x faster!) with $0 cost!

Remember Sarah: Network tuning is FREE 100x performance! ๐Ÿš€

Advertisement

Responsive Ad