Fix It Fast

Troubleshooting

Debug and solve Cassandra issues like a pro!

📖 The Troubleshooting Framework

Troubleshooting Cassandra can feel overwhelming. But with a systematic approach, you can debug ANY issue. This guide teaches you the framework used by expert DBAs to solve problems fast.

✅ The 5-Step Debug Process

  1. Identify: What exactly is broken?
  2. Isolate: Which component has the problem?
  3. Investigate: Gather evidence (logs, metrics)
  4. Fix: Apply the solution
  5. Verify: Confirm it's resolved

Follow this process every time, and you'll solve issues 10x faster! 🎯

🎯 Troubleshooting Methodology

Your systematic approach!

Step 1: Identify the Problem

Ask the right questions:

  • What's the symptom? (Errors, slow queries, crashes)
  • When did it start? (After deploy, gradually, suddenly)
  • Is it intermittent or constant?
  • Which nodes/keyspaces/tables affected?
  • Can you reproduce it?

Step 2: Check the Obvious First

-- Quick health checks: # 1. Are all nodes up? $ nodetool status # 2. Any errors in logs? $ tail -100 /var/log/cassandra/system.log | grep ERROR # 3. Disk space OK? $ df -h # 4. Memory OK? $ free -h # 5. High CPU? $ top -- Fix the obvious before digging deeper!

Step 3: Gather Evidence

  • Logs: system.log, debug.log, gc.log
  • Metrics: Grafana dashboards, JMX metrics
  • System: CPU, memory, disk, network
  • Queries: Slow query logs, tracing
  • Config: Recent changes?

Step 4: Form Hypothesis

Based on evidence, what's the likely cause?

  • Memory issue? → Check heap usage, GC logs
  • Disk issue? → Check space, I/O stats
  • Network issue? → Check connectivity, dropped messages
  • Query issue? → Check query patterns, tombstones

Step 5: Test & Fix

-- Test your hypothesis: 1. Try fix in staging/test environment 2. Monitor impact 3. If works → apply to production 4. If doesn't work → back to hypothesis -- Always verify fix worked!

🚫 Issue: Node Won't Start

Most common startup problems!

Problem 1: Port Already in Use

-- SYMPTOM in system.log: ERROR Failed to bind to: /0.0.0.0:9042 java.net.BindException: Address already in use -- CAUSE: Another process using port 9042 -- DIAGNOSE: $ sudo lsof -i :9042 # OR $ sudo netstat -tlnp | grep 9042 -- FIX Option 1: Kill old Cassandra process $ ps aux | grep cassandra $ sudo kill -9 -- FIX Option 2: Change port in cassandra.yaml native_transport_port: 9043 -- Restart: $ sudo systemctl start cassandra

Problem 2: Permission Denied

-- SYMPTOM: ERROR Unable to create commit log directory java.io.IOException: Permission denied -- CAUSE: Wrong ownership/permissions -- DIAGNOSE: $ ls -la /var/lib/cassandra/ -- FIX: Set correct ownership $ sudo chown -R cassandra:cassandra /var/lib/cassandra $ sudo chown -R cassandra:cassandra /var/log/cassandra -- Verify: $ ls -la /var/lib/cassandra/ | head drwxr-xr-x cassandra cassandra ...

Problem 3: Config Syntax Error

-- SYMPTOM: ERROR Cannot start node. Invalid yaml org.yaml.snakeyaml.error.YAMLException -- CAUSE: Typo in cassandra.yaml -- DIAGNOSE: Check YAML syntax $ sudo vim /etc/cassandra/cassandra.yaml -- Common mistakes: # Wrong indentation # Missing colon # Unquoted special characters -- FIX: Restore from backup $ sudo cp /etc/cassandra/cassandra.yaml.bak \ /etc/cassandra/cassandra.yaml -- OR validate YAML online: # Copy contents to yamllint.com

Problem 4: Disk Full

-- SYMPTOM: ERROR Cannot write to commit log java.io.IOException: No space left on device -- DIAGNOSE: $ df -h Filesystem Size Used Avail Use% /dev/sda1 100G 100G 0 100% ← FULL! -- FIX: Free up space # 1. Clear old snapshots: $ nodetool clearsnapshot --all # 2. Check what's using space: $ du -sh /var/lib/cassandra/* | sort -h # 3. Clear old logs: $ sudo find /var/log/cassandra -name "*.log.*" -mtime +7 -delete # 4. If data directory too big, add disk or scale out

Problem 5: Wrong Java Version

-- SYMPTOM: ERROR Cassandra 4.0 requires Java 11 or later -- DIAGNOSE: $ java -version openjdk version "1.8.0_292" ← Too old! -- FIX: Install Java 11 $ sudo apt install openjdk-11-jdk -- Set as default: $ sudo update-alternatives --config java # Select Java 11 -- Verify: $ java -version openjdk version "11.0.11" ← Good!

⏱️ Issue: High Latency / Slow Queries

Performance degradation troubleshooting!

Cause 1: Long GC Pauses

-- SYMPTOM: P99 latency spikes -- DIAGNOSE: Check GC logs $ grep "GC for" /var/log/cassandra/system.log | tail -20 2024-01-15 14:32:15 GC for ConcurrentMarkSweep: 5230ms → 5 second pause! 💥 -- Check heap usage: $ nodetool info | grep Heap Heap Memory (MB): 7850/8192 ← 96% full! -- FIX: Increase heap size $ sudo vim /etc/cassandra/jvm.options # Change from: -Xms8G -Xmx8G # To: -Xms16G -Xmx16G -- Restart: $ sudo systemctl restart cassandra -- Monitor GC after restart

Cause 2: Excessive Tombstones

-- SYMPTOM: Slow reads, timeouts -- DIAGNOSE: Check for tombstone warnings $ grep tombstone /var/log/cassandra/system.log Read 5000 live rows and 50000 tombstone cells → 10x more tombstones than data! 💥 -- Enable tracing on slow query: cqlsh> TRACING ON; cqlsh> SELECT * FROM users WHERE user_id='123'; /* Look for: "Scanned 50000 tombstones" */ -- FIX Option 1: Reduce gc_grace_seconds ALTER TABLE users WITH gc_grace_seconds = 86400; -- 1 day -- FIX Option 2: Run compaction $ nodetool compact my_keyspace users -- FIX Option 3: Review deletion patterns # Too many DELETEs? Consider TTL instead # Or redesign data model

Cause 3: Large Partitions

-- SYMPTOM: Specific queries very slow -- DIAGNOSE: Find large partitions $ nodetool tablehistograms my_keyspace users Partition Size: 5MB: 1000 10MB: 500 100MB: 10 ← Problem! 1000MB: 2 ← BIG PROBLEM! -- Check specific partition: $ nodetool cfstats my_keyspace.users | grep -i partition -- FIX: Redesign data model # Break large partition into multiple partitions # Add time bucket to partition key -- Example BAD: PRIMARY KEY (user_id) ← All data for user in 1 partition -- Example GOOD: PRIMARY KEY ((user_id, year_month), timestamp) ← Data spread across monthly partitions

Cause 4: Disk I/O Saturation

-- SYMPTOM: All queries slow -- DIAGNOSE: Check I/O wait $ iostat -x 1 Device %util sda 98.5 ← Disk maxed out! -- Check pending compactions: $ nodetool compactionstats pending tasks: 500 ← Backlog! -- FIX Option 1: Increase compaction throughput $ nodetool setcompactionthroughput 32 # MB/s -- FIX Option 2: Upgrade to faster disks # SSD > HDD for Cassandra -- FIX Option 3: Add more nodes # Distribute I/O load

Cause 5: Network Issues

-- SYMPTOM: Intermittent timeouts -- DIAGNOSE: Check dropped messages $ nodetool tpstats | grep -i dropped Dropped Messages: MUTATION: 1250 ← Dropping writes! -- Check network connectivity: $ nodetool netstats -- FIX: Increase timeouts $ vim /etc/cassandra/cassandra.yaml read_request_timeout_in_ms: 10000 # Increase from 5000 write_request_timeout_in_ms: 4000 # Increase from 2000 -- Check for network packet loss: $ ping -c 100 other_node_ip # Should see 0% packet loss

💥 Issue: Out of Memory (OOM)

Memory exhaustion troubleshooting!

Diagnosis

-- SYMPTOM: Node crashes, can't restart -- Check logs: $ grep -i "OutOfMemory" /var/log/cassandra/system.log ERROR OutOfMemoryError: Java heap space at org.apache.cassandra.db.Memtable -- OR check system logs: $ sudo journalctl -u cassandra | grep -i "OutOfMemory" -- Check heap dump (if exists): $ ls -lh /var/lib/cassandra/*.hprof

Common Causes

  • Heap too small: Need to increase -Xmx
  • Large queries: Fetching too much data
  • Memory leak: Bug in Cassandra or driver
  • Too many memtables: High write load
  • Large partitions: Reading huge partitions

Fixes

-- FIX 1: Increase heap size $ sudo vim /etc/cassandra/jvm.options # Rule of thumb: 1/4 to 1/2 of system RAM # Or max 32GB (due to Java compressed OOPs) # For 64GB RAM machine: -Xms16G -Xmx16G -- FIX 2: Enable heap dump on OOM -XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=/var/lib/cassandra/ -- FIX 3: Reduce memtable sizes $ vim /etc/cassandra/cassandra.yaml memtable_heap_space_in_mb: 2048 # Reduce if needed memtable_offheap_space_in_mb: 2048 -- FIX 4: Add query limits SELECT * FROM users LIMIT 1000; ← ALWAYS use LIMIT! -- FIX 5: Upgrade Cassandra version # Memory leaks fixed in newer versions

🔗 Issue: Cluster Connectivity Problems

Node communication issues!

Problem: Nodes Can't See Each Other

-- SYMPTOM: $ nodetool status # Only shows 1 node (should show all nodes) -- DIAGNOSE: Check gossip $ nodetool gossipinfo | grep STATUS -- Check network connectivity: $ telnet other_node_ip 7000 ← Gossip port $ telnet other_node_ip 9042 ← CQL port -- FIX: Check firewall $ sudo ufw status # Make sure ports 7000, 7001, 9042 are open $ sudo ufw allow 7000/tcp $ sudo ufw allow 7001/tcp $ sudo ufw allow 9042/tcp -- Check seed nodes in cassandra.yaml: seed_provider: - class_name: org.apache.cassandra.locator.SimpleSeedProvider parameters: - seeds: "node1_ip,node2_ip" ← Must be correct!

Problem: Node Flapping (UP/DOWN/UP/DOWN)

-- SYMPTOM in logs: Node /10.0.1.11 DOWN Node /10.0.1.11 UP Node /10.0.1.11 DOWN ← Flapping! -- CAUSES: # 1. Long GC pauses (node unresponsive) # 2. Network packet loss # 3. Disk I/O saturation # 4. CPU maxed out -- DIAGNOSE: # Check GC on flapping node: $ grep "GC for" /var/log/cassandra/system.log # Check system resources: $ top $ iostat -x 1 -- FIX: Address root cause # If GC → increase heap # If I/O → faster disks or reduce compaction # If network → check switches/routers

📊 Issue: Data Inconsistency

Replicas out of sync!

Diagnosis

-- SYMPTOM: Different results from different nodes -- Test consistency: cqlsh> CONSISTENCY ALL; cqlsh> SELECT * FROM users WHERE user_id='123'; /* If fails: ReadTimeout: code=1200 ... → Replicas don't agree! */ -- Check for hints: $ nodetool statusbinary $ nodetool statusgossip -- Check repair history: $ grep repair /var/log/cassandra/system.log

Fix: Run Repair

-- Option 1: Repair single table $ nodetool repair -pr my_keyspace users -- Option 2: Repair entire keyspace $ nodetool repair -pr my_keyspace -- Option 3: Full repair (all keyspaces) $ nodetool repair -pr -- Monitor progress: $ nodetool compactionstats -- Note: -pr = primary range only (prevents duplicates) -- Schedule regular repairs: # Cron job to run weekly 0 2 * * 0 nodetool repair -pr

🛠️ Troubleshooting Toolbox

Essential commands!

🔍

Health Checks

# Cluster status nodetool status # Node info nodetool info # Ring state nodetool ring # Gossip info nodetool gossipinfo
📊

Performance

# Thread pools nodetool tpstats # Table stats nodetool tablestats # Compaction stats nodetool compactionstats # GC stats nodetool gcstats
📜

Logs

# Errors grep ERROR system.log # Warnings grep WARN system.log # GC pauses grep "GC for" system.log # Tombstones grep tombstone system.log
💻

System

# Disk space df -h # Memory free -h # CPU top # Disk I/O iostat -x 1
🔧

Repairs

# Single table nodetool repair -pr ks table # Keyspace nodetool repair -pr ks # Full repair nodetool repair -pr # Monitor nodetool compactionstats
🐛

Debugging

# Enable tracing cqlsh> TRACING ON; # Query tracing SELECT * FROM ...; # Check consistency CONSISTENCY ALL; # Network stats nodetool netstats

Quick Reference Cheat Sheet

Problem Quick Check Common Fix
Node down nodetool status Check logs, restart node
Slow queries grep tombstone logs Run compaction, check GC
High latency nodetool tpstats Check pending tasks, GC
OOM errors grep OutOfMemory logs Increase heap size
Disk full df -h Clear snapshots, add disk
Dropped messages nodetool tpstats Increase timeouts, scale out
Inconsistent data CONSISTENCY ALL Run repair
Node flapping grep "GC for" logs Fix GC pauses, network

🎉 Master Cassandra Troubleshooting!

You now have a systematic approach to debug ANY issue!

🎓 What You Learned:

  • 🎯 Methodology: The 5-step debug process
  • 🚫 Node won't start: 5 common causes + fixes
  • ⏱️ High latency: GC, tombstones, large partitions, I/O, network
  • 💥 OOM errors: Diagnosis and heap tuning
  • 🔗 Cluster issues: Connectivity and flapping nodes
  • 📊 Data inconsistency: When and how to repair
  • 🛠️ Toolbox: Essential commands for every scenario

💡 The 5-Step Process:

  1. Identify: What exactly is broken?
  2. Isolate: Which component?
  3. Investigate: Gather logs, metrics, evidence
  4. Fix: Apply solution (test in staging first!)
  5. Verify: Confirm it's actually fixed

🔍 Essential Commands:

# Quick health check: nodetool status # Cluster status tail -100 system.log | grep ERROR # Recent errors df -h # Disk space nodetool tpstats # Thread pools grep "GC for" system.log # GC pauses # When things go wrong: nodetool repair -pr # Fix inconsistency nodetool compact ks table # Remove tombstones nodetool clearsnapshot --all # Free disk space

🚨 Most Common Issues:

Issue First Thing to Check
Node won't start tail system.log (look for ERROR)
Slow queries grep tombstone system.log
Node crashes grep OutOfMemory system.log
High latency grep "GC for" system.log
Disk full df -h (then clear snapshots)

🔍 Remember: Systematic approach = Fast resolution! 🎯
Follow the 5 steps every time!

Advertisement

📱 Responsive Ad 📱