Debug Like a Pro

Log Management

Master Cassandra logs for debugging and troubleshooting!

📖 The Story: David's 3AM Log Hunt

3AM. Production down. David gets the call. "Cassandra is crashing!" He SSHs to the server. Looks at `/var/log/cassandra/system.log`. File is 50GB! His `tail` command hangs. Disk is 99% full - ALL logs! Log rotation never configured. Takes 2 hours to find the issue. $200K lost. All because logs weren't managed.

😱 The Midnight Crisis

3:00 AM - The Call:

-- PagerDuty: "CRITICAL: Cassandra node1 DOWN!" $ ssh node1 $ sudo systemctl status cassandra Failed (code=exited, status=137/) -- David: "Let me check the logs..."

3:05 AM - The Horror:

$ cd /var/log/cassandra $ ls -lh -rw-r--r-- 1 cassandra cassandra 50G system.log -rw-r--r-- 1 cassandra cassandra 25G debug.log -rw-r--r-- 1 cassandra cassandra 15G gc.log -- Total: 90GB of logs! 💥 $ df -h Filesystem Size Used Avail Use% /dev/sda1 100G 99G 1G 99% -- Disk 99% FULL! 💥

3:10 AM - Searching for Needle in Haystack:

-- David tries to read the log: $ tail -100 system.log ... command hangs for 5 minutes ... ... finally shows last 100 lines ... ... but those are from yesterday! ... -- Tries to search for ERROR: $ grep ERROR system.log ... hangs for 10 minutes scanning 50GB ... ... memory runs out ... Killed -- Can't find the issue! 😱

3:30 AM - The Struggle:

  • 💥 Log files too big to read
  • 💥 grep/tail commands hang
  • 💥 Disk 99% full = Cassandra can't write
  • 💥 No log rotation configured
  • 💥 No idea WHEN the issue started
  • 💥 2 hours of trial and error

5:00 AM - Finally Found It:

-- After 2 hours, found the issue: $ journalctl -u cassandra | grep -i "out of memory" OutOfMemoryError: Java heap space -- Heap was too small! -- Issue started 3 days ago -- But logs were so messy couldn't find it!

Total Damage:

  • 💰 $200K: Lost revenue (2 hours)
  • ⏰ 2 hours: To find simple OOM error
  • 😫 David: Exhausted, frustrated
  • 📊 Boss: "Why did this take so long?"

✅ With Proper Log Management

What David Should Have Had:

-- Configured log rotation (logback.xml): 100MB 7 -- Result: Logs rotate daily, keep only 7 days $ ls -lh /var/log/cassandra/ -rw-r--r-- system.log (80 MB) ← Current -rw-r--r-- system.log.1 (100 MB) ← Yesterday -rw-r--r-- system.log.2.gz (25 MB) ← 2 days ago -rw-r--r-- system.log.3.gz (25 MB) -rw-r--r-- system.log.4.gz (25 MB) Total: ~400 MB instead of 90GB! ✅

3:05 AM - Quick Debug:

-- Check recent errors: $ tail -100 /var/log/cassandra/system.log ... instant results! ... ERROR OutOfMemoryError: Java heap space at org.apache.cassandra... -- Found in 30 seconds! ✅ -- Or use grep (fast on small file): $ grep "OutOfMemory" system.log 2024-01-15 02:45:12 ERROR OutOfMemoryError -- Issue identified in minutes!

The Better Outcome:

  • ✅ 5 minutes: Found the issue
  • ✅ 15 minutes: Fixed and restarted
  • ✅ $10K: Lost revenue (vs $200K)
  • ✅ David: Back to sleep by 3:30 AM
  • ✅ Boss: "Great response time!"

Log rotation: The difference between 2 hours and 5 minutes! 🎯

📂 Cassandra Log Types

Know your logs!

📋

system.log

Main Cassandra log

Contains:

  • Startup messages
  • Error messages
  • Warnings
  • Cluster events
  • Node state changes

Check first!

Your starting point!

🔍

debug.log

Detailed debugging info

Contains:

  • DEBUG level messages
  • Query details
  • Compaction details
  • Gossip details
  • Very verbose

Size: Very large!

Deep debugging!

♻️

gc.log

JVM garbage collection

Contains:

  • GC events
  • Pause times
  • Heap usage
  • Memory stats
  • Performance data

For: Performance tuning

JVM health!

📊

output.log

Standard output/error

Contains:

  • stdout messages
  • stderr messages
  • Early startup errors
  • JVM crashes
  • System errors

Check: If won't start

Startup issues!

📝

commitlog

Write-ahead log

Contains:

  • All writes
  • Before memtable
  • For crash recovery
  • Binary format
  • Not human-readable

Location: Different dir

Data safety!

🔐

audit logs

Security auditing

Contains:

  • All queries
  • User actions
  • Schema changes
  • Login attempts
  • Compliance data

Optional: Must enable

Security!

📍 Log File Locations

Where to find logs!

Default Log Locations

-- Main log directory: /var/log/cassandra/ -- Log files: /var/log/cassandra/system.log ← Main log /var/log/cassandra/debug.log ← Debug log /var/log/cassandra/gc.log ← GC log /var/log/cassandra/output.log ← stdout/stderr -- Commit log (different location!): /var/lib/cassandra/commitlog/ -- Audit logs (if enabled): /var/log/cassandra/audit/

How to Find Log Files

Method 1: Check cassandra-env.sh

-- Log location defined in: $ grep CASSANDRA_LOG_DIR /etc/cassandra/cassandra-env.sh CASSANDRA_LOG_DIR="/var/log/cassandra" -- Can be changed by setting environment variable

Method 2: Check logback.xml

-- Logging configuration file: $ cat /etc/cassandra/logback.xml ${cassandra.logdir}/system.log -- ${cassandra.logdir} = /var/log/cassandra

Method 3: Use find Command

-- Search for system.log: $ sudo find / -name "system.log" 2>/dev/null /var/log/cassandra/system.log -- Quick access: $ cd /var/log/cassandra $ ls -lh

Check Disk Space!

Always monitor log directory disk usage:

-- Check disk space: $ df -h /var/log/cassandra Filesystem Size Used Avail Use% /dev/sda1 100G 45G 55G 45% ← Healthy -- Check log sizes: $ du -sh /var/log/cassandra/* 250M system.log ← OK 500M debug.log ← OK 100M gc.log ← OK -- WARNING if disk > 80% full! # Cassandra needs disk space to operate

🎚️ Log Levels

Control logging verbosity!

Level Description When to Use
TRACE Every tiny detail Almost never (too verbose)
DEBUG Detailed debugging info Troubleshooting specific issues
INFO Important events Production default
WARN Warning messages Always log
ERROR Error conditions Always log

Changing Log Levels

1

Runtime (Temporary)

No restart required!

-- Via nodetool: $ nodetool setlogginglevel org.apache.cassandra DEBUG -- Check current level: $ nodetool getlogginglevels Logger Name Log Level org.apache.cassandra DEBUG -- Reset to default: $ nodetool setlogginglevel org.apache.cassandra INFO -- Enable debug for specific class: $ nodetool setlogginglevel \ org.apache.cassandra.service.StorageService DEBUG
2

Permanent (Config File)

Edit logback.xml

-- Edit logging config: $ sudo vim /etc/cassandra/logback.xml "INFO"> "SYSTEMLOG" /> "org.apache.cassandra.db.compaction" level="DEBUG"/> -- Restart to apply: $ sudo systemctl restart cassandra

WARNING: DEBUG in Production

DEBUG level creates HUGE log files!

  • ⚠️ Size: Can grow 10x-100x faster
  • ⚠️ Performance: Impacts performance
  • ⚠️ Disk: Can fill disk quickly
  • ✅ Use: Only for specific troubleshooting
  • ✅ Duration: Enable temporarily, then disable

📖 Reading and Analyzing Logs

Essential commands!

1. View Recent Log Entries

-- Last 100 lines: $ tail -100 /var/log/cassandra/system.log -- Follow in real-time (like watching live): $ tail -f /var/log/cassandra/system.log -- Follow and show last 50: $ tail -f -n 50 /var/log/cassandra/system.log -- Ctrl+C to stop following

2. Search for Errors

-- Find all ERROR messages: $ grep ERROR /var/log/cassandra/system.log -- Case-insensitive search: $ grep -i error /var/log/cassandra/system.log -- Show 5 lines before/after match: $ grep -C 5 ERROR /var/log/cassandra/system.log -- Count errors: $ grep -c ERROR /var/log/cassandra/system.log 42 ← 42 errors found -- Search for specific error: $ grep "OutOfMemoryError" /var/log/cassandra/system.log

3. Filter by Time Range

-- Find entries from specific time: $ grep "2024-01-15 14:" /var/log/cassandra/system.log -- Between 2PM and 3PM: $ awk '/2024-01-15 14:/,/2024-01-15 15:/' \ /var/log/cassandra/system.log -- Today's errors: $ grep "$(date +%Y-%m-%d)" /var/log/cassandra/system.log | \ grep ERROR

4. Search Across Multiple Files

-- Search system.log and rotated files: $ grep ERROR /var/log/cassandra/system.log* -- Search compressed files too: $ zgrep ERROR /var/log/cassandra/system.log*.gz -- Search all logs in directory: $ grep -r ERROR /var/log/cassandra/

5. Common Log Patterns

-- Find compaction issues: $ grep -i compaction /var/log/cassandra/system.log | \ grep -i error -- Find GC pauses > 1 second: $ grep "GC for" /var/log/cassandra/system.log | \ awk '$7 > 1000' -- Find connection timeouts: $ grep -i "timeout" /var/log/cassandra/system.log -- Find node joining/leaving: $ grep -E "(Joining|Leaving|Normal)" \ /var/log/cassandra/system.log | tail -20

Log Analysis Scripts

#!/bin/bash # analyze_errors.sh - Quick error summary LOG="/var/log/cassandra/system.log" echo "=== Cassandra Log Analysis ===" echo "" # Count by error type echo "Error Counts:" grep ERROR $LOG | \ awk '{for(i=1;i<=NF;i++) if($i=="ERROR") print $(i+1)}' | \ sort | uniq -c | sort -rn | head -10 echo "" # Recent errors (last hour) echo "Recent Errors:" grep ERROR $LOG | tail -20 echo "" # GC pause times echo "Recent Long GC Pauses (>1s):" grep "GC for" $LOG | \ awk '$7 > 1000 {print $1, $2, $7"ms"}' | tail -10

🔄 Log Rotation Configuration

Prevent David's nightmare!

Why Log Rotation?

  • ✅ Prevent disk from filling up (99% = crash!)
  • ✅ Keep logs manageable (50GB = unusable)
  • ✅ Faster to read/search (small files = fast grep)
  • ✅ Automatic cleanup (old logs deleted)
  • ✅ Compression (save disk space)

Configure Log Rotation (logback.xml)

-- Edit /etc/cassandra/logback.xml: $ sudo vim /etc/cassandra/logback.xml "SYSTEMLOG" class="ch.qos.logback.core.rolling.RollingFileAppender"> ${cassandra.logdir}/system.log "ch.qos.logback.core.rolling.SizeAndTimeBasedRollingPolicy"> ${cassandra.logdir}/system.log.%d{yyyy-MM-dd}.%i.gz 100MB 7 1GB %-5level [%thread] %date{ISO8601} %F:%L - %msg%n -- Restart to apply: $ sudo systemctl restart cassandra

Rotation Settings Explained

Setting Recommended Why
maxFileSize 100MB Balance between file size and number of files
maxHistory 7 days Keep one week for debugging
totalSizeCap 1-5GB Hard limit on total log size
Compression *.gz Save 80-90% disk space

Verify Log Rotation Working

-- Check log files: $ ls -lh /var/log/cassandra/ /* Should see: -rw-r--r-- system.log (80 MB) ← Current -rw-r--r-- system.log.2024-01-15.0.gz (25 MB) ← Yesterday -rw-r--r-- system.log.2024-01-14.0.gz (25 MB) ← 2 days ago -rw-r--r-- system.log.2024-01-13.0.gz (25 MB) ← 3 days ago ... */ -- Check disk usage: $ du -sh /var/log/cassandra 500M /var/log/cassandra ← Reasonable size! ✅ -- Monitor over time: $ watch -n 60 'du -sh /var/log/cassandra'

🔍 Log-Based Troubleshooting

Find issues fast!

Issue 1: Node Won't Start

-- Check system.log: $ tail -100 /var/log/cassandra/system.log -- Common errors: # Port already in use: ERROR Failed to bind to: /0.0.0.0:9042 → Fix: Kill other process or change port # Config syntax error: ERROR Cannot start node. Invalid yaml → Fix: Check cassandra.yaml syntax # Permission denied: ERROR Unable to create commit log directory → Fix: chown cassandra /var/lib/cassandra # Disk full: ERROR Cannot write to commit log. Disk full → Fix: Free up disk space

Issue 2: Performance Degradation

-- Check for long GC pauses: $ grep "GC for" /var/log/cassandra/system.log | \ awk '$7 > 1000' /* Output: 2024-01-15 14:32:15 GC for ConcurrentMarkSweep: 5230ms 2024-01-15 14:45:22 GC for ConcurrentMarkSweep: 8120ms → GC taking > 5 seconds = heap too small or memory leak! */ -- Check for dropped messages: $ grep "dropped messages" /var/log/cassandra/system.log → If many dropped = overload or timeout issues -- Check for timeouts: $ grep -i timeout /var/log/cassandra/system.log

Issue 3: Cluster Instability

-- Check node state changes: $ grep -E "(UP|DOWN|Joining|Leaving)" \ /var/log/cassandra/system.log | tail -20 /* Example output: 2024-01-15 14:30:00 Node /10.0.1.11 DOWN 2024-01-15 14:30:30 Node /10.0.1.11 UP 2024-01-15 14:31:00 Node /10.0.1.11 DOWN → Flapping node = network or hardware issue! */ -- Check for gossip issues: $ grep -i gossip /var/log/cassandra/system.log | tail -50

Issue 4: Out of Memory

-- Search for OOM: $ grep -i "OutOfMemory" /var/log/cassandra/system.log ERROR OutOfMemoryError: Java heap space at org.apache.cassandra.db.Memtable → Fix: Increase heap size in jvm.options -- Check heap usage before crash: $ grep "Heap" /var/log/cassandra/system.log | tail -20

💡 Log Management Best Practices

Don't be David!

✅

DO

  • Configure log rotation
  • Monitor disk space
  • Set reasonable retention (7 days)
  • Compress old logs
  • Use INFO level in prod
  • Ship logs to centralized system
  • Set up log alerts
  • Regular log reviews
❌

DON'T

  • Skip log rotation setup
  • Use DEBUG in production
  • Let logs fill disk
  • Delete logs manually
  • Ignore WARNING messages
  • Keep logs forever
  • Skip monitoring disk
  • Wait for logs to cause issues

Complete Log Management Checklist

Task How To Done?
Log Rotation Configure logback.xml (maxFile Size=100MB, maxHistory=7) ☐
Disk Monitoring Alert if /var/log > 80% full ☐
Log Level INFO in production (DEBUG only for troubleshooting) ☐
Compression Enable .gz compression for rotated logs ☐
Retention Keep 7 days (adjust for compliance) ☐
Centralized Logging Ship to ELK, Splunk, or CloudWatch ☐
Error Alerts Alert on ERROR/WARN patterns ☐
GC Monitoring Alert on GC pauses > 1 second ☐

Log Shipping to Central System

-- Example: Ship logs to Elasticsearch with Filebeat # 1. Install Filebeat: $ sudo apt install filebeat # 2. Configure filebeat.yml: filebeat.inputs: - type: log enabled: true paths: - /var/log/cassandra/system.log fields: type: cassandra-system - type: log enabled: true paths: - /var/log/cassandra/debug.log fields: type: cassandra-debug output.elasticsearch: hosts: ["elasticsearch:9200"] index: "cassandra-logs-%{+yyyy.MM.dd}" # 3. Start Filebeat: $ sudo systemctl start filebeat # 4. View in Kibana! # Search, filter, visualize all logs in one place

🎉 Master Log Management!

You now know how to manage Cassandra logs like a pro!

🎓 What You Learned:

  • 📖 David's nightmare: 90GB logs, 2 hours to debug
  • 📂 Log types: system, debug, gc, output, audit
  • 📍 Locations: /var/log/cassandra/
  • 🎚️ Log levels: INFO for prod, DEBUG for troubleshooting
  • 📖 Reading logs: tail, grep, search patterns
  • 🔄 Log rotation: 100MB max, 7 days retention
  • 🔍 Troubleshooting: Find issues in logs
  • 💡 Best practices: Rotate, monitor, centralize

💡 Key Takeaways:

  1. Configure log rotation - First thing after install!
  2. Monitor disk space - Alert at 80% full
  3. Keep logs manageable - 100MB files, 7 days retention
  4. INFO level in production - DEBUG only temporarily
  5. Centralize logs - Ship to ELK/Splunk/CloudWatch
  6. Know your commands - tail, grep, awk are your friends

📋 Quick Log Rotation Setup:

# Edit /etc/cassandra/logback.xml: "SizeAndTimeBasedRollingPolicy"> ${cassandra.logdir}/system.log.%d{yyyy-MM-dd}.%i.gz 100MB 7 1GB # Restart: sudo systemctl restart cassandra # Done! No more 50GB logs! ✅

🔍 Essential Debugging Commands:

# Watch logs live: tail -f /var/log/cassandra/system.log # Find errors: grep ERROR /var/log/cassandra/system.log # Last 100 lines: tail -100 /var/log/cassandra/system.log # Search for OOM: grep -i OutOfMemory /var/log/cassandra/system.log # Check GC pauses: grep "GC for" /var/log/cassandra/system.log # Count errors: grep -c ERROR /var/log/cassandra/system.log

📜 Remember David: Log rotation = 5 minutes vs 2 hours! 🎯
Set it up NOW!

Advertisement

📱 Responsive Ad 📱