Log Management
Master Cassandra logs for debugging and troubleshooting!
📖 The Story: David's 3AM Log Hunt
3AM. Production down. David gets the call. "Cassandra is crashing!" He SSHs to the server. Looks at `/var/log/cassandra/system.log`. File is 50GB! His `tail` command hangs. Disk is 99% full - ALL logs! Log rotation never configured. Takes 2 hours to find the issue. $200K lost. All because logs weren't managed.
😱 The Midnight Crisis
3:00 AM - The Call:
3:05 AM - The Horror:
3:10 AM - Searching for Needle in Haystack:
3:30 AM - The Struggle:
- 💥 Log files too big to read
- 💥 grep/tail commands hang
- 💥 Disk 99% full = Cassandra can't write
- 💥 No log rotation configured
- 💥 No idea WHEN the issue started
- 💥 2 hours of trial and error
5:00 AM - Finally Found It:
Total Damage:
- 💰 $200K: Lost revenue (2 hours)
- ⏰ 2 hours: To find simple OOM error
- 😫 David: Exhausted, frustrated
- 📊 Boss: "Why did this take so long?"
✅ With Proper Log Management
What David Should Have Had:
3:05 AM - Quick Debug:
The Better Outcome:
- ✅ 5 minutes: Found the issue
- ✅ 15 minutes: Fixed and restarted
- ✅ $10K: Lost revenue (vs $200K)
- ✅ David: Back to sleep by 3:30 AM
- ✅ Boss: "Great response time!"
Log rotation: The difference between 2 hours and 5 minutes! 🎯
📂 Cassandra Log Types
Know your logs!
system.log
Main Cassandra log
Contains:
- Startup messages
- Error messages
- Warnings
- Cluster events
- Node state changes
Check first!
Your starting point!
debug.log
Detailed debugging info
Contains:
- DEBUG level messages
- Query details
- Compaction details
- Gossip details
- Very verbose
Size: Very large!
Deep debugging!
gc.log
JVM garbage collection
Contains:
- GC events
- Pause times
- Heap usage
- Memory stats
- Performance data
For: Performance tuning
JVM health!
output.log
Standard output/error
Contains:
- stdout messages
- stderr messages
- Early startup errors
- JVM crashes
- System errors
Check: If won't start
Startup issues!
commitlog
Write-ahead log
Contains:
- All writes
- Before memtable
- For crash recovery
- Binary format
- Not human-readable
Location: Different dir
Data safety!
audit logs
Security auditing
Contains:
- All queries
- User actions
- Schema changes
- Login attempts
- Compliance data
Optional: Must enable
Security!
📍 Log File Locations
Where to find logs!
Default Log Locations
How to Find Log Files
Method 1: Check cassandra-env.sh
Method 2: Check logback.xml
Method 3: Use find Command
Check Disk Space!
Always monitor log directory disk usage:
🎚️ Log Levels
Control logging verbosity!
| Level | Description | When to Use |
|---|---|---|
| TRACE | Every tiny detail | Almost never (too verbose) |
| DEBUG | Detailed debugging info | Troubleshooting specific issues |
| INFO | Important events | Production default |
| WARN | Warning messages | Always log |
| ERROR | Error conditions | Always log |
Changing Log Levels
Runtime (Temporary)
No restart required!
Permanent (Config File)
Edit logback.xml
WARNING: DEBUG in Production
DEBUG level creates HUGE log files!
- ⚠️ Size: Can grow 10x-100x faster
- ⚠️ Performance: Impacts performance
- ⚠️ Disk: Can fill disk quickly
- ✅ Use: Only for specific troubleshooting
- ✅ Duration: Enable temporarily, then disable
📖 Reading and Analyzing Logs
Essential commands!
1. View Recent Log Entries
2. Search for Errors
3. Filter by Time Range
4. Search Across Multiple Files
5. Common Log Patterns
Log Analysis Scripts
🔄 Log Rotation Configuration
Prevent David's nightmare!
Why Log Rotation?
- ✅ Prevent disk from filling up (99% = crash!)
- ✅ Keep logs manageable (50GB = unusable)
- ✅ Faster to read/search (small files = fast grep)
- ✅ Automatic cleanup (old logs deleted)
- ✅ Compression (save disk space)
Configure Log Rotation (logback.xml)
Rotation Settings Explained
| Setting | Recommended | Why |
|---|---|---|
| maxFileSize | 100MB | Balance between file size and number of files |
| maxHistory | 7 days | Keep one week for debugging |
| totalSizeCap | 1-5GB | Hard limit on total log size |
| Compression | *.gz | Save 80-90% disk space |
Verify Log Rotation Working
🔍 Log-Based Troubleshooting
Find issues fast!
Issue 1: Node Won't Start
Issue 2: Performance Degradation
Issue 3: Cluster Instability
Issue 4: Out of Memory
💡 Log Management Best Practices
Don't be David!
DO
- Configure log rotation
- Monitor disk space
- Set reasonable retention (7 days)
- Compress old logs
- Use INFO level in prod
- Ship logs to centralized system
- Set up log alerts
- Regular log reviews
DON'T
- Skip log rotation setup
- Use DEBUG in production
- Let logs fill disk
- Delete logs manually
- Ignore WARNING messages
- Keep logs forever
- Skip monitoring disk
- Wait for logs to cause issues
Complete Log Management Checklist
| Task | How To | Done? |
|---|---|---|
| Log Rotation | Configure logback.xml (maxFile Size=100MB, maxHistory=7) | ☐ |
| Disk Monitoring | Alert if /var/log > 80% full | ☐ |
| Log Level | INFO in production (DEBUG only for troubleshooting) | ☐ |
| Compression | Enable .gz compression for rotated logs | ☐ |
| Retention | Keep 7 days (adjust for compliance) | ☐ |
| Centralized Logging | Ship to ELK, Splunk, or CloudWatch | ☐ |
| Error Alerts | Alert on ERROR/WARN patterns | ☐ |
| GC Monitoring | Alert on GC pauses > 1 second | ☐ |
Log Shipping to Central System
🎉 Master Log Management!
You now know how to manage Cassandra logs like a pro!
🎓 What You Learned:
- 📖 David's nightmare: 90GB logs, 2 hours to debug
- 📂 Log types: system, debug, gc, output, audit
- 📍 Locations: /var/log/cassandra/
- 🎚️ Log levels: INFO for prod, DEBUG for troubleshooting
- 📖 Reading logs: tail, grep, search patterns
- 🔄 Log rotation: 100MB max, 7 days retention
- 🔍 Troubleshooting: Find issues in logs
- 💡 Best practices: Rotate, monitor, centralize
💡 Key Takeaways:
- Configure log rotation - First thing after install!
- Monitor disk space - Alert at 80% full
- Keep logs manageable - 100MB files, 7 days retention
- INFO level in production - DEBUG only temporarily
- Centralize logs - Ship to ELK/Splunk/CloudWatch
- Know your commands - tail, grep, awk are your friends
📋 Quick Log Rotation Setup:
🔍 Essential Debugging Commands:
📜 Remember David: Log rotation = 5 minutes vs 2 hours! 🎯
Set it up NOW!
📱 Responsive Ad 📱