Monitoring & Alerts
Know about problems BEFORE users complain!
๐ The Story: Emily's Silent Disaster
Emily ran a 12-node Cassandra cluster. No monitoring. One node started having issues - disk 95% full, GC pauses hitting 10 seconds, CPU maxed out. Emily had NO IDEA. Eventually the node died. Then cascade failure - other nodes overloaded. Entire cluster down. 3 hours of downtime. $400K lost. All because she had no monitoring.
๐ฑ The Silent Catastrophe
Timeline of Missed Warning Signs:
How Emily Discovered the Problem:
- ๐ฑ 10:35 AM: Support ticket flood (100+ tickets in 5 min)
- ๐ 10:36 AM: CEO calls "Our site is down!"
- ๐ฑ 10:37 AM: Emily checks - entire cluster unreachable
- ๐ 10:40 AM: Starts investigating (no metrics!)
- ๐ 11:00 AM: Finally checks logs manually
- ๐ก 11:15 AM: Discovers node5 was sick for DAYS
The Recovery Nightmare:
The Damage:
- ๐ฐ $400K: Lost revenue (3 hours downtime)
- ๐ก 10,000: Angry customers
- ๐ฐ Press: "Major Outage Affects Thousands"
- ๐ Stock: Down 8%
- ๐ผ Emily: "How did you not KNOW about this?!"
โ With Proper Monitoring
What Should Have Happened:
The Better Outcome:
- โ Proactive: Fixed issue before users affected
- โ 30 minutes: Resolution time (vs 3 hours)
- โ $0 lost: Zero downtime
- โ 0 complaints: Users didn't notice
- โ Emily promoted: "Excellent monitoring!"
Monitoring cost: $0. Downtime avoided: $400K! ๐ฏ
๐ฏ Why Monitor Cassandra?
The business case for observability!
Early Detection
- Catch issues BEFORE outage
- Alert on degrading performance
- Spot trends early
- Fix proactively
- Prevent cascade failures
Stop problems early!
Performance Visibility
- Track latency trends
- Monitor throughput
- Identify bottlenecks
- Capacity planning
- Optimize queries
Know your system!
Faster MTTR
- See problem immediately
- Historical data for debugging
- Correlate events
- Root cause faster
- Reduce downtime
Debug faster!
Cost Optimization
- Right-size resources
- Avoid over-provisioning
- Spot waste
- Optimize compaction
- Reduce cloud costs
Save money!
Capacity Planning
- Growth trends
- Predict needs
- Plan scaling
- Avoid surprises
- Budget accurately
Plan ahead!
SLA Compliance
- Prove uptime
- Track SLI/SLO
- Report to stakeholders
- Demonstrate reliability
- Avoid penalties
Meet commitments!
Cost of NO Monitoring
| Scenario | Without Monitoring | With Monitoring |
|---|---|---|
| Disk Full | Find out when cluster crashes | Alert at 80%, clean up proactively |
| Memory Leak | 3-hour outage | Alert on heap growth, fix before crash |
| Performance Degradation | Users complain | See P99 spike, investigate immediately |
| Node Failure | Cascade failure | Alert when node down, handle gracefully |
๐ Key Metrics to Monitor
The essentials!
1. Cluster Health
2. Read/Write Latency
3. Throughput
4. GC Metrics
5. Disk Usage
6. Compaction
7. Dropped Messages
8. Thread Pool Stats
9. Error Rates
10. Connection Stats
๐ ๏ธ The Monitoring Stack
Popular tools!
Prometheus
Metrics collection & storage
Features:
- Pull-based scraping
- Time-series database
- PromQL query language
- Alerting rules
- Service discovery
Best for: Metrics
Industry standard!
Grafana
Visualization & dashboards
Features:
- Beautiful dashboards
- Multi-datasource
- Templating
- Annotations
- Alerting
Best for: Visualization
Perfect UI!
Alertmanager
Alert routing & deduplication
Features:
- Alert grouping
- Deduplication
- Silencing
- Routing rules
- Multi-channel
Best for: Alerts
Smart routing!
Cassandra Exporter
Metrics export for Prometheus
Features:
- JMX metrics export
- nodetool metrics
- Custom metrics
- Low overhead
- Easy setup
Best for: CassandraโPrometheus
Bridge tool!
Alternative: DataDog
All-in-one SaaS
Features:
- Full stack monitoring
- APM integration
- Log aggregation
- Built-in dashboards
- Easy setup
Best for: Quick start
$$$ but easy!
Alternative: CloudWatch
AWS native monitoring
Features:
- AWS integration
- Custom metrics
- Alarms
- Dashboards
- Logs Insights
Best for: AWS deployments
AWS users!
๐ Prometheus + Cassandra Exporter Setup
Complete setup guide!
Install Cassandra Exporter
Export JMX metrics to Prometheus format
Install Prometheus
Metrics collection server
Configure Alert Rules
Define critical alerts
๐ Grafana Dashboards
Beautiful visualizations!
Install Grafana
Dashboard platform
Add Prometheus Datasource
Connect to Prometheus
Import Cassandra Dashboard
Pre-built dashboard
Essential Dashboard Panels
| Panel | Query | Type |
|---|---|---|
| Cluster Status | up{job="cassandra"} | Stat |
| Read Latency P99 | cassandra_read_latency_seconds{quantile="0.99"} | Graph |
| Write Latency P99 | cassandra_write_latency_seconds{quantile="0.99"} | Graph |
| Read Throughput | rate(cassandra_read_total[5m]) | Graph |
| Write Throughput | rate(cassandra_write_total[5m]) | Graph |
| GC Pause Time | cassandra_gc_pause_seconds{quantile="0.99"} | Graph |
| Disk Usage | (1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) * 100 | Gauge |
| Pending Compactions | cassandra_compaction_pending | Stat |
| Dropped Messages | rate(cassandra_dropped_messages_total[5m]) | Graph |
| Heap Usage | (cassandra_jvm_memory_used_bytes / cassandra_jvm_memory_max_bytes) * 100 | Gauge |
๐จ Critical Alerts to Configure
Must-have alerts!
1. Node Down (P0 - Critical)
2. High Latency (P1 - High)
3. Disk Space Low (P1 - High)
4. Long GC Pauses (P0 - Critical)
5. Dropped Messages (P1 - High)
6. Heap Usage High (P1 - High)
7. Pending Compactions (P2 - Medium)
8. High Error Rate (P1 - High)
Alert Configuration Example (Alertmanager)
๐ก Monitoring Best Practices
Do it right!
DO
- Monitor ALL nodes
- Set up critical alerts
- Use dashboards daily
- Test alert routing
- Review metrics weekly
- Set realistic thresholds
- Document runbooks
- Automate responses
DON'T
- Skip monitoring setup
- Set too many alerts
- Ignore warning alerts
- Use default thresholds
- Alert fatigue
- Monitor without action
- Forget to test alerts
- No alert documentation
Complete Monitoring Checklist
| Component | Setup | Done? |
|---|---|---|
| Metrics Collection | Prometheus + Cassandra Exporter on all nodes | โ |
| Visualization | Grafana dashboard with key metrics | โ |
| Critical Alerts | Node down, high latency, disk spaceโ | |
| Alert Routing | Alertmanager with PagerDuty + Slack | โ |
| GC Monitoring | GC pause alerts + heap usage | โ |
| Disk Monitoring | Usage alerts at 80% and 90% | โ |
| Performance | Latency P99, throughput tracking | โ |
| Testing | Test all alert routes monthly | โ |
| Documentation | Runbooks for each alert type | โ |
| Review | Weekly dashboard review | โ |
Alert Threshold Guidelines
| Metric | Warning | Critical |
|---|---|---|
| Disk Space | > 80% full | > 90% full |
| Read Latency P99 | > 50ms | > 100ms |
| Write Latency P99 | > 20ms | > 50ms |
| GC Pause | > 500ms | > 1s |
| Heap Usage | > 80% | > 90% |
| Pending Compactions | > 100 | > 500 |
| Dropped Messages | Any | > 100/sec |
| Error Rate | > 1% | > 5% |
๐ Master Cassandra Monitoring!
You now know how to monitor like a pro!
๐ What You Learned:
- ๐ Emily's disaster: Silent failure = $400K lost
- ๐ฏ Why monitor: Early detection, faster MTTR
- ๐ Key metrics: Latency, throughput, GC, disk
- ๐ ๏ธ The stack: Prometheus + Grafana + Alertmanager
- ๐ Setup: Complete Prometheus configuration
- ๐ Dashboards: Grafana with all key panels
- ๐จ Alerts: 8 critical alerts configured
- ๐ก Best practices: Test, document, review
๐ก Key Takeaways:
- Monitor proactively - Find issues before users
- Set up alerts - Node down, latency, disk, GC
- Use dashboards - Visualize everything
- Test alerts - Make sure they work!
- Document runbooks - How to respond
- Review regularly - Weekly dashboard checks
๐ Quick Setup (30 Minutes):
๐จ Must-Have Alerts:
- Node Down โ PagerDuty (P0)
- High Latency โ Slack (P1)
- Disk > 80% โ Email (P1)
- GC > 1s โ PagerDuty (P0)
- Dropped Messages โ Slack (P1)
- Heap > 80% โ Email (P1)
- High Errors โ Slack (P1)
- Pending Compactions โ Email (P2)
๐ Remember Emily: Monitoring = Know about problems BEFORE users! ๐ฏ
Set it up TODAY!
๐ฑ Responsive Ad ๐ฑ