Proactive Operations

Monitoring & Alerts

Know about problems BEFORE users complain!

๐Ÿ“– The Story: Emily's Silent Disaster

Emily ran a 12-node Cassandra cluster. No monitoring. One node started having issues - disk 95% full, GC pauses hitting 10 seconds, CPU maxed out. Emily had NO IDEA. Eventually the node died. Then cascade failure - other nodes overloaded. Entire cluster down. 3 hours of downtime. $400K lost. All because she had no monitoring.

๐Ÿ˜ฑ The Silent Catastrophe

Timeline of Missed Warning Signs:

-- Day 1: First warning signs (Emily had no idea) node5: disk 85% full node5: GC pauses 2-3 seconds node5: CPU 70% -- Day 2: Getting worse (still no alerts) node5: disk 92% full node5: GC pauses 5-8 seconds node5: CPU 85% node5: Read latency P99 500ms (was 10ms) -- Day 3: Critical (Emily still unaware!) node5: disk 95% full โ† CRITICAL! node5: GC pauses 10+ seconds โ† CRITICAL! node5: CPU 95% โ† CRITICAL! node5: Dropping messages -- Day 4: CATASTROPHE 10:30 AM: node5 crashes (disk full + OOM) 10:35 AM: Other nodes get overloaded 10:40 AM: Cascade failure - nodes 2, 7, 9 down 10:45 AM: ENTIRE CLUSTER DOWN ๐Ÿ’ฅ -- Emily finds out from USERS!

How Emily Discovered the Problem:

  • ๐Ÿ“ฑ 10:35 AM: Support ticket flood (100+ tickets in 5 min)
  • ๐Ÿ“ž 10:36 AM: CEO calls "Our site is down!"
  • ๐Ÿ˜ฑ 10:37 AM: Emily checks - entire cluster unreachable
  • ๐Ÿ” 10:40 AM: Starts investigating (no metrics!)
  • ๐Ÿ“œ 11:00 AM: Finally checks logs manually
  • ๐Ÿ’ก 11:15 AM: Discovers node5 was sick for DAYS

The Recovery Nightmare:

-- Emily's 3-hour scramble: 11:30 AM: Free disk space on node5 11:45 AM: Restart node5 (won't start - OOM) 12:00 PM: Increase heap size 12:15 PM: Finally node5 starts 12:30 PM: Restart nodes 2, 7, 9 12:45 PM: Cluster slowly recovering 1:00 PM: Run repairs 1:30 PM: Back online (finally!) Total downtime: 3 hours! ๐Ÿ’ฅ

The Damage:

  • ๐Ÿ’ฐ $400K: Lost revenue (3 hours downtime)
  • ๐Ÿ˜ก 10,000: Angry customers
  • ๐Ÿ“ฐ Press: "Major Outage Affects Thousands"
  • ๐Ÿ“‰ Stock: Down 8%
  • ๐Ÿ’ผ Emily: "How did you not KNOW about this?!"

โœ… With Proper Monitoring

What Should Have Happened:

-- Day 1: Alert triggers automatically 9:15 AM: โš ๏ธ ALERT: node5 disk > 80% Email + Slack + PagerDuty sent Emily sees alert immediately 9:30 AM: Emily investigates Checks Grafana dashboard Sees disk usage trend 9:45 AM: Emily cleans up old snapshots Disk drops to 60% Alert clears Crisis averted! โœ… -- Alternative: Caught GC issue 2:30 PM: โš ๏ธ ALERT: node5 GC pause > 1s Emily sees alert 2:45 PM: Checks heap usage in Grafana Realizes heap too small 3:00 PM: Increases heap + restart Problem solved! โœ… -- Result: Zero downtime! # Users never noticed anything # Emily fixed issues proactively # Boss: "Great operational excellence!"

The Better Outcome:

  • โœ… Proactive: Fixed issue before users affected
  • โœ… 30 minutes: Resolution time (vs 3 hours)
  • โœ… $0 lost: Zero downtime
  • โœ… 0 complaints: Users didn't notice
  • โœ… Emily promoted: "Excellent monitoring!"

Monitoring cost: $0. Downtime avoided: $400K! ๐ŸŽฏ

๐ŸŽฏ Why Monitor Cassandra?

The business case for observability!

๐Ÿšจ

Early Detection

  • Catch issues BEFORE outage
  • Alert on degrading performance
  • Spot trends early
  • Fix proactively
  • Prevent cascade failures

Stop problems early!

๐Ÿ“Š

Performance Visibility

  • Track latency trends
  • Monitor throughput
  • Identify bottlenecks
  • Capacity planning
  • Optimize queries

Know your system!

โฑ๏ธ

Faster MTTR

  • See problem immediately
  • Historical data for debugging
  • Correlate events
  • Root cause faster
  • Reduce downtime

Debug faster!

๐Ÿ’ฐ

Cost Optimization

  • Right-size resources
  • Avoid over-provisioning
  • Spot waste
  • Optimize compaction
  • Reduce cloud costs

Save money!

๐Ÿ“ˆ

Capacity Planning

  • Growth trends
  • Predict needs
  • Plan scaling
  • Avoid surprises
  • Budget accurately

Plan ahead!

โš–๏ธ

SLA Compliance

  • Prove uptime
  • Track SLI/SLO
  • Report to stakeholders
  • Demonstrate reliability
  • Avoid penalties

Meet commitments!

Cost of NO Monitoring

Scenario Without Monitoring With Monitoring
Disk Full Find out when cluster crashes Alert at 80%, clean up proactively
Memory Leak 3-hour outage Alert on heap growth, fix before crash
Performance Degradation Users complain See P99 spike, investigate immediately
Node Failure Cascade failure Alert when node down, handle gracefully

๐Ÿ“Š Key Metrics to Monitor

The essentials!

1. Cluster Health

-- Node Status: โœ… All nodes UP/NORMAL โœ… No nodes DN (down) โœ… No nodes UNโ†’DN flapping -- How to check: nodetool status -- Alert on: # Any node not in UN state # Node state changes (flapping)

2. Read/Write Latency

-- Critical metrics: P50 (median) latency P95 latency P99 latency โ† Most important! P999 latency -- Good targets: Read P99 < 10ms Write P99 < 5ms -- Alert on: # P99 > 50ms (reads) # P99 > 20ms (writes) # Sudden spikes

3. Throughput

-- Track: Reads/second Writes/second Total operations/second -- Alert on: # Sudden drop (indicates issues) # Approaching capacity

4. GC Metrics

-- Monitor: GC pause time GC frequency Heap usage -- Good targets: GC pause < 100ms Young GC < 50ms Full GC: NEVER -- Alert on: # GC pause > 1 second # Heap usage > 80% # Any Full GC events

5. Disk Usage

-- Track: Disk space used/available Disk I/O utilization SSTable count -- Alert on: # Disk > 80% full # Disk I/O > 80% # Rapid growth

6. Compaction

-- Monitor: Pending compactions Compaction throughput SSTable count growth -- Alert on: # Pending compactions > 100 # Stuck compactions

7. Dropped Messages

-- Track: Dropped READ messages Dropped WRITE messages Dropped MUTATION messages -- Alert on: # Any dropped messages # Indicates overload or timeout

8. Thread Pool Stats

-- Monitor pools: ReadStage (active/pending) MutationStage (active/pending) CompactionExecutor (active/pending) -- Alert on: # High pending tasks # Blocked threads

9. Error Rates

-- Track: Timeouts Unavailable exceptions Read/write failures -- Alert on: # Error rate > 1% # Sudden spike

10. Connection Stats

-- Monitor: Active connections Connection errors Connection pool saturation -- Alert on: # Connection failures # Pool exhaustion

๐Ÿ› ๏ธ The Monitoring Stack

Popular tools!

๐Ÿ“ˆ

Prometheus

Metrics collection & storage

Features:

  • Pull-based scraping
  • Time-series database
  • PromQL query language
  • Alerting rules
  • Service discovery

Best for: Metrics

Industry standard!

๐Ÿ“‰

Grafana

Visualization & dashboards

Features:

  • Beautiful dashboards
  • Multi-datasource
  • Templating
  • Annotations
  • Alerting

Best for: Visualization

Perfect UI!

๐Ÿ””

Alertmanager

Alert routing & deduplication

Features:

  • Alert grouping
  • Deduplication
  • Silencing
  • Routing rules
  • Multi-channel

Best for: Alerts

Smart routing!

๐Ÿ“Š

Cassandra Exporter

Metrics export for Prometheus

Features:

  • JMX metrics export
  • nodetool metrics
  • Custom metrics
  • Low overhead
  • Easy setup

Best for: Cassandraโ†’Prometheus

Bridge tool!

๐Ÿ“ฆ

Alternative: DataDog

All-in-one SaaS

Features:

  • Full stack monitoring
  • APM integration
  • Log aggregation
  • Built-in dashboards
  • Easy setup

Best for: Quick start

$$$ but easy!

โ˜๏ธ

Alternative: CloudWatch

AWS native monitoring

Features:

  • AWS integration
  • Custom metrics
  • Alarms
  • Dashboards
  • Logs Insights

Best for: AWS deployments

AWS users!

๐Ÿ“ˆ Prometheus + Cassandra Exporter Setup

Complete setup guide!

1

Install Cassandra Exporter

Export JMX metrics to Prometheus format

-- Download Cassandra Exporter: $ cd /opt $ wget https://github.com/criteo/cassandra_exporter/releases/download/v2.3.8/cassandra_exporter-2.3.8-all.jar -- Create systemd service: $ sudo vim /etc/systemd/system/cassandra-exporter.service [Unit] Description=Cassandra Exporter After=cassandra.service [Service] Type=simple User=cassandra ExecStart=/usr/bin/java -jar \ /opt/cassandra_exporter-2.3.8-all.jar \ --host 127.0.0.1:7199 Restart=always [Install] WantedBy=multi-user.target -- Start exporter: $ sudo systemctl daemon-reload $ sudo systemctl start cassandra-exporter $ sudo systemctl enable cassandra-exporter -- Test it: $ curl localhost:8080/metrics | grep cassandra /* Should see metrics like: cassandra_read_latency_seconds{quantile="0.99"} 0.005 cassandra_write_latency_seconds{quantile="0.99"} 0.002 */
2

Install Prometheus

Metrics collection server

-- Install Prometheus: $ cd /opt $ wget https://github.com/prometheus/prometheus/releases/download/v2.45.0/prometheus-2.45.0.linux-amd64.tar.gz $ tar xvf prometheus-2.45.0.linux-amd64.tar.gz $ cd prometheus-2.45.0.linux-amd64 -- Configure prometheus.yml: $ vim prometheus.yml global: scrape_interval: 15s evaluation_interval: 15s scrape_configs: - job_name: 'cassandra' static_configs: - targets: - 'node1:8080' - 'node2:8080' - 'node3:8080' labels: cluster: 'production' -- Start Prometheus: $ ./prometheus --config.file=prometheus.yml -- Access UI: $ open http://localhost:9090
3

Configure Alert Rules

Define critical alerts

-- Create alert rules file: $ vim /opt/prometheus/cassandra_alerts.yml groups: - name: cassandra interval: 30s rules: # Node down - alert: CassandraNodeDown expr: up{job="cassandra"} == 0 for: 1m labels: severity: critical annotations: summary: "Cassandra node {{ $labels.instance }} is down" # High latency - alert: HighReadLatency expr: cassandra_read_latency_seconds{quantile="0.99"} > 0.05 for: 5m labels: severity: warning annotations: summary: "High read latency on {{ $labels.instance }}" # Disk space - alert: DiskSpaceLow expr: (node_filesystem_avail_bytes / node_filesystem_size_bytes) < 0.2 for: 5m labels: severity: warning annotations: summary: "Disk space < 20% on {{ $labels.instance }}" # GC pauses - alert: LongGCPauses expr: cassandra_gc_pause_seconds{quantile="0.99"} > 1 for: 5m labels: severity: critical annotations: summary: "GC pauses > 1s on {{ $labels.instance }}" -- Update prometheus.yml: rule_files: - 'cassandra_alerts.yml' -- Restart Prometheus

๐Ÿ“‰ Grafana Dashboards

Beautiful visualizations!

1

Install Grafana

Dashboard platform

-- Install on Ubuntu: $ sudo apt-get install -y software-properties-common $ sudo add-apt-repository "deb https://packages.grafana.com/oss/deb stable main" $ wget -q -O - https://packages.grafana.com/gpg.key | sudo apt-key add - $ sudo apt-get update $ sudo apt-get install grafana -- Start Grafana: $ sudo systemctl start grafana-server $ sudo systemctl enable grafana-server -- Access UI: # URL: http://localhost:3000 # Default login: admin/admin
2

Add Prometheus Datasource

Connect to Prometheus

-- In Grafana UI: 1. Go to Configuration โ†’ Data Sources 2. Click "Add data source" 3. Select "Prometheus" 4. Configure: URL: http://localhost:9090 Access: Server (default) 5. Click "Save & Test" -- Should see: "Data source is working"
3

Import Cassandra Dashboard

Pre-built dashboard

-- Option 1: Import from Grafana.com 1. Go to Dashboards โ†’ Import 2. Enter Dashboard ID: 11862 3. Click Load 4. Select Prometheus datasource 5. Click Import -- Option 2: Create custom dashboard 1. Click "+ Create Dashboard" 2. Add panels for key metrics 3. Use PromQL queries -- Example queries: # Read latency P99: cassandra_read_latency_seconds{quantile="0.99"} # Throughput: rate(cassandra_read_total[5m]) # Disk usage: (1 - (node_filesystem_avail_bytes / node_filesystem_size_bytes)) * 100

Essential Dashboard Panels

Panel Query Type
Cluster Status up{job="cassandra"} Stat
Read Latency P99 cassandra_read_latency_seconds{quantile="0.99"} Graph
Write Latency P99 cassandra_write_latency_seconds{quantile="0.99"} Graph
Read Throughput rate(cassandra_read_total[5m]) Graph
Write Throughput rate(cassandra_write_total[5m]) Graph
GC Pause Time cassandra_gc_pause_seconds{quantile="0.99"} Graph
Disk Usage (1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) * 100 Gauge
Pending Compactions cassandra_compaction_pending Stat
Dropped Messages rate(cassandra_dropped_messages_total[5m]) Graph
Heap Usage (cassandra_jvm_memory_used_bytes / cassandra_jvm_memory_max_bytes) * 100 Gauge

๐Ÿšจ Critical Alerts to Configure

Must-have alerts!

1. Node Down (P0 - Critical)

-- Alert when node is down: - alert: CassandraNodeDown expr: up{job="cassandra"} == 0 for: 1m labels: severity: critical annotations: summary: "Node {{ $labels.instance }} is DOWN" description: "Immediate attention required!" -- Send to: PagerDuty, Slack, Email -- Response: Investigate immediately

2. High Latency (P1 - High)

-- Alert on P99 latency spike: - alert: HighReadLatency expr: cassandra_read_latency_seconds{quantile="0.99"} > 0.05 for: 5m labels: severity: warning annotations: summary: "High read latency: {{ $value }}s" - alert: HighWriteLatency expr: cassandra_write_latency_seconds{quantile="0.99"} > 0.02 for: 5m labels: severity: warning -- Response: Investigate query patterns, check GC

3. Disk Space Low (P1 - High)

-- Alert at 80% full: - alert: DiskSpaceLow expr: (node_filesystem_avail_bytes / node_filesystem_size_bytes) < 0.2 for: 5m labels: severity: warning annotations: summary: "Disk < 20% free on {{ $labels.instance }}" -- Critical at 90%: - alert: DiskSpaceCritical expr: (node_filesystem_avail_bytes / node_filesystem_size_bytes) < 0.1 for: 1m labels: severity: critical -- Response: Clean snapshots, add disk, scale out

4. Long GC Pauses (P0 - Critical)

-- Alert on GC > 1 second: - alert: LongGCPauses expr: cassandra_gc_pause_seconds{quantile="0.99"} > 1 for: 5m labels: severity: critical annotations: summary: "GC pauses > 1s: {{ $value }}s" -- Response: Check heap usage, tune GC, increase heap

5. Dropped Messages (P1 - High)

-- Alert on any dropped messages: - alert: DroppedMessages expr: rate(cassandra_dropped_messages_total[5m]) > 0 for: 5m labels: severity: warning annotations: summary: "Dropping messages on {{ $labels.instance }}" -- Response: Check overload, timeouts, thread pools

6. Heap Usage High (P1 - High)

-- Alert when heap > 80%: - alert: HeapUsageHigh expr: (cassandra_jvm_memory_used_bytes / cassandra_jvm_memory_max_bytes) > 0.8 for: 10m labels: severity: warning -- Response: Investigate memory leak, increase heap

7. Pending Compactions (P2 - Medium)

-- Alert when compactions backing up: - alert: PendingCompactions expr: cassandra_compaction_pending > 100 for: 30m labels: severity: warning -- Response: Check compaction throughput, tune settings

8. High Error Rate (P1 - High)

-- Alert on error spike: - alert: HighErrorRate expr: rate(cassandra_errors_total[5m]) / rate(cassandra_requests_total[5m]) > 0.01 for: 5m labels: severity: critical -- Response: Check logs, investigate errors

Alert Configuration Example (Alertmanager)

-- /opt/prometheus/alertmanager.yml: global: resolve_timeout: 5m route: group_by: ['alertname', 'cluster'] group_wait: 10s group_interval: 10s repeat_interval: 12h receiver: 'default' routes: # Critical alerts to PagerDuty - match: severity: critical receiver: 'pagerduty' continue: true # All alerts to Slack - receiver: 'slack' receivers: - name: 'default' email_configs: - to: 'ops@company.com' from: 'alerts@company.com' smarthost: 'smtp.company.com:587' - name: 'pagerduty' pagerduty_configs: - service_key: '' - name: 'slack' slack_configs: - api_url: 'https://hooks.slack.com/services/...' channel: '#cassandra-alerts' title: '{{ .GroupLabels.alertname }}' text: '{{ range .Alerts }}{{ .Annotations.summary }} {{ end }}'

๐Ÿ’ก Monitoring Best Practices

Do it right!

โœ…

DO

  • Monitor ALL nodes
  • Set up critical alerts
  • Use dashboards daily
  • Test alert routing
  • Review metrics weekly
  • Set realistic thresholds
  • Document runbooks
  • Automate responses
โŒ

DON'T

  • Skip monitoring setup
  • Set too many alerts
  • Ignore warning alerts
  • Use default thresholds
  • Alert fatigue
  • Monitor without action
  • Forget to test alerts
  • No alert documentation

Complete Monitoring Checklist

Component Setup Done?
Metrics Collection Prometheus + Cassandra Exporter on all nodes โ˜
Visualization Grafana dashboard with key metrics โ˜
Critical Alerts Node down, high latency, disk spaceโ˜
Alert Routing Alertmanager with PagerDuty + Slack โ˜
GC Monitoring GC pause alerts + heap usage โ˜
Disk Monitoring Usage alerts at 80% and 90% โ˜
Performance Latency P99, throughput tracking โ˜
Testing Test all alert routes monthly โ˜
Documentation Runbooks for each alert type โ˜
Review Weekly dashboard review โ˜

Alert Threshold Guidelines

Metric Warning Critical
Disk Space > 80% full > 90% full
Read Latency P99 > 50ms > 100ms
Write Latency P99 > 20ms > 50ms
GC Pause > 500ms > 1s
Heap Usage > 80% > 90%
Pending Compactions > 100 > 500
Dropped Messages Any > 100/sec
Error Rate > 1% > 5%

๐ŸŽ‰ Master Cassandra Monitoring!

You now know how to monitor like a pro!

๐ŸŽ“ What You Learned:

  • ๐Ÿ“– Emily's disaster: Silent failure = $400K lost
  • ๐ŸŽฏ Why monitor: Early detection, faster MTTR
  • ๐Ÿ“Š Key metrics: Latency, throughput, GC, disk
  • ๐Ÿ› ๏ธ The stack: Prometheus + Grafana + Alertmanager
  • ๐Ÿ“ˆ Setup: Complete Prometheus configuration
  • ๐Ÿ“‰ Dashboards: Grafana with all key panels
  • ๐Ÿšจ Alerts: 8 critical alerts configured
  • ๐Ÿ’ก Best practices: Test, document, review

๐Ÿ’ก Key Takeaways:

  1. Monitor proactively - Find issues before users
  2. Set up alerts - Node down, latency, disk, GC
  3. Use dashboards - Visualize everything
  4. Test alerts - Make sure they work!
  5. Document runbooks - How to respond
  6. Review regularly - Weekly dashboard checks

๐Ÿ“‹ Quick Setup (30 Minutes):

# 1. Install Cassandra Exporter (5 min) Download + configure on each node# 2. Install Prometheus (5 min) Download + configure scraping# 3. Configure Alerts (10 min) Create alert rules file# 4. Install Grafana (5 min) Download + start service# 5. Create Dashboard (5 min) Import dashboard 11862 OR create custom# Done! You're monitoring! โœ…

๐Ÿšจ Must-Have Alerts:

  1. Node Down โ†’ PagerDuty (P0)
  2. High Latency โ†’ Slack (P1)
  3. Disk > 80% โ†’ Email (P1)
  4. GC > 1s โ†’ PagerDuty (P0)
  5. Dropped Messages โ†’ Slack (P1)
  6. Heap > 80% โ†’ Email (P1)
  7. High Errors โ†’ Slack (P1)
  8. Pending Compactions โ†’ Email (P2)

๐Ÿ“Š Remember Emily: Monitoring = Know about problems BEFORE users! ๐ŸŽฏ
Set it up TODAY!

Advertisement

๐Ÿ“ฑ Responsive Ad ๐Ÿ“ฑ